Ancient book character recognition method based on topological mask and deformable strip pooling

CN122780969APending Publication Date: 2026-09-18GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611129428.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

其根本原因在于现有网络的多尺度特征调度机制无法适应古代文献特有的领域特性,主要存在以下两个致命瓶颈 :一是密集排版导致的“特征混叠”,古代汉字古籍普遍采用密集垂直的排版方式,且经常在正文旁嵌套微小的双列夹注 ,由于现有技术的标准卷积算子均呈各向同性,其方形感受野缺乏对带状、条状文本布局的几何约束 ,在提取跨尺度细小行间注释特征时,感受野不可避免地会侵入相邻的正文列或深层的纸张背景噪声中,导致下游解码器接收到的特征流受到严重污染和混淆 ;二是严重物理退化导致的“特征平滑陷阱”,由于长期保存,古代文献往往伴随着严重的虫蛀、墨水渗透、字迹褪色以及微观笔画断裂 ,如果直接采用现有的去噪平滑算法或强行在主特征流上进行掩码重构来修复残缺笔画,模型在优化过程中会本能地倾向于使用模糊的局部均值去填充笔画断裂处 ,这虽然在宏观上有利于界定连续的文本行边界,但却在微观上完全破坏和抹平了区分形近字所必需的高频锐利细节,最终导致模型陷入“定位精准但字符识别错误频繁”的技术困境

Benefits of technology

[0025]First, addressing the severe physical degradation and microscopic stroke breakage commonly found in ancient Chinese texts, the HATFM module employs an implicit "erasure-repair" training mechanism to force the network to learn and internalize the topological priors of Chinese character structures. This design successfully overcomes the "feature smoothing trap" caused by conventional denoising methods, achieving precise repair of incomplete stroke features without polluting the main feature stream, thus greatly preserving the high-frequency sharp details needed to distinguish similar-looking characters. Second, to overcome the "feature aliasing" defect that existing models easily encounter when dealing with dense vertical layouts of ancient texts, this technical solution uses the MFDSP module, which injects directional spatial priors from the layout of ancient texts. This module utilizes a combination of deformable feature alignment and orthogonal bar pooling to construct rigid geometric constraints, precisely filtering out lateral interference and multi-scale background noise from adjacent text columns, much like "blinds," thereby providing an extremely pure visual feature stream for the downstream decoder. Finally, through a unique bypass reconstruction strategy, this method completely discards the HATFM bypass branches used for training constraints when deployed for actual inference tasks, relying solely on the refined main path for prediction. This design ingeniously achieves zero additional computational cost and zero parameter overhead, enabling the model to maintain strong robustness against severe physical degradation while achieving fast and stable practical applications, ultimately resulting in a high-precision and high-efficiency end-to-end ancient book recognition solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780969A_ABST
    Figure CN122780969A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on topological mask and deformable strip pool's ancient book character recognition method, comprising the following steps: 1) preliminary visual feature extraction;2) combined with the multiscale feature fusion module of deformable strip pool;3) auxiliary training module: high activation topological feature mask module;4) query modeling and joint loss optimization;5) output recognition character;6) optimization.This method does not lose and pollute high-frequency visual details under the premise, topological structure repair of incomplete stroke is implicitly realized on feature level, while using layout direction spatial perception accurately isolates the interference of adjacent text column, suppresses background noise, and greatly improves the character recognition accuracy of severely degraded ancient books.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and pattern recognition technology, specifically a method for recognizing ancient texts based on topological masking and deformable strip pooling. Background Technology

[0002] Ancient Chinese texts are invaluable carriers of historical, cultural, and scientific heritage. In the digital humanities era, large-scale digitization of these historical archives is crucial for the preservation, retrieval, and in-depth text analysis of cultural relics. However, traditional fully manual transcription methods are extremely time-consuming and labor-intensive. Therefore, automated page-level text recognition technology has become a fundamental and highly sought-after means of unlocking the knowledge contained within ancient classics. In recent years, existing page-level recognition methods can be broadly categorized into two types: segmentation-based methods and segmentation-free or implicit segmentation methods. Among existing representative technologies, segmentation-based methods often handle text localization and recognition through decoupling or joint optimization. The paper "A computationally efficient pipeline approach to full-page offline handwritten text recognition" (published in Proceedings of International Conference on Document Analysis and Recognition Workshops) constructs independent text localization and transcription modules respectively. In view of the special characteristics of complex ancient Chinese texts, the paper "Joint layout analysis, character detection and recognition for historical document digitization" (published in Proceedings of international conference on frontiers inhandwriting recognition) develops a historical document digitization system that integrates layout analysis, character detection and recognition.In segmentation-free or implicit segmentation methods, researchers tend to directly read full-page features. The paper "OrigamiNet: Weakly-supervised, segmentation-free, one-step, full-page text recognition by learning to unfold" (published in *Proceedings of IEEE conference on computer vision and pattern recognition*) achieves weakly supervised, segmentation-free, one-step full-page text recognition through a feature unfolding mechanism. Among recently emerging frameworks, the paper "Unitext: A unified framework for Chinese text detection, recognition, and restoration in ancient document and inscription images" (published in *Applied Sciences*) proposes a multi-task unified framework that attempts to perform glyph restoration concurrently with text detection and recognition.

[0003] Despite the progress made by the aforementioned existing technologies, overall performance still suffers a significant decline when faced with the dual challenges of "severe physical degradation" and "irregular nested layouts" in the real world. The fundamental reason lies in the inability of existing multi-scale feature scheduling mechanisms to adapt to the unique domain characteristics of ancient texts, primarily exhibiting two fatal bottlenecks: First, "feature aliasing" caused by dense typesetting. Ancient Chinese texts generally employ dense vertical typesetting, often nesting tiny double-column annotations alongside the main text. Since the standard convolution operators in existing technologies are isotropic, their square receptive fields lack geometric constraints on strip-shaped and bar-shaped text layouts. When extracting cross-scale fine interline annotation features, the receptive field inevitably intrudes into adjacent text columns or deep paper background noise, resulting in severe contamination of the feature stream received by the downstream decoder. The first problem is confusion; the second is the "feature smoothing trap" caused by severe physical degradation. Due to long-term preservation, ancient documents often suffer from severe insect damage, ink penetration, fading of characters, and microscopic stroke breaks. If existing denoising and smoothing algorithms are directly used or mask reconstruction is forcibly performed on the main feature flow to repair the missing strokes, the model will instinctively tend to use fuzzy local means to fill in the stroke breaks during the optimization process. Although this is beneficial for defining the boundaries of continuous text lines macroscopically, it completely destroys and smooths out the high-frequency sharp details necessary for distinguishing similar-looking characters microscopically, ultimately leading the model into the technical dilemma of "accurate positioning but frequent character recognition errors." Therefore, it is necessary to invent a recognition method that can effectively adapt to the complex layout and degradation features of ancient books. Summary of the Invention

[0004] The purpose of this invention is to address the problem of recognizing ancient Chinese characters in highly degraded and densely formatted texts. It proposes a method for recognizing ancient text characters based on topological masking and deformable strip pooling. This method provides a highly robust single-stage end-to-end text recognition framework. It aims to implicitly restore the topological structure of incomplete strokes at the feature level without losing or contaminating high-frequency visual details. Simultaneously, it utilizes spatial awareness of the layout direction to accurately isolate interference from adjacent text columns, suppressing background noise and significantly improving the character recognition accuracy of severely degraded ancient texts.

[0005] The technical solution to achieve the objective of this invention is:

[0006] A method for recognizing ancient texts based on topological masking and deformable strip pooling includes the following steps:

[0007] 1) Preliminary visual feature extraction,

[0008] The input image of the ancient book is fed into a convolutional neural network encoder to extract preliminary, multi-level visual feature maps F. Feature map F contains the basic structure and texture information of the characters and is used as subsequent input.

[0009] 2) Combining a multi-scale feature fusion module with deformable strip pooling,

[0010] The multi-scale feature fusion module, as the core path for backbone feature refinement, takes the input feature map F and first unifies the channel dimensions through a 1×1 horizontal convolution. Then, it uses a top-down feature pyramid with deformable bar pooling channel attention to process the low-level features of the current level. and the high-level features obtained by upsampling from the previous layer The fused features are obtained through deformable feature alignment, orthogonal strip pooling and recalibration, and top-down filtering fusion. , will integrate features As input to the transformer encoder, the output features ;

[0011] 3) Auxiliary training module: High activation topological feature mask module.

[0012] First, for feature maps The spatial activation intensity is calculated by averaging along the channel dimension. :

[0013] (1.6);

[0014] Set the base discard rate Define the dynamic discard probability With retention probability And a binary mask is generated through Bernoulli sampling;

[0015] Then pixel-level scale compensation is performed to obtain the damage features. The damaged features Input multi-scale feature fusion module to obtain Reconstruction loss calculate;

[0016] 4) Query modeling and joint loss optimization, based on features output by the transformer encoder. Initialize detection query and identification query ,Bundle Input a linear layer to generate text classification scores, select the top K features based on the text classification scores, and feed these top K features into a linear layer to initialize the detection query. The proposal is for identifying queries From the detection query Extracting features from proposals And take the average along the height dimension to initialize the identification query. ,Will Projected onto semantic probability space :

[0017] (1.10)

[0018] W is a learnable matrix, and Concatenation to generate language vectors Then combine with task query and Perform cross-modal semantic interaction and query the task after the interaction. As input to the transformer decoder;

[0019] 5) Finally, the transformer decoder outputs detection query and recognition query. The detection query enters the detection head, which consists of two feedforward layers, and outputs the position of the detection box. The recognition query enters a linear layer, namely the recognition head, and outputs the recognized character.

[0020] 6) Use of joint mission losses And the cosine similarity reconstruction loss calculated for the bypass branch of the high-activation topological feature mask module. Perform overall joint optimization:

[0021] (1.12)

[0022] (1.13)

[0023] in and These are real label classes and bounding boxes. The prediction representing the bounding box, yes Predicted probability of class , , and These are the loss weights for classification, bounding boxes, polygons, and recognition, respectively.

[0024] The specific effects of this technical solution are as follows:

[0025] First, addressing the severe physical degradation and microscopic stroke breakage commonly found in ancient Chinese texts, the HATFM module employs an implicit "erasure-repair" training mechanism to force the network to learn and internalize the topological priors of Chinese character structures. This design successfully overcomes the "feature smoothing trap" caused by conventional denoising methods, achieving precise repair of incomplete stroke features without polluting the main feature stream, thus greatly preserving the high-frequency sharp details needed to distinguish similar-looking characters. Second, to overcome the "feature aliasing" defect that existing models easily encounter when dealing with dense vertical layouts of ancient texts, this technical solution uses the MFDSP module, which injects directional spatial priors from the layout of ancient texts. This module utilizes a combination of deformable feature alignment and orthogonal bar pooling to construct rigid geometric constraints, precisely filtering out lateral interference and multi-scale background noise from adjacent text columns, much like "blinds," thereby providing an extremely pure visual feature stream for the downstream decoder. Finally, through a unique bypass reconstruction strategy, this method completely discards the HATFM bypass branches used for training constraints when deployed for actual inference tasks, relying solely on the refined main path for prediction. This design ingeniously achieves zero additional computational cost and zero parameter overhead, enabling the model to maintain strong robustness against severe physical degradation while achieving fast and stable practical applications, ultimately resulting in a high-precision and high-efficiency end-to-end ancient book recognition solution.

[0026] This technical solution implicitly restores the topological structure of incomplete strokes at the feature level without losing or polluting high-frequency visual details. At the same time, it uses spatial perception of the layout direction to accurately isolate interference from adjacent text columns, suppress background noise, and significantly improve the character recognition accuracy of severely degraded ancient books. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the overall process of the embodiment;

[0028] Figure 2 This is a schematic diagram of the internal structure of the deformable strip pooling and multi-scale feature fusion (MFDSP) module in the embodiment. Figure a shows the overall structure of MFDSP, and Figure b shows the structure of the deformable channel attention part in MFDSP.

[0029] Figure 3 This is a schematic diagram of the internal structure of the High Activation Topology Mask (HATFM) module in the embodiment;

[0030] Figure 4 The specific examples identified in the embodiments include Case A and Case B;

[0031] Figure 5 This is a line graph comparing the performance of the example with other methods in detecting text P (precision), R (recall), and F (F1 score);

[0032] Figure 6 is a line graph comparing the embodiment with other recognition methods on two key performance indicators, AR*(%) and CR*(%). Detailed Implementation

[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this is not intended to limit the scope of the invention.

[0034] Example:

[0035] Reference Figure 1 A method for recognizing ancient texts based on topological masking and deformable strip pooling includes the following steps:

[0036] 1) Preliminary visual feature extraction,

[0037] An image of an ancient Chinese document containing severely degraded and densely typed characters is input into a convolutional neural network encoder to extract preliminary, multi-level visual features. These features include basic stroke and texture information of the characters and are used as subsequent inputs.

[0038] 2) Combine with a deformable strip pooling multi-scale feature fusion (MFDSP) module, such as Figure 2 As shown,

[0039] Visual features are input into MFDSP, which serves as the core path for backbone feature refinement. MFDSP first unifies the channel dimensions of the visual features through 1×1 horizontal convolutions, and then processes the low-level features of the current layer using a top-down feature pyramid with deformable bar pooling channel attention. and the high-level features obtained by upsampling from the previous layer :

[0040] 2.1): Deformable feature alignment, using advanced features to predict spatial offsets. In standard normalized grid The above is achieved through bilinear resampling Correcting curved text:

[0041] (1.1)

[0042] 2.2): Orthogonal strip pooling and recalibration, in Vertical and horizontal bar pooling are applied to extract prior information about the vertical text columns and horizontal line spacing, respectively, and then aggregated into a global descriptor. and :

[0043] (1.2)

[0044] Combining one-dimensional convolution (ECA) with the Sigmoid function Generate channel attention weights :

[0045] (1.3)

[0046] 2.3): Top-down filtering and fusion, As a soft mask for filtering low-level features Background noise in the data, followed by high-level features The features are added together and then smoothly output as the main path fusion feature through 3×3 convolution and ReLU activation function. :

[0047] (1.4)

[0048] (1.5)

[0049] Then, As input to the transformer encoder, the output features ;

[0050] 3) Auxiliary Training Module: High Activation Topological Feature Mask (HATFM) module, such as... Figure 3 As shown,

[0051] First, for feature maps The spatial activation intensity is calculated by averaging along the channel dimension. :

[0052] (1.6)

[0053] To eliminate numerical fluctuations between different images, local Min-Max normalization is applied to obtain... , :

[0054] (1.7)

[0055] Set the base discard rate Define the dynamic discard probability With retention probability A binary mask is generated through Bernoulli sampling. :

[0056] (1.8)

[0057] To prevent training distribution shift, pixel-level scale compensation is performed to obtain damaged features. :

[0058] (1.9)

[0059] Damaged features Input MFDSP to get , and Reconstruction loss The HATFM module is only used to assist in learning structural priors during the training phase and is completely discarded during the inference phase.

[0060] 4) Query modeling and joint loss optimization.

[0061] Features based on transformer encoder output Initialize detection query and identification query .Bundle Input a linear layer to generate text classification scores, select the top K features based on the text classification scores, and feed these top K features into a linear layer to initialize the detection query. The proposal is for identifying queries From the detection query Extracting features from proposals And take the average along the height dimension to initialize the identification query. . Will Projected onto semantic probability space :

[0062] (1.10)

[0063] W is a learnable matrix, and Concatenation to generate language vectors Then, combined with task query and Perform cross-modal semantic interaction:

[0064] (1.11)

[0065] Where PE is the position code. It is an attention mask, followed by the task query after interaction. As input to the transformer decoder;

[0066] 5) Finally, the transformer decoder outputs detection query and recognition query. The detection query enters the detection head, which consists of two feedforward layers, and outputs the position of the detection box. The recognition query enters a linear layer, namely the recognition head, and outputs the recognized character.

[0067] 6) Use of joint mission losses (classification loss is focal loss. Bounding box loss includes l1 loss and GIoU loss, polygon loss also uses l1 loss, and recognition loss is standard cross-entropy loss.) as well as the cosine similarity reconstruction loss calculated for the bypass branch of the high activation topological feature mask module perform overall joint optimization:

[0068] (1.12)

[0069] (1.13)

[0070] where and are the real label classes and bounding boxes, represents the prediction of bounding boxes, is predicted probability of the class, , , and are the loss weights of classification, bounding box, polygon and recognition respectively.

[0071] In this example, two specific recognition cases A and B in Figure 4 are identified and analyzed in depth, which can clearly reveal the significant technical advantages of the method of this example compared with the baseline method.

[0072] In Case A, the real label starting from the second column from the left is "若無願増語是菩薩摩訶薩善現汝", the prediction of baseline is "右無願増語是菩薩摩訶薩善現汝", and the prediction of this example is "若無願増語是菩薩摩訶薩善現汝". The core difference between the two lies in the first character "若", which the baseline model incorrectly recognizes as "右", a character with very similar shape. In ancient books, due to micro stroke breakage or ink fading caused by physical degradation, such similar-shaped characters are extremely similar at the feature level. The method of this example can correctly recognize, which highlights the powerful capability of its core high activation topological feature mask (HATFM) auxiliary module. By internalizing the topological prior structure of Chinese characters, it effectively repairs incomplete strokes at the feature level, enabling the model to resist the interference of visual blurring and make accurate judgments when distinguishing characters with similar shapes.

[0073] This advantage is more prominent in the complex typesetting scenario of Case B, where typical missed recognition and truncation errors caused by "feature aliasing" are presented. First, in the third column from the left of Case B, there is a unique typesetting of ancient books called "small-character double-line interlinear annotation". The conventional square receptive field of the baseline model, when extracting these tiny and crowded characters, is seriously interfered by adjacent text columns, resulting in the missed recognition of key characters such as "御之父 (Yu Zhi Fu)" in the interlinear annotation; however, the method proposed in this example can clearly and accurately separate and recognize complete double-line small characters. Second, starting from the eighth column on the right of Case B, the ground truth label is "岳富貴輕浮雲我願陶靖節𦎟墻若見聞窮蘆樂真趣", the prediction of the baseline is only "岳富貴輕浮雲我願陶靖節𦎟墻若見", while the prediction of the method in this example is "岳富貴輕浮雲我願陶靖節𦎟墻若見聞窮蘆樂真趣". The baseline model omits the second half of the text, resulting in a serious premature recognition truncation.

[0074] The ability of the method in this example to overcome the above two types of problems is entirely attributed to its Multi-scale Feature Fusion module combined with Deformable Strip Pooling (MFDSP) on the main path. This module uses orthogonal strip pooling to construct a "louver" geometric constraint, which extremely accurately filters the horizontal interference from adjacent text columns. It clearly separates the tiny and crowded interlinear annotation text and retrieves the missed recognition characters "御之父".

[0075] These two cases prove from multiple dimensions, such as easily confused similar characters, missed characters in dense double-line interlinear annotations, and premature truncation caused by curved text columns, that compared with traditional methods, the method in this example has significant robustness and accuracy when dealing with typical challenges of ancient book OCR (physical degradation and feature aliasing).

[0076] In terms of quantitative performance, the performance advantage of this example is also fully verified on the MTHv2 benchmark with severe physical degradation and complex typesetting. As Figure 5 shows, comparing the method of this example (marked as "Ours") with a variety of mainstream text detectors (such as Mask R-CNN, DBNet++, etc.) and special historical document analysis frameworks (such as HisDoc R-CNN, DB-SegHist, etc.), the method of this example achieves the highest detection precision of up to 98.80% on the MTHv2 dataset, and the overall F-measure score reaches 97.69%. This result is significantly better than many widely used well-known models, which proves the effectiveness and advancement of the multi-scale feature fusion architecture proposed in this example in overcoming dense typesetting aliasing. As Figure 6As shown, this example method is compared with other advanced end-to-end page-level recognition methods (such as two-stage Det+Recog, FOTS, PageNet, etc.). This example method achieves 94.25% and 95.70% in the two core metrics of average recognition rate (AR*) and character recognition rate (CR*), respectively. This performance significantly outperforms two-stage pipeline methods and models such as Star Flow-Read, and surpasses the leading end-to-end system PageNet in the crucial character recognition rate (CR*). This result strongly demonstrates that this example method successfully addresses the shortcomings of existing technologies when dealing with degenerative features such as microscopic stroke breaks, achieving top-tier performance in the field of end-to-end recognition of complex ancient texts.

Claims

1. A method for recognizing ancient texts based on topological masking and deformable strip pooling, characterized in that, Includes the following steps: 1) Preliminary visual feature extraction, The input image of the ancient book is fed into a convolutional neural network encoder to extract preliminary, multi-level visual feature maps F. Feature map F contains the basic structure and texture information of the characters and is used as subsequent input. 2) Combining a multi-scale feature fusion module with deformable strip pooling, The multi-scale feature fusion module, as the core path for backbone feature refinement, takes the input feature map F and first unifies the channel dimensions through a 1×1 horizontal convolution. Then, it uses a top-down feature pyramid with deformable bar pooling channel attention to process the low-level features of the current level. and the high-level features obtained by upsampling from the previous layer The fused features are obtained through deformable feature alignment, orthogonal strip pooling and recalibration, and top-down filtering fusion. , will integrate features As input to the transformer encoder, the output features ; 3) Bypass-assisted training module: High-activation topological feature mask module, First, for feature maps The spatial activation intensity is calculated by averaging along the channel dimension. : (1.6); Set the base discard rate Define the dynamic discard probability With retention probability And a binary mask is generated through Bernoulli sampling; Then pixel-level scale compensation is performed to obtain the damage features. Damage characteristics Input multi-scale feature fusion module to obtain Reconstruction loss The high-activation topological feature mask module is only used to assist in learning structural priors during the training phase and is completely discarded during the inference phase. 4) Query modeling and joint loss optimization. Features based on transformer encoder output Initialize detection query and identification query ,Bundle Input a linear layer to generate text classification scores, select the top K features based on the text classification scores, and feed these top K features into a linear layer to initialize the detection query. The proposal is for identifying queries From the detection query Extracting features from proposals And take the average along the height dimension to initialize the identification query. ,Will Projected onto semantic probability space : (1.10) W is a learnable matrix, and Concatenation to generate language vectors Then combine with task query and Perform cross-modal semantic interaction and query the task after the interaction. As input to the transformer decoder; 5) Finally, the transformer decoder outputs detection query and recognition query. The detection query enters the detection head, which consists of two feedforward layers, and outputs the position of the detection box. The recognition query enters a linear layer, namely the recognition head, and outputs the recognized character. 6) Use joint mission losses And the cosine similarity reconstruction loss calculated for the bypass branch of the high-activation topological feature mask module. Perform overall joint optimization: (1.12) (1.13) in and These are real label classes and bounding boxes. The prediction representing the bounding box, yes Predicted probability of class , , and These are the loss weights for classification, bounding boxes, polygons, and recognition, respectively.

2. The method for ancient text recognition based on topological masking and deformable strip pooling according to claim 1, characterized in that, The specific steps of step 2) are as follows: 2.1): Deformable feature alignment, using advanced features to predict spatial offsets. In standard normalized grid The above is achieved through bilinear resampling Correcting curved text: (1.1) 2.2): Orthogonal strip pooling and recalibration, in Vertical and horizontal bar pooling are applied to extract the vertical and horizontal structural priors, respectively, and then aggregated into a global descriptor. and The vertical axis represents the text column, and the horizontal axis represents the line spacing. (1.2) Combining one-dimensional convolution with the Sigmoid function Generate channel attention weights : (1.3) 2.3): Top-down filtering and fusion, As a soft mask for filtering low-level features Background noise in the data, followed by high-level features The features are added together and then smoothed out by convolution and ReLU activation to fuse the main path. : (1.4) (1.5) Then, As input to the transformer encoder.

3. The method for ancient text recognition based on topological masking and deformable strip pooling according to claim 1, characterized in that, In step 3), to eliminate numerical fluctuations between different images, local Min-Max normalization is applied to obtain... , : (1.7) Set the base discard rate Define the dynamic discard probability With retention probability A binary mask is generated through Bernoulli sampling. : (1.8) To prevent training distribution shift, pixel-level scale compensation is performed to obtain damaged features. : (1.9) Damaged features Input multi-scale feature fusion module to obtain , and Reconstruction loss The computationally intensive, highly activated topological feature masking module is only used during the training phase to assist in learning structural priors and is completely discarded during the inference phase.

4. The method for ancient book text recognition based on topological masking and deformable strip pooling according to claim 1, characterized in that, In step 4), and Concatenation to generate language vectors Then combine with task query and Perform cross-modal semantic interaction: (1.11) Where PE is the position code. It is an attention mask, followed by the task query after interaction. As input to the transformer decoder.

5. The method for ancient text recognition based on topological masking and deformable strip pooling according to claim 1, characterized in that, Joint mission loss in step 6) Including classification loss This refers to focus loss and bounding box loss. It includes L1 loss and GIoU loss, the polygon loss also uses L1 loss, and the recognition loss is the standard cross-entropy loss.