A method for intelligently identifying and automatically filling key record information of a document archive

CN122821574APending Publication Date: 2026-09-25NEIJIANG GUOXING CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610999794.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明提供一种文书档案关键著录信息智能识别及自动入库填充方法,解决相关技术中退化扫描件关键著录字段识别失准、批量自动入库效率低且缺乏分级复核机制的技术问题

Benefits of technology

[0036]将背景退化补偿所得背景色调偏移量用于印章检测色相区间的动态补偿,并结合色彩加权分离、同文书字体笔画密度比值驱动的实例化字符模板及著录字段候选字符集约束,针对政务印章遮挡成文日期等场景改善字符恢复与定位稳定性,减少因纸张老化、字体差异造成的误识与漏识,使部分遮挡字段具备自动写入或高可信度辅助确认条件,缓解通用OCR在遮挡区域输出乱码或空白带来的全量人工重录压力,提升印章相关著录字段的可处理比例;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821574A_ABST
    Figure CN122821574A_ABST
Patent Text Reader

Abstract

The application relates to the fields of archive informatization and document image processing technology, and discloses a method for intelligently identifying and automatically filling key recording information of official documents and archives, which comprises the following steps: sequentially performing background degradation compensation, geometric correction and effective resolution evaluation on a scanning image to obtain a normalized official document image with a quality mark vector; performing multi-level segmentation of a seal, color weighted separation and font parameterized template matching to generate a pure text image and seal processing metadata; performing edition intergenerational discrimination and layout structure analysis to obtain a recording field candidate area set and a full-text character stream; obtaining recording field extraction values and confidence scores through multi-path collaborative extraction and cross-field consistency verification; and performing double-threshold grading storage routing and feedback loop. The application can improve the field recognition accuracy of batch recording of historical official documents and reduce the workload of manual review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of archival information technology and document image processing technology, and more specifically, it relates to a method for intelligent identification and automatic data entry of key bibliographic information in document archives. Background Technology

[0002] In the digitization of government archives, the extraction and storage of key bibliographic information from scanned paper documents has long relied on optical character recognition (OCR) engines combined with manual proofreading. Existing general OCR solutions are geared towards modern document designs with relatively uniform layouts. After performing character recognition on the scanned images, they directly output the full text, which then requires manual searching and filling in fields such as title, document number, date of issuance, and issuing authority. This approach is insufficient to meet the timeliness requirements of large-scale archiving.

[0003] Provincial-level archives management institutions need to process a large number of scanned government documents spanning multiple historical periods each year. Within the same batch, there are documents formatted according to the old version before 1993, as well as documents showing typical signs of degradation such as yellowing backgrounds due to paper aging and stamps covering the date areas. The header structure and field positions of these documents differ from current national standards, making the cataloging tasks more demanding in terms of identification, disambiguation, and data entry routing than those for general office documents.

[0004] The existing technologies described above have the following problems: Fixed hue thresholds for seal detection are difficult to adapt to differences in background color shifts between different batches, resulting in both false positives and false negatives; general character templates cannot adapt to individual font differences in historical documents, and character recognition in areas obscured by seals is easily confused; single-format templates are insufficient to cover documents spanning different eras, leading to insufficient reliability in field location; the same format for the date of issue, date of publication, and date of citation makes it impossible to determine the recording target based solely on character content, resulting in low batch recording efficiency, a large workload for manual verification, and an increased risk of omissions and errors in key fields. Furthermore, there is a lack of field-level confidence grading and review task generation mechanisms that integrate with the document management system. Summary of the Invention

[0005] This invention provides a method for intelligent identification and automatic data entry of key bibliographic information in document archives, which solves the technical problems of inaccurate identification of key bibliographic fields in degraded scanned documents, low efficiency of batch automatic data entry, and lack of hierarchical review mechanism in related technologies.

[0006] This invention provides a method for intelligent identification and automatic data entry of key bibliographic information in documents and archives, comprising the following steps:

[0007] S1, acquire the scanned image of the original archival document, and sequentially perform background degradation compensation, geometric correction and effective resolution evaluation on the image to obtain a normalized document image with attached quality annotation vector;

[0008] S2, based on the normalized document image with attached quality annotation vector, performs prior multi-level segmentation of seal color and shape, color weighted separation, font parameterized template matching and candidate character set constraints to generate a plain text image after seal occlusion repair and seal processing metadata.

[0009] S3, based on the plain text image after the seal occlusion repair, performs global character recognition and layout type recognition, and combines flexible prior template and anchor point positioning to analyze the layout structure and obtain the candidate area set of bibliographic fields, the layout functional area mapping map and the full text character stream;

[0010] S4 utilizes the candidate region set of bibliographic fields, the page functional area mapping map, the full-text character stream and the stamp processing metadata. Through dual-path result alignment, field heterogeneity perception multi-path collaborative extraction and cross-field consistency verification, the extracted values ​​of each bibliographic field and the field confidence score are obtained.

[0011] S5, based on the extracted values ​​of each bibliographic field and the field confidence score, performs dual-threshold hierarchical data entry routing in combination with the business importance of the fields, and performs a closed loop of character confusion matrix update and page layout parsing model incremental fine-tuning, outputting archived bibliographic information and a list of low-confidence field review tasks.

[0012] In a preferred embodiment, the background degradation compensation is handled for three types of degradation respectively: for overall hue shift, limited contrast adaptive histogram equalization is performed on the luminance channel in the LAB color space, and compensation is applied to the chroma channel with the pixel mean of the edge blank area, and the background hue shift is written into the quality label vector.

[0013] For stroke edge blurring caused by ink diffusion, anisotropic diffusion filtering is performed; for physically damaged areas, only the damage coordinate range is recorded, and no repair is performed.

[0014] In a preferred embodiment, the prior multi-level segmentation of the seal color morphology is performed in three stages in sequence: the initial color screening stage extracts the background hue offset from the quality annotation vector, uses the offset to perform dynamic compensation on the hue interval of the seal color prior distribution in the HSV color space, performs pixel-level screening on the whole image and performs connected component analysis, and filters out fragments with an area lower than the minimum seal area threshold.

[0015] In the morphological verification stage, the roundness index and size of the candidate region are calculated, and the center and radius of the circle are fitted by Hough circle transform. Regions whose fitting residuals exceed the first morphological tolerance range are excluded. In the semantic verification stage, the curved text inside the ring is extracted from the region that has passed the morphological verification. The character-level similarity is compared with the administrative agency name dictionary. When the similarity exceeds the first semantic similarity threshold, it is identified as a valid seal. The center coordinates, radius and agency attributes are written into the seal processing metadata.

[0016] In a preferred embodiment, before performing occluded character matching, the font parameterized template matching calculates the ratio of the average stroke width to the average character height among the recognizable digital characters outside the seal boundary, and uses this ratio as the font stroke density ratio of the current document.

[0017] Using the ratio of the standard stroke width to the standard character height of the general character template as a benchmark, the two ratios are divided to obtain the adaptation factor; using the adaptation factor as the anisotropic scaling coefficient in the stroke width direction, the general template of each character in the candidate character set is scaled and deformed in the width direction to generate an instantiated character template set;

[0018] When the number of recognizable numeric characters is lower than the preset minimum sampling threshold, the adaptation factor is set to 1 and the template adaptation level is marked as unadapted in the recovery confidence.

[0019] In a preferred embodiment, the candidate character set constraint is based on the prior characteristic that the position of the seal page is concentrated in the date area of ​​the document, and the candidate character set of the corresponding character grid is narrowed to the Gregorian calendar year, month and day character set.

[0020] For character cells whose page position is uncertain and which may cross the document date area and document number area, the matching results and confidence scores of the date candidate set and document number candidate set are retained, and stored together with the character page coordinates in the seal processing metadata, so that the corresponding candidate set can be selected according to the field affiliation in subsequent steps.

[0021] For completely occluded grids where the remaining stroke information is close to zero, no character recovery results are output, the confidence level is recorded as zero, and the relevant fields are sent to the manual review route.

[0022] In a preferred embodiment, the layout type recognition comprehensively detects the shape of the header dividing line, the completeness of the footer area, and the height distribution of characters in the header text line, classifies the document according to the layout regularity, and outputs the layout recognition confidence score.

[0023] For documents with high confidence in layout recognition and complete layout structure, the normalized expected positions of each bibliographic field are obtained from the pre-stored layout template library, and a search area is constructed with the expected positions as the center. The difference obtained by subtracting the layout recognition confidence from 1 is multiplied by the horizontal expansion coefficient and the vertical expansion coefficient respectively, and then multiplied by the benchmark search radius to obtain the expansion amount. The lower the layout recognition confidence, the wider the search range, and the vertical expansion coefficient is smaller than the horizontal expansion coefficient.

[0024] In a preferred embodiment, the anchor point positioning is for documents with low confidence in layout recognition or incomplete layout structure. Three types of visual anchor points are used to replace fixed coordinates for prior field positioning: the signature area anchor point is located by detecting abnormal line spacing in the lower half of the text, and the candidate area for the document date is limited to its neighborhood.

[0025] The font size mutation point anchor point locates the title candidate area by detecting the local maxima of the average character height distribution of each text line; the document number structure anchor point locates the document number candidate area by searching for character substrings with fixed combinations of bracket-type symbols and year numbers in the full text character stream.

[0026] In a preferred embodiment, the dual-path result alignment cross-compares the candidate region coordinates of each field with the dual-path character position coordinates stored in the seal processing metadata, selects the recovery result of the corresponding candidate set according to the field type, and puts the character sequence of each field into a single-path determination state.

[0027] In the heterogeneity-aware multi-path collaborative extraction of the fields, regular fields use a finite state machine to perform state transition matching on the character sequence within the field candidate area to verify the fixed combination structure of agency codes, bracket symbols, year numbers, serial numbers and number characters. When the types of left and right brackets in historical documents are inconsistent, they are still accepted and marked as non-standard bracket forms.

[0028] Ambiguity-resolving fields are assigned to candidate fields based on their functional zoning in the layout functional area mapping diagram. The printing date falling in the imprint area and the reference date in the main text paragraph are excluded, and the document date is locked in the adjacent position of the signature area for extraction.

[0029] In a preferred embodiment, the cross-field consistency verification executes the following rule: the name of the agency referred to by the agency code in the document number should be semantically consistent with the value of the issuing agency field;

[0030] The year in the document date should be the same as the year in parentheses in the document number. If they are inconsistent, the confidence level of both should be reduced. The security classification field and the urgency level field should be present or absent at the same time. If one is present and the other is absent, the confidence level of the present field should be reduced. The overall confidence level of each bibliographic field is obtained by weighting and summing three components: the credibility of the extraction path, the content compliance score, and the cross-field consistency score, according to the specific weight of the field type.

[0031] In a preferred embodiment, the dual-threshold hierarchical inbound routing compares the comprehensive value of a field with two preset thresholds: when the comprehensive value is higher than the first inbound threshold, it is submitted to the file management system in an automatic write state; when the comprehensive value is between the two thresholds, it is written to the database in a pending confirmation state and included in the batch quality check.

[0032] If the comprehensive value is lower than the second review threshold, it will not be written to the database, but a review task entry containing candidate results and confidence component details will be generated and pushed to the manual review queue.

[0033] The character confusion matrix update and layout parsing model incremental fine-tuning feedback loop includes two channels: the character confusion matrix channel adjusts the convolutional similarity score in real time when systematic confusion is detected;

[0034] The incremental fine-tuning channel of the layout analysis model is triggered after accumulating manually corrected records to a preset threshold number. The weights of the backbone network are kept frozen, and only the parameters of the classification head and regression head are updated.

[0035] The beneficial effects of this invention are as follows:

[0036] The background hue offset obtained from background degradation compensation is used for dynamic compensation of the hue range of seal detection. Combined with color weighted separation, instantiated character templates driven by the stroke density ratio of the same document font, and candidate character set constraints of the recording field, the stability of character recovery and positioning is improved for scenarios such as government seals obscuring the document date. This reduces misidentification and omission caused by paper aging and font differences, and enables some obscured fields to have automatic writing or high-confidence auxiliary confirmation conditions. It alleviates the pressure of full manual re-entry caused by the output of garbled characters or blanks in the obscured area by general OCR, and improves the processable ratio of seal-related recording fields.

[0037] By combining page layout generational discrimination with flexible prior templates and Class A visual anchors for page structure analysis, and relying on the page functional area mapping map to resolve ambiguities between the document date, the issuance date, and the citation date, and cooperating with cross-field consistency verification and field business importance weighting dual-threshold hierarchical data entry routing, along with a dual-channel feedback loop of real-time character confusion matrix updates and incremental fine-tuning of the page layout analysis model, the risk of incorrect data entry is controlled while reducing the proportion of fields requiring manual verification or modification in batch processing. This improves the overall operational efficiency of digitized cataloging of government archives and the availability of archived data, and supports human-machine collaborative verification methods for low-confidence field review tasks and reference verification of identified fields. Attached Figure Description

[0038] Figure 1 This is a flowchart of a method for intelligent identification and automatic data entry of key bibliographic information in document archives according to the present invention;

[0039] Figure 2 This invention provides a flowchart of a method for intelligent identification and automatic data entry of key bibliographic information in document archives. Figure 1 ;

[0040] Figure 3 This invention provides a flowchart of a method for intelligent identification and automatic data entry of key bibliographic information in document archives. Figure 2 . Detailed Implementation

[0041] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0042] At least one embodiment of the present invention discloses a method for intelligent identification and automatic database filling of key bibliographic information in documents and archives, such as... Figures 1 to 3 As shown, it includes the following steps:

[0043] S1, acquire the scanned image of the original archival document, and sequentially perform background degradation compensation, geometric correction and effective resolution evaluation on the image to obtain a normalized document image with attached quality annotation vector;

[0044] In one embodiment of the present invention, the image quality defects of government archive scans are not limited to a single type. Within the same batch, multiple issues such as low resolution, page tilt, yellowing background, and localized stains often occur simultaneously, and each type of defect interferes with subsequent steps in different ways: yellowing background introduces a systematic shift into the red channel, interfering with the color-priority-based seal positioning in S2; page tilt causes a deviation between the projected position of the prior coordinates of the page layout and the actual position; stroke diffusion blurs the seal edge outline, affecting the fitting accuracy of the circular outline. Given that background degradation compensation can effectively improve the signal-to-noise ratio of linear features in the image, thus providing a clearer header dividing line signal for subsequent geometric correction, and that the horizontal text line layout after geometric correction can make the stroke statistics for resolution evaluation more accurate, this step is executed sequentially in the order of background degradation processing, geometric correction, and effective resolution evaluation, forming a positive dependency chain where the preprocessing creates favorable conditions for subsequent processing.

[0045] Specifically, background degradation processing is designed with different methods for three common types of archival degradation. The overall tone shift caused by paper oxidation is handled in the LAB color space: a limited contrast adaptive histogram equalization operation is performed on the luminance channel to make the local contrast of the image more uniform; simultaneously, using the pixel mean of the 5% width edge blank area on each of the four sides of the image (usually excluding stamps and text) as a benchmark, compensation is applied to the chroma channel to make the background color return to a neutral color direction, and the hue shift of the pixels in this area in the HSV color space is recorded as the background tone shift and written into the quality annotation vector. For the blurring of stroke edges caused by ink diffusion, anisotropic diffusion filtering is used. This filter automatically reduces the filtering strength at stroke edges with high gradient strength and increases the smoothing effect in flat background areas with low gradient strength, which can suppress diffusion noise while preserving the clarity of character edges. Physical damage (creases, tears, stains, etc.) only records the spatial extent of the damaged area without attempting repair, because artificial repair can easily introduce non-existent stroke information into the damaged area, which is more misleading for character recognition than the original blank space. The coordinates of the damaged area are written into the quality annotation vector in the form of a list of normalized coordinate rectangles.

[0046] After background degradation processing, geometric distortion correction is performed on the image with improved contrast. The correction adopts a two-stage scheme: the first stage targets page tilt, using Hough linear transformation to detect two types of directional structures: header dividing lines (horizontal solid or thin lines below the header area, which have more significant contrast after background compensation) and text line baselines (approximately parallel horizontal clusters formed by the bottom edges of adjacent text lines). The weighted median of the deviation of the measured angles of the two from the horizontal direction is taken (the header dividing lines are given a higher weight due to their more regular shape), and an affine transformation matrix is ​​constructed accordingly to complete the overall tilt correction.

[0047] The second stage addresses perspective distortion by detecting the straight lines on the four sides of the document, solving for the inverse perspective transformation matrix, and restoring the trapezoidal distortion at the four corners caused by improper scanning to a normal viewing state. Perspective distortion correction markers are recorded in the quality annotation vector. Geometric correction restores the horizontal alignment of text lines, providing an accurate pixel distribution basis for subsequent horizontal projection statistics in resolution evaluation.

[0048] After geometric correction is completed, an effective resolution assessment is performed on the corrected image. This invention does not directly use the DPI parameters recorded by the scanner because, under the same DPI setting, documents with different font sizes show significant differences in stroke pixel density. DPI itself cannot accurately reflect the degree of character detail retention. Instead, the effective resolution level is estimated by statistically analyzing the pixel ratio of stroke width to character height within a horizontal text line area. The specific method is as follows:

[0049] Edge detection is performed on the corrected image to distinguish between densely populated areas of stroke edges and flat background areas. Within horizontal text lines, the mean ratio of stroke width to character height in pixels is calculated for densely populated edge areas. If this ratio is lower than a first resolution threshold, partitioned upsampling is performed on the image. Sharpening interpolation is prioritized for text regions to preserve stroke outline details, while bilinear smoothing interpolation is performed on background regions to avoid noise amplification. When the degradation is too severe, causing densely populated edge areas to be unable to be effectively recognized, uniform interpolation upsampling across the entire image is used as a degradation strategy. Simultaneously, the confidence level of the effective resolution assessment in the quality annotation vector is recorded as low reliability, instructing subsequent steps to apply additional confidence reduction to the overall character recognition results of the document.

[0050] The output quality annotation vector stores all evaluation results in structured fields: tilt correction angle, a floating-point number in degrees, with positive values ​​indicating clockwise deflection and a range within ±15 degrees; perspective distortion correction identifier, a Boolean type; effective resolution level, taking one of four discrete levels: high, medium, low, and unevaluable; effective resolution evaluation confidence, a normalized floating-point number; background hue offset, represented by the hue channel offset in the HSV color space, calculated from the difference between the average hue value of pixels in the edge blank area and the neutral color reference value; ink diffusion intensity, a normalized floating-point number; a list of physically damaged areas, stored as a list of normalized coordinate rectangles, with each rectangle represented by a quadruple of the x and y coordinates of its top-left and bottom-right corners; and a comprehensive degradation reliability score, calculated based on the above degradation indicators, for subsequent steps to perceive the overall quality status of the current document image.

[0051] S2, based on the normalized document image with attached quality annotation vector, performs prior multi-level segmentation of seal color and shape, color weighted separation, font parameterized template matching and candidate character set constraints to generate a plain text image after seal occlusion repair and seal processing metadata.

[0052] In one embodiment of the present invention, the interference of official seals in government documents with the identification of cataloging fields is fundamentally different from that of general image noise. This is because noise is random and unstructured, while the color distribution, geometric shape, text content, and layout position of the seal are all subject to real-world constraints. Prior knowledge of three dimensions of government seals—color range (vermilion to orange-red), shape (perfect circle, diameter constrained by seal production standards), and text content (full name of the organization + seal / fixed format for special seals)—can construct a progressively layered positioning verification mechanism: color prior knowledge is used for rapid initial screening, eliminating the vast majority of non-seal areas; geometric constraints eliminate elements with similar colors but inconsistent shapes; and textual semantic verification serves as the final confirmation. The layout position of official seals in government documents is highly concentrated in the lower right corner of the document's date area, typically more than 40% of the page height away from the document number position slightly to the right of the center. This highly stable layout feature is an important prior basis for determining the attribution of character grid fields.

[0053] The core problem in the initial color screening stage is that the aging degree of paper from different eras in a batch of historical archives varies, resulting in differences in the amount of background hue shifting towards red. If a fixed hue interval threshold is used for screening, documents with yellowish backgrounds will be misdetected as candidate areas for seals, while documents with whitish backgrounds will be missed when the seal color is too light. Before the color screening is performed, this step extracts the background hue offset from the quality annotation vector passed in from S1. This offset is used to apply equal compensation to the upper and lower bounds of the hue interval in the HSV color space where the seal color is prior to its distribution, forming a dynamic seal detection hue interval that adapts to the current document background color state. Then, pixel-level screening is performed on the entire image, and connected component analysis is performed on the screening results to filter out fragmented connected components with an area lower than the minimum seal area threshold, retaining a number of candidate regions with the required area.

[0054] In the morphological verification stage, the contours of candidate regions are extracted, and a roundness index (a standardized value of the ratio of area to the square of the perimeter, which approaches 1 for a perfect circle) is calculated. The geometric dimensions of the candidate regions are then filtered based on the fixed proportional relationship between the standard specifications of government seals and the document layout. Regions filtered by both roundness and size are further fitted with the center coordinates and radius using Hough circle transform. Regions whose fitting residuals exceed the first morphological tolerance range are excluded, and the center coordinates and radius values ​​of the remaining regions are recorded. In the semantic verification stage, specialized character recognition is performed on the regions that pass morphological verification. Curved text within the annular region is extracted, and the extracted content is compared with a pre-set administrative agency name dictionary for character-level similarity. If the similarity exceeds the first semantic similarity threshold, it is considered a valid government seal, and its center coordinates, radius, agency attribute, and seal type are written into the seal processing metadata.

[0055] After the three-stage verification is completed, the process of restoring the obscured characters begins. Character grids within the stamp's coverage area are categorized into partially obscured and completely obscured grids, each following a different processing path. For partially obscured grids, stamp strokes (components with relatively high red channel intensity) and residual text strokes (gray-black components with relatively low red channel intensity) coexist in the same grid. The color-weighted separation step assigns stamp stroke weights and text stroke weights based on the red channel intensity of each pixel. This weighted processing yields a residual text stroke feature map, preserving the incomplete stroke morphology information of the partially obscured characters. Directly matching the residual stroke feature map using a general character template library suffers from systematic failures in historical document scenarios: the printed fonts in historical documents are highly individualized due to differences in age, region, and institution. The thickness and proportions of the same character's strokes may vary significantly across different documents. The general template cannot adapt to these individual font differences, leading to easy confusion between similar characters such as the number 6 and 0, or the characters for "month" and "use" when the residual stroke signal is insufficient.

[0056] In one embodiment of the present invention, to address the systemic failure problem, before performing occluded character matching, the font stroke parameters of the current document are extracted from identifiable numeric characters outside the seal boundary using the established circular boundary of the seal. Numeric characters appear in multiple places in government documents, such as document number annotations, cited clauses, and text serial numbers; as long as they are outside the seal coverage area, they can be used as sampling sources without waiting for the field positioning results of S3. The ratio of the average stroke width to the average character height of these numeric characters is calculated to obtain the font stroke density ratio of the current document. Then, using the ratio of the standard stroke width to the standard character height of the general character template as a benchmark, the two ratios are divided, and the resulting quotient is the adaptation factor. Using this adaptation factor as the anisotropic scaling coefficient in the stroke width direction, the general template of each character in the candidate character set is scaled and deformed in the width direction to generate an instantiated character template set adapted to the font style of the current document, which serves as the matching target for subsequent similarity calculations. When the number of recognizable numeric characters outside the seal boundary is lower than the preset minimum sampling threshold, the adaptation factor is set to 1, that is, no deformation is performed on the general template, the general template is used directly, and the template adaptation level is marked as unadapted in the recovery confidence of the relevant characters.

[0057] The candidate character set constraint mechanism further narrows the matching search space by leveraging the limited and predictable domain characteristics of the bibliographic field's value range. Since the seal's position is highly concentrated in the document date area, the candidate character set for character cells within the document date field can be reduced to the Gregorian calendar year, month, and day character set (i.e., Arabic numerals from 0 to 9 and Chinese characters for year, month, and day). For character cells with uncertain positions or the possibility of crossing both the document date and document number areas, the matching results and confidence scores of both candidate character sets (date candidate set and document number candidate set) are retained and stored along with the character's page coordinates in the seal processing metadata. This allows S4 to select the corresponding candidate set's recovery result based on the field type after confirming the field's attribution. Under the combined effect of the candidate character set constraint and instantiated character templates, the residual stroke feature map is convolved with each instantiated character template in the candidate set to calculate convolutional similarity. The candidate character with the highest score is used as the recovery result, and the maximum similarity score is used as the character-level recovery confidence score and stored in the seal processing metadata.

[0058] For completely occluded grids where the residual stroke information is close to zero due to excessive occlusion, the convolutional similarity scores of each candidate character tend to be uniform, making effective differentiation impossible. For such grids, no character recovery results are output; the confidence level is directly recorded as zero, and the relevant fields are routed for manual review in S5 without forced inference. This process clearly distinguishes between honest rejection when residual signals are insufficient and constraint-matched recovery when residual signals are valid, preventing low-confidence inferences from entering the automatic database process. After all character grids are processed, the stamp pixels are replaced with background texture fill (expanding inwards from the local background texture outside the stamp circumference).

[0059] S3, based on the plain text image after the seal occlusion repair, performs global character recognition and layout type recognition, and combines flexible prior template and anchor point positioning to analyze the layout structure and obtain the candidate area set of bibliographic fields, the layout functional area mapping map and the full text character stream;

[0060] After S2 outputs the plain text image, before proceeding to the page layout analysis, a global character recognition is performed on the image to generate a full-text character stream, which contains the recognized content of all characters in the text and the coordinate information of each character on the page. The result of this global recognition is shared in S3 and S4: S3 uses the coordinate information of each text line to assist in field region localization, and S4 uses the string content to perform semantic feature extraction and field boundary determination. The two steps share the same recognition result, avoiding redundant calculations caused by repeated recognition of the same image, and eliminating processing ambiguities that may be caused by inconsistencies in the character streams when the two steps are recognized separately.

[0061] The basic approach to page layout analysis is to first determine whether the current document layout has a reliable, fixed field layout, and then select a matching positioning strategy. Within the same batch of documents, the page positions of the recorded fields vary significantly: some documents have complete headers and footers, and relatively stable field positions; while some early archived documents have varied layouts and lack a unified layout standard. If the same set of fixed coordinates is used to analyze documents with significant layout differences, the field positioning deviation will increase.

[0062] Layout type recognition comprehensively judges based on visual features such as the shape of the header dividing line, the completeness of the footer area, and the distribution of header font size. Documents are classified according to their layout regularity, and a layout recognition confidence score is output. Then, a template matching path or anchor point positioning path is selected based on the layout recognition confidence score and the completeness of the layout structure. For documents with high layout recognition confidence and complete layout structure, the template matching path is applicable: the normalized expected position of each bibliographic field is obtained from a pre-stored layout template library, and a search area is constructed centered on the expected position. The search range narrows as the layout recognition confidence increases and widens as the confidence decreases. Specifically, the difference between 1 and the layout recognition confidence is multiplied by the horizontal expansion coefficient and the vertical expansion coefficient, and then multiplied by the baseline search radius to obtain the expansion amount. The vertical expansion coefficient is smaller than the horizontal expansion coefficient to reflect the relatively more stable layout regularity of the field's vertical coordinate.

[0063] For documents with low confidence in layout recognition or incomplete layout structure, fixed coordinate templates are not applicable. Instead, fields are located based on visual relative relationships within the page. Three types of anchor points exist in these documents that can be detected visually: signature area anchor points, where the visual characteristics are a small number of text lines (usually no more than 3 lines) and significantly wider line spacing than the body text. These can be located by detecting anomalies in the statistical values ​​of line spacing in the lower half of the page, thus limiting the candidate date area to the vicinity of the signature area; font size abrupt change anchor points, where the font size in the title area is always larger than that in the body text area and is nearly centered horizontally. These can be located by detecting local maxima in the average character height distribution of each text line; and document number structure anchor points, where substrings containing a fixed combination of bracket-like symbols and year numbers are retrieved from the string structure of the global character stream and used as positional anchor points for the document number candidate area. All of these anchor points are derived from the relative relationships within the page and do not rely on fixed coordinate priors.

[0064] In one embodiment of the present invention, within the elastic search range of the template matching path or the anchor neighborhood of the anchor positioning path, three types of layout element detection methods are comprehensively used to accurately locate the candidate region of the field: pixel density projection is performed on the horizontal direction of the image, and the pixel density distribution of each row is statistically analyzed to obtain the position and spacing information of the text lines; horizontal solid lines (header separator lines, border lines of the imprint area) and significant vertical blank bands in the image are detected to delineate the boundaries of each functional area. The layout functional area mapping map is generated at this stage, and the range of the header area, the main text area (including the boundary coordinates of the signature area), the imprint area, and each sub-area is marked; the character height distribution in each local area is statistically analyzed, and the positions where the font size is significantly larger than the surrounding area are marked as the basis for the title candidate area. The three types of detection results are jointly inferred with the expected position features in the template matching path or the neighborhood constraints in the anchor positioning path. When the detection result of a certain area is consistent with the expected features, it is confirmed as the candidate area of ​​the corresponding field. When there is a deviation between the detection result and the expectation, the measured visual evidence is given priority, and the deviation amount is recorded in the field type annotation. The document number field adds a round of structural verification on top of the elastic search: it retrieves character substrings that satisfy bracket symbols, year numbers, serial numbers and number characters from the global character stream, compares their page positions with the candidate areas determined by the elastic search for spatial consistency, and marks them as high confidence when the positions are consistent. When there is a discrepancy in the position, the candidate area is updated with priority based on the position verified by the string structure.

[0065] The output includes three items: a layout functional area mapping map (marking the boundary coordinates of the header area, main text area including the signature area, the imprint area and each sub-area), a candidate area set of bibliographic fields (each field contains candidate area coordinates, field type annotations and layout deviation markers), and a full-text character stream generated by global character recognition (the recognized content of each character and its layout coordinates).

[0066] S4 utilizes the candidate region set of bibliographic fields, the page functional area mapping map, the full-text character stream and the stamp processing metadata. Through dual-path result alignment, field heterogeneity perception multi-path collaborative extraction and cross-field consistency verification, the extracted values ​​of each bibliographic field and the field confidence score are obtained.

[0067] The input to S4 comes from two sources: the candidate region set of bibliographic fields, the layout function area mapping map, and the full-text character stream passed by S3, and the seal processing metadata passed directly by S2.

[0068] In one embodiment of the present invention, before starting the extraction of each field, step S4 performs a two-path result alignment processing first: cross-comparing the coordinates of each field candidate region with the two-path character position coordinates stored in the seal processing metadata, and for two-path characters whose coordinates fall within the current field candidate region, selecting the restoration result of the corresponding candidate set from the metadata according to the current field type. The character set of the date candidate set is selected for the writing date field, and the character set of the document number candidate set is selected for the document number field, which serves as the character stream input for the field extraction. After this alignment processing is completed, the character sequence of each field is in a single-path determined state, and the subsequent four extraction paths can be executed independently and mutually verified;

[0069] Regular fields are represented by the issuing document number, whose value structure is organ code + bracket-type symbols + 4-digit year number + serial number + the Chinese character "号" (meaning "No."). The extraction method uses a finite state machine to perform state transition matching on the character sequence in the field candidate region. The states and transition conditions are as follows: in the initial state, Chinese characters are continuously received to transition to the organ code accumulation state; in the organ code accumulation state, after accumulating 2 to 6 Chinese characters, when a left bracket-type character (historical variants such as square brackets, title brackets, etc.) is encountered, the state enters the bracket opening state and the bracket type is recorded; in the bracket opening state, after receiving 4 consecutive digits and verifying that the year value ranges from a preset starting year to the current year of identification plus 1, the state enters the year confirmation state; in the year confirmation state, after receiving a positive integer serial number, the state enters the serial number state; in the serial number state, after receiving a right bracket character corresponding to the recorded left bracket type, the state enters the bracket closing state (in historical documents, the situation where the left and right bracket types are inconsistent is still accepted and marked as non-standard bracket form); in the bracket closing state, after receiving the Chinese character "号", the state enters the acceptance state, and the document number string is completely extracted; in any state, when a character that does not meet the transition condition is encountered, the process returns to the initial state for re-matching. The value set of the security classification field is closed under the framework of confidentiality regulations, and full dictionary matching is performed on the character stream in the candidate region; for the situation where the recognition result and the valid entry have an edit distance within the first edit distance threshold due to minor OCR errors, correction is completed through fuzzy matching.

[0070] Boundary determination fields are represented by the title; S3's layout parsing has roughly located the candidate area for the title, and S4 performs fine boundary determination on the text lines provided by the global character stream within this rough range. The semantic features of the first character of the line used for determination come directly from the existing character recognition results of the global character stream and do not depend on the output of S4 itself, so there is no circular dependency. A feature vector is established for each text line, including the ratio of the font size of the line to the average font size of the main text, the horizontal alignment, and whether the sequence of characters at the beginning of the line belongs to the title guiding word dictionary (including verb phrase guiding words such as "about", "for", "then") or the addressee guiding word dictionary (including title guiding words such as "each", "to", "present"). Text lines with a font size higher than the average of the main text, horizontally centered, and whose first character belongs to the title guiding word dictionary are marked as title body lines; text lines with a font size close to the average of the main text and whose first character belongs to the addressee guiding word dictionary are marked as addressee lines; adjacent lines starting with parentheses are marked as parenthesis lines. Consecutive title rows are merged into title field values, while parenthetical rows are recorded separately and marked as optional additional fields.

[0071] Disambiguation-resolving fields are exemplified by the document date. An official document image typically contains three identical date sequences: the document date (located at the end of the main body area in the signature section, serving as the recording target), the issuance date (located in the imprint area), and the citation date (appearing in the text paragraphs when referencing other documents). These dates cannot be distinguished solely by their character content. Disambiguation is primarily based on the layout functional area mapping map output by S3: the document date falls at the end of the main body area adjacent to the signature section, the issuance date falls in the imprint area, and the citation dates are scattered throughout the text paragraphs of the main body area. The document date region is located based on the field candidate region's placement within the layout functional areas, and then the date string is extracted from the character stream within that region. If characters that have been repaired by S2 seal occlusion exist within the candidate region, the confidence levels of each recovered character are weighted and averaged into the overall confidence level of the document date field. When the confidence level of a single recovered character is lower than the second character recovery threshold, the overall confidence level of the document date field is correspondingly lowered, ensuring that the impact of seal occlusion on the field's reliability is accurately reflected.

[0072] Semantic generation fields are represented by keywords. These keywords do not exist in the original document text and must be generated through semantic analysis of the document content. The main text area is located from the layout function area mapping map output by S3. The global character stream text of this area is extracted as the source of the generated text. Based on the product of word frequency and inverse document frequency, type weights are introduced: administrative terms (policy and regulation names, project names, institution names) are given a higher weight than general words, and administrative verb phrases (issue, approve, strengthen, standardize, etc.) are given a higher weight than general verbs, reflecting the domain rules for selecting keywords in government documents. After candidate keywords are determined, they are compared and matched with the archival keyword list issued by the State Archives Administration: if a direct corresponding entry can be found in the list, that entry is used as the output keyword; if no direct correspondence exists, the semantically closest superordinate word is found from the list based on character similarity and semantic superordinate relationship as a candidate, and the source is marked as superordinate word replacement in the confidence score, with a correspondingly lower confidence score than in the direct matching case. Typically, 3 to 5 keywords are output.

[0073] Cross-field consistency verification is performed after the extraction of the four types of fields. The core verification rules include: the name of the agency referred to by the agency code in the document number should be semantically consistent with the value of the issuing agency field; the year in the date of issuance field should be the same as the year number in parentheses of the document number; if they are different, it means that at least one field has an identification error; the security classification field and the urgency level field are required to be marked or both missing in the standard official document format. If one exists and the other is missing, the confidence level of the existing field is reduced and a missing item prompt record is generated. Field pairs that pass the consistency test have their confidence level increased to a certain extent; for field pairs that are found to be inconsistent, the field with the lower confidence level is marked as pending verification and is prioritized for manual review in S5.

[0074] In one embodiment of the present invention, the overall confidence level of each bibliographic field is obtained by weighted summation of three items: extraction path confidence level, content compliance score, and cross-field consistency score. Different field types correspond to different weight allocations: for fields with strong regularity (such as document number and security classification), the content compliance score has a higher weight, with weight ratios of 0.3 (extraction path confidence level), 0.5 (content compliance), and 0.2 (cross-field consistency) in the example; for fields mainly based on boundary judgment (such as title), the extraction path confidence level is emphasized, with weight ratios of 0.5, 0.2, and 0.3 in the example; for fields that need to resolve ambiguity (such as the date of writing), the three weights are similar, with weight ratios of 0.35, 0.3, and 0.35 in the example; for semantic generation fields (such as keywords), due to their non-fixed format, consistency assessment is emphasized more, with weight ratios of 0.2, 0.1, and 0.7 in the example.

[0075] S5, based on the extracted values ​​of each bibliographic field and the field confidence score, combined with the business importance of the field, performs double threshold hierarchical data entry routing, and performs character confusion matrix update and page layout parsing model incremental fine-tuning feedback loop, outputting archived bibliographic information and a list of low-confidence field review tasks;

[0076] Each bibliographic field carries a quantified overall confidence score, which is the core basis for the inclusion decision, but not the only basis. Different fields have different levels of importance to archival retrieval and business use: title and document number are the most critical retrieval identifier fields, and errors will directly result in the archive becoming unsearchable; the date of issuance affects the accuracy of the annual filing of files; the classification and urgency level have a direct impact on the protection and management of archives; the importance of keywords is relatively lower; the inclusion routing mechanism combines the confidence score with the business importance weight of the field into a comprehensive value. High-importance fields are more likely to enter the review queue under the same confidence level. A more cautious attitude is maintained towards core fields in balancing the benefits of automation and the quality of bibliographic entry.

[0077] In one embodiment of the present invention, hierarchical routing assigns each field to three paths according to two preset thresholds. Fields with a comprehensive value higher than the first entry threshold are submitted to the document management system interface in an automatic write state. The written field value includes an identification source identifier, and the system records the confidence score details of the field in the audit log for reference in subsequent quality checks. Fields with a comprehensive value between the first entry threshold and the second review threshold are written to the database in a pending confirmation state. Field records in this state can be retrieved and used normally without affecting the basic usability of the current document entry, but will be included in the system's periodic batch quality check plan, prompting the document administrator to confirm in the next check cycle. Fields with a comprehensive value lower than the second review threshold are not written to the database, but instead generate a structured review task entry, which includes the unique identifier of the current document, the name of the problem field, the identification candidate result, the confidence component details, and a screenshot of the candidate area of ​​the field in the plain text image output by S2, and is pushed to the manual review queue. The plain text image is saved along with the document record throughout the entire processing flow for retrieval when generating the screenshot.

[0078] In batch processing scenarios, a single document often exhibits a mixed situation where some fields are automatically written while others are in the review queue. When the review task is delivered to a human reviewer, the system simultaneously displays the results of the other automatically written fields for the reviewer's reference. A typical scenario is that the document number (including year information) and issuing authority have been automatically written, while the date of issuance, due to insufficient confidence in restoration of the seal, enters the review queue. When processing the date of issuance, the reviewer can directly narrow down the scope of judgment using the displayed document number and year, usually without needing to revisit the original paper document to complete the verification. This design, which uses identified fields to assist human review, essentially extends the already completed cross-field consistency constraints at the human-computer collaboration level.

[0079] In one embodiment of the present invention, the feedback loop is divided into two independently operating channels, targeting different types of optimization objectives. The first channel is the real-time rule update of character similarity weights, which does not depend on model retraining: the system maintains a character confusion statistics matrix, and each element of the matrix records the cumulative number of times character i was identified by the system as character j in the historical correction records and then manually corrected. When the number of times character i is systematically misidentified as the same character j in the correction records within the preset statistics window M exceeds M multiplied by a preset ratio threshold, it is identified as systematic confusion. In the convolutional similarity calculation of S2, a preset reduction amount is applied to the similarity score of character j, while an equal amount of similarity boost is applied to character i, so that the system automatically avoids this confusion mode in subsequent matching. The adjustment takes effect synchronously after each manual correction operation, without waiting for the model fine-tuning cycle.

[0080] The second channel is for incremental fine-tuning of the layout analysis model, executed over a longer period. This model uses a convolutional neural network as its backbone feature extraction network. The input normalized document image is processed through multiple convolutional and pooling layers to obtain feature maps containing visual features at different scales. These feature maps are fed into two parallel output heads. The layout region classification head consists of several fully connected layers, taking the flattened vector of the feature map as input and outputting the category probabilities of each functional region, such as the header area, main text area (including the signature area), and footer area. The field boundary regression head also consists of several fully connected layers, taking the flattened vector of the feature map as input and outputting the normalized coordinate offset of each candidate region of the bibliographic field relative to its desired position. During the pre-training phase, the model uses scanned images of historical documents with manually labeled layout functional regions and field boundary coordinates as training samples. The layout region classification head uses the labeled region categories as supervision signals, and the field boundary regression head uses the labeled boundary coordinates as supervision signals. The losses of the two output heads are calculated as the sum of the cross-entropy of each category and the absolute value of the difference between the predicted coordinates and the labeled coordinates, weighted 2:3 to obtain the total loss. During the incremental fine-tuning phase, the backbone network weights are kept frozen, and only the parameters of the two output heads are allowed to participate in gradient updates. The trigger condition is that the manually corrected page area location records accumulate to a preset threshold K. No more than 20 training iterations are performed, with the learning rate set to one-tenth of that in the pre-training phase. After fine-tuning, the location accuracy of key fields (title, document number, and date of issuance) is evaluated on the validation set. If the accuracy improves compared to before fine-tuning, the online model weights are updated; otherwise, the original weights are retained. Manually corrected records for the date of issuance field are simultaneously incorporated into the character confusion matrix of the first channel, enabling collaborative updates between the two channels in the date of issuance scenario. After each batch of documents is processed, the system generates a batch quality statistics report, indicating deviations of the automatic writing rate, pending confirmation rate, and review rate of each field from the historical average, serving as a reference for the model maintenance plan.

[0081] In one embodiment of the present invention, the application scenario is applied to the batch digitization and cataloging of official internal documents of various units in a large enterprise's document archive center between 2025 and 2026. When a comprehensive archive management agency carried out a document digitization and cataloging project for its subordinate units, the document sources covered the headquarters' comprehensive management departments and multiple business branches. The document types were mainly work notices, approval opinions, business letters, application reports, and meeting minutes. Based on the document format characteristics, the above documents were mainly of categories with high format regularity, with some historical documents with incomplete format structures. A portion of these documents came from early archives of branch offices, with lower format regularity. Approximately a certain proportion of the scanned documents showed circular stamps covering the date area, and the paper quality was mainly light to moderate yellowing.

[0082] As shown in Table 1, the five representative documents selected from the above batches cover different processing difficulties such as no obscuring, partial obscuring, complete obscuring, and low regularity layout. Their document types, layout regularity, main quality problems, and seal obscuring are as follows.

[0083] Table 1. Basic Information of 5 Representative Document Samples

[0084]

[0085] As shown in Table 2, after the above 5 documents were processed by the method of the present invention, the recognition results of each major bibliographic field and the number of fields automatically written and entered into the review queue are as follows;

[0086] Table 2. Examples of identification results for the bibliographic fields in 5 documents.

[0087]

[0088] The above identification results reflect the processing behavior of the method of the present invention under different difficulty conditions. In serial number 2 (documents without seal obscuring and with a regular layout), all 7 bibliographic fields reached the automatic writing confidence level, which is the benchmark performance of the method under ideal conditions. In serial number 3 (the date of writing is completely obscured), S2 classifies the relevant cells of the date of writing into the completely obscured category and records the confidence level as zero. The date field enters the manual review queue; the title, document number, and other 5 fields are not affected by the seal obscuring and still complete the automatic writing normally. When the reviewer processes the date of writing in serial number 3, the system interface simultaneously displays the automatically written document number A, Unit Letter

[2026] 71. The reviewer can directly limit the year of writing to 2026 and only needs to confirm the month and date on the original document. The processing time of a single review task is thus significantly shortened. In document number 5 (a historically formatted document with low regularity), S3 identified it as a document with incomplete layout structure and switched to a field positioning mechanism based on visual anchors. Fields such as title, document number, and issuing unit were accurately positioned under the constraints of anchors in the signature area and font size mutation points. However, the document's content involved multiple work issues, and the correspondence between the candidate keywords and the thesaurus largely relied on the replacement of higher-level words. The confidence level was lower than the entry threshold, so it was entered into the review queue along with the document date field.

[0089] A comparison was made between the full processing results of this batch and the baseline solution using general OCR plus full manual verification: the baseline solution required manual confirmation for all fields, with a manual intervention rate of 100%; in this batch, the method of this invention automatically wrote approximately 83% of all fields, approximately 8% were pending confirmation, approximately 9% entered the review queue, and the proportion of fields requiring manual verification or modification was approximately 17%. For documents where the date of issue was obscured by a seal, the general OCR solution typically outputs garbled characters or blanks in the obscured area, requiring 100% manual re-entry of the field; the method of this invention, through the partial obscured character recovery mechanism of S2, enabled the date fields in serial numbers 1 and 5 to be automatically written or to assist in rapid manual confirmation with high reliability, reducing the workload of re-entry from scratch. Overall, manual intervention in this batch of cataloging processing mainly focused on the date fields completely obscured by seals and the subject term fields with low thesaurus matching, achieving improved cataloging efficiency while ensuring cataloging quality.

[0090] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A method for intelligent identification and automatic data entry of key bibliographic information in archival documents, characterized in that, Includes the following steps: S1, acquire the scanned image of the original archival document, and sequentially perform background degradation compensation, geometric correction and effective resolution evaluation on the image to obtain a normalized document image with attached quality annotation vector; S2, based on the normalized document image with attached quality annotation vector, performs prior multi-level segmentation of seal color and shape, color weighted separation, font parameterized template matching and candidate character set constraints to generate a plain text image after seal occlusion repair and seal processing metadata. S3, based on the plain text image after the seal occlusion repair, performs global character recognition and layout type recognition, and combines flexible prior template and anchor point positioning to analyze the layout structure and obtain the candidate area set of bibliographic fields, the layout functional area mapping map and the full text character stream; S4 utilizes the candidate region set of bibliographic fields, the page functional area mapping map, the full-text character stream and the stamp processing metadata. Through dual-path result alignment, field heterogeneity perception multi-path collaborative extraction and cross-field consistency verification, the extracted values ​​of each bibliographic field and the field confidence score are obtained. S5, based on the extracted values ​​of each bibliographic field and the field confidence score, performs dual-threshold hierarchical data entry routing in combination with the business importance of the fields, and performs a closed loop of character confusion matrix update and page layout parsing model incremental fine-tuning, outputting archived bibliographic information and a list of low-confidence field review tasks.

2. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The background degradation compensation addresses three types of degradation separately: for overall hue shift, a limited contrast adaptive histogram equalization is performed on the luminance channel in the LAB color space, and compensation is applied to the chroma channel using the pixel mean of the edge blank area, and the background hue shift is written into the quality label vector. For stroke edge blurring caused by ink diffusion, anisotropic diffusion filtering is performed; for physically damaged areas, only the damage coordinate range is recorded, and no repair is performed.

3. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The prior multi-level segmentation of the seal color morphology is performed in three stages: the initial color screening stage extracts the background hue offset from the quality annotation vector, uses the offset to perform dynamic compensation on the hue range of the seal color prior distribution in the HSV color space, performs pixel-level screening on the whole image and performs connected component analysis, and filters out fragments with an area lower than the minimum seal area threshold. In the morphological verification stage, the roundness index and size of the candidate region are calculated, and the center and radius of the circle are fitted by Hough circle transform. Regions whose fitting residuals exceed the first morphological tolerance range are excluded. In the semantic verification stage, the curved text inside the ring is extracted from the region that has passed the morphological verification. The character-level similarity is compared with the administrative agency name dictionary. When the similarity exceeds the first semantic similarity threshold, it is identified as a valid seal. The center coordinates, radius and agency attributes are written into the seal processing metadata.

4. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, Before performing occluded character matching, the font parameterized template matching calculates the ratio of the average stroke width to the average character height among the recognizable digital characters outside the seal boundary, and uses this ratio as the font stroke density ratio of the current document. The adaptation factor is obtained by dividing the ratio of the standard stroke width to the standard character height of the general character template. Using the adaptation factor as the anisotropic scaling coefficient in the stroke width direction, the general template of each character in the candidate character set is scaled and deformed in the width direction to generate an instantiated character template set; When the number of recognizable numeric characters is lower than the preset minimum sampling threshold, the adaptation factor is set to 1 and the template adaptation level is marked as unadapted in the recovery confidence.

5. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The candidate character set constraint is based on the prior characteristic that the position of the seal page is concentrated in the date area of ​​the document, and shrinks the candidate character set of the corresponding character grid to the Gregorian calendar year, month and day character set; For character cells whose page position is uncertain and which may cross the document date area and document number area, the matching results and confidence scores of the date candidate set and document number candidate set are retained, and stored together with the character page coordinates in the seal processing metadata, so that the corresponding candidate set can be selected according to the field affiliation in subsequent steps. For completely occluded grids where the remaining stroke information is close to zero, no character recovery results are output, the confidence level is recorded as zero, and the relevant fields are sent to the manual review route.

6. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The layout type recognition comprehensively detects the shape of the header dividing line, the completeness of the footer area, and the height distribution of characters in the header text line. Based on the layout regularity, the document is classified and the layout recognition confidence score is output. For documents with high confidence in layout recognition and complete layout structure, the normalized expected positions of each bibliographic field are obtained from the pre-stored layout template library, and a search area is constructed with the expected positions as the center. The difference obtained by subtracting the layout recognition confidence from 1 is multiplied by the horizontal expansion coefficient and the vertical expansion coefficient respectively, and then multiplied by the benchmark search radius to obtain the expansion amount. The lower the layout recognition confidence, the wider the search range, and the vertical expansion coefficient is smaller than the horizontal expansion coefficient.

7. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The anchor point positioning is for documents with low confidence in layout recognition or incomplete layout structure. It uses three types of visual anchor points to replace fixed coordinates for prior field positioning: the signature area anchor point is located by detecting abnormal line spacing in the lower half of the text on the page, and the candidate area for the document date is limited to its neighborhood. The font size mutation point anchor point locates the title candidate area by detecting the local maxima of the average character height distribution of each text line; the document number structure anchor point locates the document number candidate area by searching for character substrings with fixed combinations of bracket-type symbols and year numbers in the full text character stream.

8. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The dual-path result alignment cross-compares the candidate region coordinates of each field with the dual-path character position coordinates stored in the seal processing metadata, selects the recovery result of the corresponding candidate set according to the field type, and puts the character sequence of each field into a single-path determination state. In the heterogeneity-aware multi-path collaborative extraction of the fields, regular fields use a finite state machine to perform state transition matching on the character sequence within the field candidate area to verify the fixed combination structure of agency codes, bracket symbols, year numbers, serial numbers and number characters. When the types of left and right brackets in historical documents are inconsistent, they are still accepted and marked as non-standard bracket forms. Ambiguity-resolving fields are assigned to candidate fields based on their functional zoning in the layout functional area mapping diagram. The printing date falling in the imprint area and the reference date in the main text paragraph are excluded, and the document date is locked in the adjacent position of the signature area for extraction.

9. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The cross-field consistency verification follows these rules: the name of the issuing authority referred to by the agency code in the document number should be semantically consistent with the value of the issuing authority field. The year in the document date should be the same as the year in parentheses in the document number. If they are inconsistent, the confidence level should be reduced for both. The security classification field and the urgency level field should be present or absent simultaneously. If one is present while the other is absent, the confidence level of the present field is reduced. The overall confidence level of each bibliographic field is obtained by weighting and summing three components: extraction path credibility, content compliance score, and cross-field consistency score, according to the specific weights of the field type.

10. The method for intelligent identification and automatic data entry of key bibliographic information in document archives according to claim 1, characterized in that, The dual-threshold hierarchical inbound routing compares the comprehensive value of the field with two preset thresholds: when the comprehensive value is higher than the first inbound threshold, it is submitted to the file management system in an automatic write state. When the composite value is between the two thresholds, it is written to the database in a pending confirmation state and included in the batch quality check. If the comprehensive value is lower than the second review threshold, it will not be written to the database, but a review task entry containing candidate results and confidence component details will be generated and pushed to the manual review queue. The character confusion matrix update and layout parsing model incremental fine-tuning feedback loop includes two channels: the character confusion matrix channel adjusts the convolutional similarity score in real time when systematic confusion is detected; The incremental fine-tuning channel of the layout analysis model is triggered after accumulating manually corrected records to a preset threshold number. The weights of the backbone network are kept frozen, and only the parameters of the classification head and regression head are updated.