An automatic classification method for archival data based on machine learning
Patent Information
- Application Number
- CN202610945158.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
AI Technical Summary
不同来源文件页序不稳定,扫描件、PDF和电子文档混合输入时容易出现页码缺失、重复和跨件混排;现有分类多依赖关键字、标签或单页结果,难以同时约束切件状态、档案类别、保管期限和案卷关系;题名、文号、印章和表格栏位等关键证据对分类结果的贡献程度缺少量化校验,导致无关正文或低可信OCR内容容易影响分类;同时,复核结果通常只用于人工更正,难以反向更新证据区权重、状态转移代价和分类约束关系,影响后续批次档案自动分类的一致性和可持续优化能力
本发明通过将档案批次数据预处理为带有页序关系、文本块识别可信度和页序低可信标记的档案数据帧,解决了扫描件、PDF和电子文档混合输入时页码缺失、重复、跨件混排以及低可信OCR内容直接参与分类的问题,使后续切件、分类和保管期限判定能够基于统一页面对象和可追溯页序关系执行,降低批量档案处理中因输入形态不一致造成的误分件和误归类风险;
Smart Images

Figure CN122838631A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent processing technology for archival data, and in particular to an automatic classification method for archival data based on machine learning. Background Technology
[0002] With the increasing demand for digital management of archives in government agencies, enterprises, and engineering projects, technologies for batch receiving, automatic classification, intelligent file assembly, and retention period determination of archives have received widespread attention. Among existing technologies, CN114117171A discloses an intelligent method for organizing engineering archives, involving intelligent classification, retention period division, intelligent file assembly, file sorting, and file sorting within a file; CN116663549B discloses a method for digital management of enterprise archives, involving keyword extraction, classification tag generation, and archiving organization; CN113536182A also discloses an inference processing method based on content block sequences and type sequences. While these technologies can improve archive organization efficiency to some extent, the following problems still exist in actual batch archive processing scenarios: The page order of documents from different sources is unstable, and when scanned documents, PDFs, and electronic documents are mixed, page numbers are easily missing, duplicated, and mixed across documents. Existing classifications mostly rely on keywords, tags, or single-page results, making it difficult to simultaneously constrain the status of cut-off documents, archive categories, retention periods, and file relationships. The contribution of key evidence such as titles, document numbers, seals, and table fields to the classification results lacks quantitative verification, making it easy for irrelevant text or low-reliability OCR content to affect the classification. At the same time, the review results are usually only used for manual correction, making it difficult to update the evidence area weights, state transition costs, and classification constraints in reverse, affecting the consistency and sustainable optimization capabilities of automatic classification of subsequent batches of archives.
[0003] Therefore, how to provide an automatic classification method for archival data based on machine learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] One objective of this invention is to propose an automatic classification method for archival data based on machine learning. This invention fully utilizes the improved LayoutLMv3 archival classification model, evidence occlusion contribution analysis, and Viterbi dynamic programming algorithm to perform page order sorting, evidence area identification, component classification matching, and review feedback updates on batch archival data. It has the advantages of high classification accuracy, strong rule consistency, low manual review workload, and high batch processing efficiency.
[0005] An automatic classification method for archival data based on machine learning according to an embodiment of the present invention includes the following steps: Step 1: Receive batch data of archives, perform preprocessing on the batch data of archives and write the page sequence relationship to generate archive data frames; Step 2: Read the archive classification configuration data and generate a segmentation-classification constraint diagram based on archive category, file relationship, retention period rules, and segmentation status; Step 3: Input the archive data frame into the improved LayoutLMv3 archive classification model, identify the title, document number, seal and table fields as evidence areas, generate candidate results based on the complete evidence area and the obscured evidence area respectively, and generate evidence contribution value based on the difference between the two candidate results; Step 4: Based on the credibility of the evidence area and the evidence contribution value, align the text features and image layout features with the credibility of the evidence to generate enhanced archival features; Step 5: Generate segmentation-classification candidate states based on the archive enhancement features, match the segmentation-classification candidate states with the segmentation-classification constraint graph, and generate path residuals; Step 6: Configure the candidate states of segmentation-classification for the same batch of archives as the classification states of the Viterbi dynamic programming algorithm. Configure the state cost according to the model confidence, evidence contribution value and path residual. Configure the state transition cost according to the page order relationship, segmentation state, case file continuity relationship and classification rule matching relationship. Solve for the segmentation-classification state sequence with the lowest cumulative cost. Step 7: Write the final classification result according to the segmentation-classification state sequence, write the archive objects whose cumulative cost exceeds the preset cost threshold or whose evidence contribution value is lower than the preset contribution threshold into the review pool, and update the evidence area weight, state transition cost and segmentation-classification constraint graph according to the review result.
[0006] Optionally, step one specifically includes: Receive batch data of archives, read the source identifier, file directory level and file arrangement order, generate batch identifier, and generate original file identifier and original file sequence number for the original file; Perform file type identification on the batch data of archives, write file type tags, and uniformly convert image files, PDF files and electronic document files into page-level archive objects, and generate page identifiers for page-level archive objects; Perform page standardization processing on page-level archive objects to generate standard page images, and perform OCR recognition and layout region recognition on the standard page images to generate page text, text block coordinates, text block recognition confidence, page content index, layout region index and page number candidate region recognition results; Write an identification confidence tag based on the identification confidence level of the text block; Page sequence numbers are generated based on the original file sequence number, page number, table of contents record order, and page number candidate area recognition results, and page sequence relationship markers are written according to the page sequence numbers; when the page number recognition results are missing, duplicated, or inconsistent with the page order in the original file, a low confidence page sequence marker is written. The batch identifier, original file identifier, page identifier, file type marker, standard page image index, page sequence number, page content index, layout area index, text block recognition confidence, recognition confidence marker, page sequence low confidence marker, and page sequence relationship marker are encapsulated into an archive data frame.
[0007] Optionally, step two specifically includes: Read the file classification configuration data to obtain the file category table, file rule table, retention period rule table, and component cutting rule table; Generate file category nodes based on the file category table, generate file relationship nodes based on the file rule table, generate retention period nodes based on the retention period rule table, and generate cut status nodes based on the cut rule table. The segmentation status node includes a continuation of the previous status node, a new status node, and a new case file status node. Establish a time limit constraint edge between the archive category node and the retention period node, establish an attribution constraint edge between the archive category node and the case file relationship node, and establish a switching constraint edge between the cut status node and the case file relationship node. Write the applicable conditions and constraint weights for the deadline constraint edge, the attribution constraint edge, and the switching constraint edge to generate the segment-classification constraint graph.
[0008] Optionally, step three specifically includes: An improved LayoutLMv3 document classification model is constructed, which includes an input embedding module, a joint encoding module, an evidence localization and candidate output module, and an evidence masking contribution module. The input embedding module reads the page content index, standard page image index, text block coordinates, and layout area index from the archive data frame. It performs word segmentation on the page text to generate a text word sequence, performs image block segmentation on the standard page image to generate an image block sequence, generates a two-dimensional layout position code based on the text block coordinates and image block coordinates, and merges the text word sequence, image block sequence, and two-dimensional layout position code into a text-image-layout joint input sequence. The joint encoding module performs unified Transformer encoding on the text-image-layout joint input sequence, enabling text words and image blocks to participate in self-attention calculation in the same sequence, generating text word hiding features, image block hiding features, and page global hiding features; The evidence location and candidate output module generates evidence type probabilities based on text word hidden features, image block hidden features, and two-dimensional layout position encoding. Evidence type probabilities include title probability, document number probability, seal probability, table field probability, and non-evidence area probability. When the maximum probability among title, document number, seal, and table field probabilities reaches a preset evidence threshold and is greater than the non-evidence area probability, the corresponding text block or image block is identified as the title evidence area, document number evidence area, seal evidence area, or table field evidence area, and the maximum probability is written into the evidence area recognition credibility. When the non-evidence area probability is the highest, or when the title, document number, seal, and table field probabilities do not reach the preset evidence threshold, the corresponding text block or image block is excluded from the evidence area. An evidence area set is generated based on the title evidence area, document number evidence area, seal evidence area, and table field evidence area. The evidence location and candidate output module generates the first candidate result under the complete evidence area based on the page's global hidden features and the evidence area set. The evidence masking contribution module generates a text-image-layout joint input sequence after masking the evidence area based on the evidence area set. The text-image-layout joint input sequence after masking the evidence area is input into the same joint encoding module and the same evidence location and candidate output module to generate a second candidate result. The evidence contribution value is generated based on the difference between the first candidate result and the second candidate result.
[0009] Optionally, the evidence masking contribution module generates an evidence contribution value based on the difference between the first candidate result and the second candidate result; The first candidate result includes the first candidate cut status, the first candidate file category, the first candidate retention period, and the first candidate confidence level; The evidence areas for title, document number, seal, and table fields are to be obscured separately, one evidence area at a time, while leaving the other evidence areas unchanged. A masking mask is generated based on the coordinates of the masked evidence area. Text words in the masked evidence area are replaced with preset masking words, and image blocks in the masked evidence area are replaced with preset masking image blocks. The corresponding two-dimensional layout position code is retained, and a text-image-layout joint input sequence is generated after the masked evidence area is masked. The combined input sequence of text, image and layout after the evidence area is covered is input into the same joint encoding module and the same evidence location and candidate output module to generate a second candidate result. The second candidate result includes the second candidate piece status, the second candidate file category, the second candidate retention period and the second candidate confidence level. When the state of the first candidate cut piece is inconsistent with the state of the second candidate cut piece, the difference in cut piece state is recorded as 1; otherwise, it is recorded as 0. When the first candidate file category is different from the second candidate file category, the difference in file category is recorded as 1; otherwise, it is recorded as 0. When the retention period of the first candidate is inconsistent with that of the second candidate, the difference in retention period is recorded as 1; otherwise, it is recorded as 0. The difference between the first candidate confidence level and the second candidate confidence level is taken as the candidate confidence level change, and is set to 0 when the candidate confidence level change is less than 0. The differences in the status of the cut pieces, the differences in the archive categories, the differences in the retention period, and the changes in the candidate confidence level are weighted and summed according to preset weights to generate the evidence contribution value of the corresponding evidence area.
[0010] Optionally, step four specifically includes: Read the evidence area to identify credibility, evidence contribution value, and text block identification credibility in the archive data frame; The initial weight of evidence is obtained by multiplying the credibility of the evidence area identification and the evidence contribution value of the same evidence area. Then, the initial weight of evidence for each evidence area within the same archive object is normalized to obtain the credibility weight of evidence. The hidden features of text terms in the evidence area are used as text features, and the hidden features of image blocks in the evidence area and the corresponding two-dimensional layout position codes are fused into image layout features. The positional overlap between text features and image layout features is determined based on the coordinates of the evidence area, and the text features and image layout features are aligned according to the positional overlap. When the text block recognition confidence is lower than the preset text confidence threshold, the fusion weight of text features is reduced and the fusion weight of image layout features is increased. The text features and image layout features after region alignment are weighted and fused according to the credibility weight of the evidence to generate archive enhancement features.
[0011] Optionally, step five specifically includes: Candidate archive categories are generated based on the similarity between archive category nodes in the archive enhancement feature and segmentation-classification constraint diagram. Candidate retention periods are generated based on the retention period constraints corresponding to the candidate archive categories. Candidate segmentation status is generated based on the page order relationship, title changes, document number changes and format continuity between the current archive object and adjacent archive objects. Candidate segmentation status, candidate file category, and candidate retention period are combined into segmentation-classification candidate status, and model confidence is generated based on the probability corresponding to the candidate segmentation status, the probability corresponding to the candidate file category, and the probability corresponding to the candidate retention period. The candidate states of the segmentation-classification are matched with the segmentation-classification constraint graph. When the candidate file category cannot match the file category node, the file category node residual is 1, otherwise it is 0. When the candidate retention period cannot match the retention period node, the retention period node residual is 1, otherwise it is 0. When the candidate segmentation state cannot match the segmentation state node, the segmentation state node residual is 1, otherwise it is 0. Perform graph edge matching between the candidate states of segmentation-classification and the constraint graph of segmentation-classification. When there is no corresponding period constraint edge between the candidate file category and the candidate retention period, the period constraint residual is 1; otherwise, it is 0. When the candidate file category and the case file relationship do not satisfy the affiliation constraint edge, the affiliation constraint residual is 1; otherwise, it is 0. When the candidate segmentation status and the page sequence relationship or the case file continuity relationship do not satisfy the switching constraint edge, the switching constraint residual is 1; otherwise, it is 0. Based on the constraint weights of the corresponding edges in the segmentation-classification constraint graph, a weighted summation is performed on the residuals of the archive category node, the residual of the retention period node, the residual of the segmentation status node, the residual of the period constraint, the residual of the attribution constraint, and the residual of the switching constraint to generate the path residual.
[0012] Optionally, step six specifically includes: Sort the candidate status of cut-and-classify within the same batch of archives according to page order; Configure the segment-classification candidate state corresponding to each sorted file object as the classification state of the Viterbi dynamic programming algorithm; Configure the state cost of each classification state based on the model confidence, evidence contribution value and path residual, and configure the state transition cost between adjacent classification states based on page order relationship, piece status, case file continuity relationship and classification rule matching relationship. Establish a Viterbi recursive table and a predecessor state table. Calculate the cumulative cost of each category of state based on the state cost and state transition cost, and record the corresponding predecessor state. In the last file object, select the category state with the lowest cumulative cost as the termination state. Then, backtrack from the termination state according to the predecessor state table to obtain the segment-category state sequence with the lowest cumulative cost.
[0013] Optionally, the classification status includes candidate segmentation status, candidate file category, candidate retention period, model confidence, evidence contribution value, and path residual; The model confidence score is processed in reverse to generate a reverse model confidence score value, and the evidence contribution value is processed in reverse to generate a reverse evidence contribution value value. The reverse model confidence score value, the reverse evidence contribution value value, and the path residual are weighted and summed according to preset weights to obtain the state cost. The page order relationship generates a page order relationship cost. When the page order of adjacent file objects is continuous and there is no low confidence marker for the page order, the page order relationship cost takes a low value; otherwise, it takes a high value. The segmentation state generates a segmentation state cost. When the segmentation state of a subsequent file object is a continuation of the previous one, and the preceding and following file objects satisfy the conditions of continuous file relationship, consistent file category, and consistent retention period, the segmentation state cost is set to a low value. When the segmentation state of a subsequent file object is a continuation of the previous one, but the preceding and following file objects do not satisfy the conditions of continuous file relationship, consistent file category, or consistent retention period, the segmentation state cost is set to a high value. When the segmentation state of a subsequent file object is a newly created file with a new title, document number, or responsible entity characteristic, the segmentation state cost is set to a low value. When the segmentation state of a subsequent file object is a newly created file, and the file continuity relationship changes and satisfies the classification rule matching relationship, the segmentation state cost is set to a low value. The case file continuity relationship generates a case file continuity relationship cost. When adjacent file objects satisfy the continuity rule within the same case file, the case file continuity relationship cost takes a low value; otherwise, it takes a high value. The classification rule matching relationship generates a classification rule matching relationship cost. When the deadline constraint edge, ownership constraint edge, or switching constraint edge corresponding to the state transition matches, the classification rule matching relationship cost takes a low value; when the deadline constraint edge, ownership constraint edge, or switching constraint edge corresponding to the state transition does not match, the classification rule matching relationship cost is increased according to the constraint weight of the corresponding graph edge. The page order relationship cost, the segmentation state cost, the case file continuity relationship cost, and the classification rule matching relationship cost are weighted and summed according to preset weights to generate the state transition cost.
[0014] Optionally, step seven specifically includes: The cut-classification status sequence determines the cut status, archive category, and retention period of each archive object, and the item-level affiliation and file affiliation are determined based on the cut status to generate the final classification result; Read the cumulative cost and evidence contribution value corresponding to each category status in the segmentation-classification status sequence. When the cumulative cost exceeds the preset cost threshold or the evidence contribution value is lower than the preset contribution threshold, write the corresponding file object into the review pool. Read the review results and compare them with the final classification results. When the review results are consistent with the final classification results, mark the corresponding archive object as a confirmed sample. When the review results modify the cut status, archive category or retention period, mark the corresponding archive object as a corrected sample. For confirmed samples, the weight of the evidence region corresponding to the evidence region whose evidence contribution value is higher than the preset contribution threshold and is consistent with the final classification result is increased by a preset step size; for corrected samples, the weight of the evidence region corresponding to the evidence region whose evidence contribution value is higher than the preset contribution threshold and corresponds to the classification result before correction is decreased by a preset step size. For the correct adjacent classification state transition relationship in the confirmed sample, the corresponding state transition cost is reduced by a preset step size; for the modified adjacent classification state transition relationship in the corrected sample, the state transition cost before correction is increased by a preset step size, and the state transition cost after correction is reduced by a preset step size. When the review results confirm that the correspondence between the file category and the retention period is correct, increase the constraint weight of the corresponding period constraint edge in the cut-classification constraint graph; when the review results modify the file category, retention period, or cut status, update the file category node, retention period node, cut status node, and corresponding graph edge weight in the cut-classification constraint graph.
[0015] The beneficial effects of this invention are: This invention solves the problems of missing, duplicate, cross-document mixed layout, and low-confidence OCR content directly participating in classification when scanning documents, PDFs, and electronic documents are mixed input by preprocessing batch data of archives into archive data frames with page order relationship, text block recognition confidence, and low-confidence page order markers. This enables subsequent segmentation, classification, and retention period determination to be performed based on a unified page object and traceable page order relationship, reducing the risk of missegmentation and misclassification caused by inconsistent input formats in batch archive processing. This invention constructs a segmentation-classification constraint graph and matches the segmentation-classification candidate states generated by the enhanced archival features with archival category nodes, retention period nodes, segmentation status nodes, and period constraint edges, affiliation constraint edges, and switching constraint edges. This solves the problem that existing classification methods rely only on keywords or single-page labels and are difficult to constrain segmentation status, archival category, retention period, and file relationship simultaneously. It enables candidate classification results to be validated by rule paths and form path residuals, thereby improving the classification consistency of cross-page continuous archives and multi-file mixed archives. This invention identifies evidence areas such as titles, document numbers, seals, and table fields using an improved LayoutLMv3 archival classification model. It generates evidence contribution values based on the differences between candidate results of complete evidence areas and obscured evidence areas. Then, it incorporates model confidence, evidence contribution values, and path residuals into the state cost and state transition cost of Viterbi dynamic programming. This solves the problems of unquantifiable key evidence contributions, the susceptibility of sequence decoding to local misjudgments, and the difficulty in reverse optimization of review results. The final classification results possess evidentiary interpretability, sequence stability, and continuous updating capability, which is of great significance for improving the automation level of archival organization and reducing the cost of manual review. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an automatic classification method for archival data based on machine learning proposed in this invention; Figure 2 This is a schematic diagram of an automatic classification method for archival data based on machine learning proposed in this invention; Figure 3 This is a framework diagram of the improved LayoutLMv3 archive classification model in the automatic archive data classification method based on machine learning proposed in this invention. Detailed Implementation
[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0018] refer to Figures 1-3 An automatic classification method for archival data based on machine learning includes the following steps: Step 1: Receive batch data of archives, perform preprocessing on the batch data of archives and write the page sequence relationship to generate archive data frames; Step 2: Read the archive classification configuration data and generate a segmentation-classification constraint diagram based on archive category, file relationship, retention period rules, and segmentation status; Step 3: Input the archive data frame into the improved LayoutLMv3 archive classification model, identify the title, document number, seal and table fields as evidence areas, generate candidate results based on the complete evidence area and the obscured evidence area respectively, and generate evidence contribution value based on the difference between the two candidate results; Step 4: Based on the credibility of the evidence area and the evidence contribution value, align the text features and image layout features with the credibility of the evidence to generate enhanced archival features; Step 5: Generate segmentation-classification candidate states based on the archive enhancement features, match the segmentation-classification candidate states with the segmentation-classification constraint graph, and generate path residuals; Step 6: Configure the candidate states of segmentation-classification for the same batch of archives as the classification states of the Viterbi dynamic programming algorithm. Configure the state cost according to the model confidence, evidence contribution value and path residual. Configure the state transition cost according to the page order relationship, segmentation state, case file continuity relationship and classification rule matching relationship. Solve for the segmentation-classification state sequence with the lowest cumulative cost. Step 7: Write the final classification result according to the segmentation-classification state sequence, write the archive objects whose cumulative cost exceeds the preset cost threshold or whose evidence contribution value is lower than the preset contribution threshold into the review pool, and update the evidence area weight, state transition cost and segmentation-classification constraint graph according to the review result.
[0019] In this embodiment, step one specifically includes: Receive batch data of archives to be processed, read the source identifier, file directory level and file arrangement order of the batch data, and generate batch identifier; generate original file identifier for each original file in the batch data, and generate original file sequence number according to the receiving order, scanning order or directory order; File type identification is performed on the batch data of archives, file type tags are written, and image files, PDF files and electronic document files are uniformly converted into page-level archive objects, and page identifiers are generated for page-level archive objects; among them, page splitting processing is performed on PDF files, layout preservation conversion processing is performed on electronic document files, and page reading processing is performed on image files; Page standardization processing is performed on page-level archive objects to generate standard page images; OCR recognition and layout region recognition are performed on the standard page images to generate page text, text block coordinates, text block recognition confidence, page content index, layout region index, and page number candidate region recognition results; the layout region recognition results include title candidate regions, body text candidate regions, table candidate regions, seal candidate regions, and page number candidate regions; Write an identification confidence mark based on the text block identification confidence; generate page sequence numbers based on the original file sequence number, page number, table of contents record order and page number candidate area identification results, and write page sequence relationship marks based on the page sequence numbers; when the page number identification results are missing, duplicated or inconsistent with the page order in the original file, write a page sequence low confidence mark. The batch identifier, original file identifier, page identifier, file type marker, standard page image index, page sequence number, page content index, layout area index, text block recognition confidence, recognition confidence marker, page sequence low confidence marker, and page sequence relationship marker are encapsulated into an archive data frame; In this embodiment, the archive object is a page-level archive object; multiple consecutive page-level archive objects are grouped into the same item-level archive record or file record according to the segmentation status. The current archive object refers to the page-level archive object being processed, and adjacent archive objects refer to the page-level archive objects located before and after the current archive object in the page sequence relationship.
[0020] In this embodiment, step two specifically includes: Read the archive classification configuration data to obtain the archive category table, file rule table, retention period rule table, and component cutting rule table; perform field validation and version marking on each configuration table; when there are null values, duplicate category codes, or invalid retention periods in the configuration table, write a configuration exception flag and write the configuration exception item to the configuration validation record; configuration exception items do not participate in the component cutting-classification constraint diagram generation; Based on the archive category table, archive category nodes are generated; based on the case file rule table, case file relationship nodes are generated; based on the retention period rule table, retention period nodes are generated; and based on the segmentation rule table, segmentation status nodes are generated. The segmentation status nodes include a continuation of the previous status node, a new status node, and a new case file status node. Among them, the continuation of the previous status node corresponds to conditions such as title continuation, page number continuation, or attachment following; the new status node corresponds to conditions such as new title, new document number, or new responsible entity; and the new case file status node corresponds to conditions such as changes in case file number, changes in category level, or changes in retention period. Establish a time limit constraint edge between the archive category node and the retention period node, establish an attribution constraint edge between the archive category node and the case file relationship node, and establish a switching constraint edge between the segmentation status node and the case file relationship node; write the applicable conditions and constraint weights for the time limit constraint edge, attribution constraint edge and switching constraint edge, and generate a segmentation-classification constraint graph; Specifically, the initial constraint weight of the time limit constraint edge is set to 0.40, the initial constraint weight of the attribution constraint edge is set to 0.30, and the initial constraint weight of the switching constraint edge is set to 0.30. The basis for this setting is that the retention period rules usually have strong normative constraints, and the attribution of files and the switching of files are more affected by the page context. Therefore, the time limit constraint edge is set with a higher initial weight.
[0021] In this embodiment, step three specifically includes: An improved LayoutLMv3 archival classification model is constructed, which includes an input embedding module, a joint encoding module, an evidence localization and candidate output module, and an evidence masking contribution module. The input embedding module and the joint encoding module are constructed based on the text-image-layout joint encoding structure of LayoutLMv3, and the evidence localization and candidate output module and the evidence masking contribution module are improved parts set up for archival classification tasks. The input embedding module reads the page content index, standard page image index, text block coordinates, and layout area index from the archive data frame; performs word segmentation on the page text to generate a text word sequence; performs non-overlapping segmentation on the standard page image according to a preset image block size to generate an image block sequence; generates a two-dimensional layout position code based on the text block coordinates and image block coordinates, and merges the text word sequence, image block sequence, and two-dimensional layout position code into a text-image-layout joint input sequence. The joint encoding module performs unified Transformer encoding on the text-image-layout joint input sequence, enabling text words and image blocks to participate in self-attention calculation in the same sequence, generating text word hiding features, image block hiding features, and page global hiding features; The evidence location and candidate output module generates evidence type probabilities based on text word hidden features, image block hidden features, and two-dimensional layout position encoding. The evidence type probabilities include title probability, document number probability, seal probability, table field probability, and non-evidence area probability. When the maximum probability among title probability, document number probability, seal probability, and table field probability reaches a preset evidence threshold and is greater than the non-evidence area probability, the corresponding text block or image block is identified as the title evidence area, document number evidence area, seal evidence area, or table field evidence area, and the maximum probability is written into the evidence area recognition credibility. When the non-evidence area probability is the highest, or when the title probability, document number probability, seal probability, and table field probability all fail to reach the preset evidence threshold, the corresponding text block or image block is excluded from the evidence area. Specifically, the preset evidence threshold is set to 0.65. The basis for this setting is that when the probability of evidence type is lower than 0.65 in historical labeled samples, the proportion of text paragraphs, page numbers and ordinary table content being misidentified as titles, document numbers or table columns increases; the probability of non-evidence areas participates in the screening as a negative probability, which can reduce irrelevant text blocks from entering the evidence area set. An evidence area set is generated based on the title evidence area, document number evidence area, seal evidence area, and table column evidence area; adjacent evidence areas with the same evidence type are merged; overlapping evidence areas with different evidence types retain the area with the higher probability of evidence type, and the evidence area type, evidence area coordinates, and evidence area recognition credibility are written in. In this embodiment, the complete evidence area is the set of evidence areas that have not undergone masking processing; the masked evidence area is the input state formed by selecting an evidence area from the evidence area set and replacing the text words and image blocks in the evidence area with preset masking content. The evidence localization and candidate output module generates a first candidate result under the complete evidence area based on the page's global hidden features and the evidence area set. Specifically, the evidence localization and candidate output module performs mapping and normalization processing on the page's global hidden features to generate candidate probabilities for slice status, archive category, and retention period, respectively; and generates a first candidate slice status, a first candidate archive category, a first candidate retention period, and a first candidate confidence score based on the highest probability among the candidate probabilities, thus obtaining the first candidate result. The first candidate confidence score is obtained by weighting the candidate probabilities for slice status, archive category, and retention period. Specifically, in the first candidate confidence score, the candidate probability weight for cut status is 0.30, the candidate probability weight for archive category is 0.45, and the candidate probability weight for retention period is 0.25. The basis for this setting is that the archive category determines the archiving direction, the cut status affects the item-level attribution, and the retention period is highly constrained by the archive category. Therefore, the candidate probability weight for archive category is the highest. The evidence masking contribution module generates a text-image-layout joint input sequence after masking the evidence area based on the evidence area set. The text-image-layout joint input sequence after masking the evidence area is input into the same joint encoding module and the same evidence location and candidate output module to generate a second candidate result. The evidence contribution value is generated based on the difference between the first candidate result and the second candidate result. Specifically, the evidence masking contribution module performs masking processing on the title evidence area, document number evidence area, seal evidence area, and table column evidence area respectively; one evidence area is masked at a time, while other evidence areas remain unchanged; a masking mask is generated based on the coordinates of the masked evidence area, text words in the masked evidence area are replaced with preset masking words, image blocks in the masked evidence area are replaced with preset masking image blocks, and the corresponding two-dimensional layout position code is retained, generating a text-image-layout joint input sequence after masking the evidence area; this joint input sequence is re-input into the same joint encoding module and the same evidence positioning and candidate output module to generate the second candidate piece status, the second candidate file category, the second candidate retention period, and the second candidate confidence level, thus obtaining the second candidate result; When the status of the first candidate segment is inconsistent with that of the second candidate segment, the difference in segment status is recorded as 1; otherwise, it is recorded as 0. When the category of the first candidate archive is inconsistent with that of the second candidate archive, the difference in archive category is recorded as 1; otherwise, it is recorded as 0. When the retention period of the first candidate archive is inconsistent with that of the second candidate archive, the difference in retention period is recorded as 1; otherwise, it is recorded as 0. The difference between the confidence level of the first candidate archive and the confidence level of the second candidate archive is used as the change in candidate confidence level. When the change in candidate confidence level is less than 0, it is set to 0. The differences in cut status, file category, retention period, and candidate confidence level are weighted and summed according to preset weights to generate the evidence contribution value for the corresponding evidence area. Specifically, the weight of the difference in cut status is 0.25, the weight of the difference in file category is 0.35, the weight of the difference in retention period is 0.20, and the weight of the change in candidate confidence level is 0.20, with a sum of 1. The basis for this setting is that an incorrect file category will directly lead to an incorrect filing direction, so it is given a high weight; the cut status and retention period affect the file-level attribution and retention rules, respectively; and the change in candidate confidence level reflects the change in stability after obscuring, so it is used as a supplementary weight. Normalize the evidence contribution values of each evidence area within the same archival object to obtain the normalized evidence contribution values.
[0022] In this embodiment, step four specifically includes: Read the evidence area recognition credibility, evidence contribution value, and text block recognition credibility in the archive data frame; multiply the evidence area recognition credibility and evidence contribution value of the same evidence area to obtain the initial evidence weight; if there is evidence area weight in the configuration data, correct the initial evidence weight according to the evidence area weight; perform normalization processing on the initial evidence weight of each evidence area within the same archive object to obtain the evidence credibility weight. In this embodiment, the weight of the evidence area is the basic weight for the title evidence area, document number evidence area, seal evidence area and table column evidence area to participate in the calculation of the evidence credibility weight; before the review results are received, the weights of each type of evidence area are configured according to the preset initial value; after the review results are received, the weights of the evidence areas are updated according to the confirmed sample and the revised sample, and participate in the calculation of the evidence credibility weight in the subsequent batch processing of archives. Specifically, when the confidence level of the evidence region is lower than 0.60 or the evidence contribution value is lower than 0.12, the corresponding evidence confidence weight is multiplied by 0.50 and then normalized again. The basis for this setting is that low confidence evidence regions and low contribution evidence regions have a weak supporting effect on the classification results, and direct participation in fusion is likely to introduce noise. The hidden features of text terms in the evidence area are used as text features, and the hidden features of image blocks in the evidence area and the corresponding two-dimensional layout position codes are fused into image layout features. The positional overlap relationship between text features and image layout features is determined according to the coordinates of the evidence area, and the text features and image layout features are aligned according to the positional overlap relationship. When the confidence level of text block recognition is lower than the preset text confidence threshold, the fusion weight of text features is reduced, while the fusion weight of image layout features is increased. Specifically, the preset text confidence threshold is set to 0.80; when the confidence level of text block recognition is not lower than 0.80, the fusion weight of text features in the title evidence area and document number evidence area is set to 0.70, and the fusion weight of image layout features is set to 0.30; when the confidence level of text block recognition is lower than 0.80, the fusion weight of text features is adjusted to 0.45, and the fusion weight of image layout features is adjusted to 0.55. The basis for these settings is that in low-confidence OCR text, the seal, table lines, and layout position have higher stability in the determination of the evidence area. The text features and image layout features after region alignment are weighted and fused according to the credibility weight of the evidence to generate enhanced archival features. Specifically, the image layout feature fusion weight is increased by default for the seal evidence area and the table field evidence area, with the image layout feature fusion weight set to 0.65 and the text feature fusion weight set to 0.35; the basis for this setting is that seals and table fields are more dependent on image shape, position, and line structure.
[0023] In this embodiment, step five specifically includes: The process involves reading the archive enhancement features and the segmentation-classification constraint graph; generating candidate archive categories based on the similarity between the archive enhancement features and the archive category nodes in the segmentation-classification constraint graph. Specifically, the similarity is calculated using cosine similarity, with a preset candidate category threshold of 0.72. When multiple archive categories meet the criteria, the top three candidate archive categories are retained in descending order of similarity. This setting is based on the fact that candidate categories with similarity below 0.72 are prone to cross-category confusion, and retaining the top three candidates balances candidate recall and decoding complexity. Candidate retention periods are generated based on the retention period constraints corresponding to the candidate archive categories. When a candidate archive category corresponds to only one retention period node, that retention period node is determined as the candidate retention period. When a candidate archive category corresponds to multiple retention period nodes, the title, document number, and table column features in the archive enhancement features are combined for matching, and the retention period with the higher matching degree is selected as the candidate retention period. Candidate file segmentation states are generated based on the page sequence relationship, title changes, document number changes, and format continuity between the current file and adjacent file objects. Specifically, when the current file and the previous file have consecutive page sequences and the title similarity is higher than 0.85 or the document numbers match consecutively, a candidate file segmentation state for continuing the previous file is generated; when the title similarity is lower than 0.50 and a new document number or responsible entity characteristic appears, a candidate file segmentation state for creating a new file is generated; when the file category changes or the file relationship switches, a candidate file segmentation state for creating a new file is generated. The basis for this setting is that when the same file spans multiple pages, the title and document number usually remain highly consistent, while new files usually have new titles or document numbers. Candidate segmentation status, candidate file category, and candidate retention period are combined into segmentation-classification candidate statuses. Model confidence is then generated based on the probabilities corresponding to the candidate segmentation status, candidate file category, and candidate retention period. Specifically, the model confidence is obtained by weighting the probabilities corresponding to the candidate segmentation status, candidate file category, and candidate retention period, with weights of 0.30, 0.45, and 0.25, respectively. The weighting criteria are consistent with those of the first candidate confidence setting, ensuring consistency in the confidence scale between the candidate status generation stage and the candidate output stage. Node matching and edge matching are performed between the candidate states of segmentation-classification and the segmentation-classification constraint graph to generate path residuals. When a candidate file category cannot match a file category node, the file category node residual is 1; otherwise, it is 0. When a candidate retention period cannot match a retention period node, the retention period node residual is 1; otherwise, it is 0. When a candidate segmentation state cannot match a segmentation state node, the segmentation state node residual is 1; otherwise, it is 0. When there is no corresponding retention period constraint edge between a candidate file category and a candidate retention period, the retention period constraint residual is 1; otherwise, it is 0. When there is no hierarchical constraint edge between a candidate file category and a file relationship, the hierarchical constraint residual is 1; otherwise, it is 0. When there is no switching constraint edge between a candidate segmentation state and page order relationship or file continuity relationship, the switching constraint residual is 1; otherwise, it is 0. Based on the constraint weights of the corresponding edges in the segmentation-classification constraint graph, a weighted summation is performed on the residuals of the archive category node, retention period node, segmentation status node, period constraint residual, attribution constraint residual, and switching constraint residual to generate path residuals. Specifically, the weights of the archive category node residual, retention period node residual, and segmentation status node residual are set to 0.10, 0.10, and 0.10, respectively, while the weights of the period constraint residual, attribution constraint residual, and switching constraint residual are set to 0.25, 0.20, and 0.25, respectively. The rationale for this setting is that edge residuals reflect whether candidate states conform to archive classification rules and are better able to reflect rule conflicts than node residuals; therefore, the weight of edge residuals is higher than that of node residuals.
[0024] In this embodiment, step six specifically includes: Read the segmentation-classification candidate status, model confidence, evidence contribution value and path residual within the same archive batch, and sort the segmentation-classification candidate status according to the page order; configure the segmentation-classification candidate status corresponding to each archive object after being sorted according to the page order as the classification status of the Viterbi dynamic programming algorithm; The state cost of the classification state is configured based on model confidence, evidence contribution value, and path residual. The model confidence and evidence contribution values are inversely processed to obtain inverse values for model confidence and evidence contribution, respectively. The inverse values for model confidence, evidence contribution, and path residual are then weighted and summed according to preset weights to obtain the state cost of the classification state. Specifically, the inverse value for model confidence is 1 minus the model confidence, and the inverse value for evidence contribution is 1 minus the evidence contribution value. The weights for the inverse values for model confidence and evidence contribution are 0.40, 0.25, and 0.35, respectively. These settings are based on the principle that the state cost should reflect not only the confidence of the model output but also the stability of the evidence and the consistency of the rules. Model confidence and path residual have a more direct impact on the final state selection, hence their higher weights. The state transition cost is configured based on the page order relationship, cut status, file continuity relationship and classification rule matching relationship between adjacent file objects; the state transition cost includes page order relationship cost, cut status cost, file continuity relationship cost and classification rule matching relationship cost; In this embodiment, the case file continuity relationship indicates whether adjacent file objects satisfy the continuity rule within the same case file or the cross-case file switching rule; the classification rule matching relationship indicates whether the term constraint edge, attribution constraint edge and switching constraint edge corresponding to adjacent classification states satisfy the applicable conditions in the segmentation-classification constraint diagram. Specifically, the low-value cost is set to 0.05, and the high-value cost is set to 1.00. The basis for this setting is that the low-value cost represents a normal state transition, while the high-value cost represents a state transition that violates page order, segmentation, or classification rules. Maintaining a clear interval between the two can improve the effect of Viterbi dynamic programming in suppressing erroneous transitions. When adjacent archival objects have consecutive page sequences and no low-confidence page sequence markers exist, the page sequence relationship cost is low; when the page sequence relationship is missing, duplicated, abrupt, or has low-confidence page sequence markers, the page sequence relationship cost is high. When the cut status of the next archival object is a continuation of the previous one, and the preceding and following archival objects satisfy the conditions of continuous file relationship, consistent file category, and consistent retention period, the cut status cost is low; when the cut status of the next archival object is a continuation of the previous one, but the preceding and following archival objects do not satisfy the conditions of continuous file relationship, consistent file category, or consistent retention period, the cut status cost is high; when the cut status of the next archival object is a newly created file, and there are new titles, document numbers, or responsible entity characteristics, the cut status cost is low; when the cut status of the next archival object is a newly created file, and the file continuity relationship changes and satisfies the classification rule matching relationship, the cut status cost is low. When adjacent file objects satisfy the continuity rule within the same file, the file continuity relationship cost is set to a low value; otherwise, a high value is set. When the deadline constraint edge, attribution constraint edge, or switching constraint edge corresponding to the state transition matches, the classification rule matching relationship cost is set to a low value. When the corresponding graph edges do not match, the classification rule matching relationship cost is increased according to the constraint weight of the corresponding graph edges. Specifically, the weights of page order relationship cost, component status cost, file continuity relationship cost, and classification rule matching relationship cost are 0.20, 0.35, 0.20, and 0.25, respectively. The basis for this setting is that a component status error will directly cause a component-level attribution error, so the component status cost has the highest weight, followed by the classification rule matching relationship cost, while the page order relationship cost and the file continuity relationship cost participate in the correction as sequence continuity constraints. Establish a Viterbi recursive table and a predecessor state table. The Viterbi recursive table records the minimum cumulative cost when reaching each category state of each file object, and the predecessor state table records the category state of the previous file object corresponding to the minimum cumulative cost. Initialization is performed on the first file object, using the state costs of each category state of that file object as the initial cumulative cost, and setting the corresponding predecessor state to empty. Starting from the second file object, recursive processing is performed. For each category state of the current file object, the cumulative cost of transitioning from each category state of the previous file object to the current category state is calculated. This cumulative cost is obtained by adding the minimum cumulative cost of the previous file object's category state, the state transition cost, and the state cost of the current category state. In each category state of the previous archive object, the category state with the lowest cumulative cost is selected as the predecessor state of the current category state, and the corresponding lowest cumulative cost is written into the Viterbi recursive table, and the corresponding predecessor state is written into the predecessor state table; the recursive calculation of all archive objects in the same archive batch is completed in sequence; in each category state of the last archive object, the category state with the lowest cumulative cost is selected as the termination state; starting from the termination state, the category state corresponding to each archive object is determined by backtracking step by step according to the predecessor state table; the backtracking results are arranged in a forward order according to the page order relationship, and the segment-category state sequence with the lowest cumulative cost is obtained.
[0025] In this embodiment, step seven specifically includes: Read the segmentation-classification status sequence obtained in step six, and parse the classification status corresponding to each archive object in turn according to the page order; determine whether the current archive object is a continuation of the previous one, a new one, or a new file based on the segmentation status in the classification status, and write the corresponding classification information according to the archive category and retention period in the classification status; When the segmentation status is "Continue from the previous one," the current file object is assigned to the previous file. When the segmentation status is "Create a new one," a new file-level record is generated for the current file object. When the segmentation status is "Create a new case file," a new case file record is generated for the current file object, and a corresponding file-level record is generated under the new case file. The final classification result is generated based on the segmentation-classification status sequence. The final classification result includes batch identifier, page identifier, page order relationship, segmentation status, file category, retention period, file-level attribution relationship, and case file attribution relationship. The cumulative cost and evidence contribution value corresponding to each classification state in the segmentation-classification state sequence are read. When the cumulative cost exceeds a preset cost threshold or the evidence contribution value is lower than a preset contribution threshold, the corresponding file object is written into the review pool. Specifically, the preset cost threshold is the average cumulative cost of historical confirmed samples plus twice the standard deviation. In one embodiment, the average cumulative cost of historical confirmed samples is 0.92 and the standard deviation is 0.31, so the preset cost threshold is: 0.92 + 2 × 0.31 = 1.54, and rounded up to 1.60. The preset contribution threshold is 0.12. The basis for setting this threshold is that a high cumulative cost indicates an unstable state sequence, and a low evidence contribution value indicates that the candidate result lacks stable evidence support. The review pool records the page sequence, cut status, archive category, retention period, cumulative cost, evidence contribution value, and reason for triggering review for the archive objects to be reviewed; when the same archive object has both abnormal cumulative cost and insufficient evidence contribution value, the archive object is marked as a high-priority review object; Read the review results and compare them with the final classification results; when the review results are consistent with the final classification results, mark the corresponding archive object as a confirmed sample; when the review results modify the cut status, archive category or retention period, mark the corresponding archive object as a corrected sample, and record the cut status, archive category and retention period before and after the correction. For confirmed samples, the weights of evidence regions whose evidence contribution values exceed the preset contribution threshold and are consistent with the final classification result are increased by a preset step size. For corrected samples, the weights of evidence regions whose evidence contribution values exceed the preset contribution threshold and correspond to the classification result before correction are decreased by a preset step size. Specifically, the preset step size is 0.05. The rationale for this setting is that a single review result only reflects feedback from a small number of samples, and using a step size of 0.05 can gradually adjust the weights of the evidence regions, avoiding drastic fluctuations in weights caused by a single mislabeling. For correct adjacent classification state transition relationships in confirmed samples, the corresponding state transition cost is reduced by a preset step size. For modified adjacent classification state transition relationships in corrected samples, the state transition cost before correction is increased by a preset step size, and the state transition cost after correction is reduced by a preset step size. Specifically, the state transition cost update step size is 0.05, consistent with the evidence area weight update step size. The basis for this setting is that both the evidence area weight and the state transition cost are review feedback update parameters, and using a consistent step size facilitates version management and rollback. The segmentation-classification constraint graph is updated based on the review results. When the review results confirm the correct correspondence between the archive category and the retention period, the constraint weight of the corresponding period constraint edge in the segmentation-classification constraint graph is increased. When the review results modify the archive category, retention period, or segmentation status, the archive category node, retention period node, segmentation status node, and corresponding graph edge weights in the segmentation-classification constraint graph are updated. Specifically, each confirmed sample increases the constraint weight of the corresponding graph edge by 0.03, and each corrected sample decreases the constraint weight of the corresponding graph edge before correction by 0.03 and increases the constraint weight of the corresponding graph edge after correction by 0.03. The graph edge constraint weight is limited to between 0.05 and 0.95. The setting basis is that the constraint weight needs to reflect the review feedback trend, but should not be adjusted to 0 or 1 due to local sample feedback, so as to avoid the segmentation-classification constraint graph losing its generalization ability. Write the updated evidence region weights, state transition costs, and segment-classification constraint graph into the configuration data, and write the updated version tag; when processing the next batch of files, read the corresponding configuration data according to the updated version tag.
[0026] Example 1: To verify the feasibility of this invention in practice, it was applied to the historical project archive organization scenario of a municipal construction project archive center. The center imported batches of archive data from mixed sources, including scanned images, PDF files, and electronic documents, totaling 120 project batches and 18,600 page-level archive objects. Manual review and annotation were used to form baseline results. The archive content covered categories such as project initiation, contract management, construction process, completion acceptance, and financial settlement, with retention periods including permanent, 30-year, and 10-year. During implementation, the system preprocessed the batch archive data and wrote page sequence relationships, generating archive data frames. Based on the archive classification configuration data, a segmentation-classification constraint graph was generated, where the initial constraint weight for the term constraint edge was 0.40, the initial constraint weight for the attribution constraint edge was 0.30, and the initial constraint weight for the switching constraint edge was 0.30.
[0027] In the comparison, Method A uses file titles, document number keywords, and fixed classification rules for categorization, suitable for pages with regular layouts, complete keywords, and clear retention period rules, but it is not sensitive to continuous files spanning multiple pages, attachment pages, and page number anomalies. Method B uses ordinary LayoutLMv3 for multimodal classification of single-page files, which can improve the file category recognition effect by utilizing text and layout information, but it does not perform joint verification of the segmentation status, file continuity relationship, and retention period constraints. The unimproved model adds ordinary Viterbi sequence decoding to ordinary LayoutLMv3, which can improve the judgment of continuous files spanning multiple pages, but the state cost mainly depends on the model output probability, without introducing evidence contribution values and path residuals, which can easily lead to local pages being over-smoothed by the preceding and following states. In this invention, the preset evidence threshold is set to 0.65, and the weights of the differences in piece status, file category, retention period, and candidate confidence change are set to 0.25, 0.35, 0.20, and 0.20, respectively. When the text block recognition confidence is lower than 0.80, the text feature fusion weight is adjusted from 0.70 to 0.45, and the image layout feature fusion weight is adjusted from 0.30 to 0.55. In the state cost, the weights of the model confidence reverse value, the evidence contribution reverse value, and the path residual are set to 0.40, 0.25, and 0.35, respectively. The low value cost of the state transition cost is set to 0.05, the high value cost is set to 1.00, the preset cost threshold is set to 1.60, and the preset contribution threshold is set to 0.12.
[0028] Table 1. Comparison of the Implementation Effects of Automatic Classification of Archival Data
[0029] As shown in Table 1, Method A achieves a retention period accuracy of 90.4%, higher than Method B, indicating that fixed classification rules still have certain advantages in scenarios where retention period and file category are clearly linked. However, Method A's segmentation accuracy and file attribution consistency rate are only 84.8% and 81.9%, respectively, indicating that it struggles to handle issues such as missing page numbers, attachments, continuous cross-page layouts, and mixed file arrangement. Method B achieves a file category accuracy of 94.1%, higher than the unimproved model, indicating that the standard LayoutLMv3 has good category recognition capabilities for single-page text and layout features. However, Method B's retention period accuracy and file attribution consistency rate are only 88.7% and 85.3%, respectively, indicating that single-page classification results are difficult to stably constrain the relationship between retention period and file.
[0030] After incorporating ordinary Viterbi sequence decoding into the unimproved model, the segmentation accuracy improved to 92.3%, and the case file attribution consistency rate improved to 90.7%, indicating that sequence decoding can improve cross-page continuity judgment. However, its file category accuracy was 93.4%, slightly lower than Method B. This is because ordinary Viterbi sequence decoding tends to maintain consistency between previous and subsequent states, which may lead to over-smoothing when there is a real switch in the category of a local page. This invention quantifies the impact of the title evidence area, document number evidence area, seal evidence area, and table column evidence area on the candidate results through evidence contribution values, and constrains the consistency between the segmentation-classification candidate state and the segmentation-classification constraint graph through path residual constraints, achieving a segmentation accuracy of 95.4%, a file category accuracy of 95.6%, a retention period accuracy of 94.2%, and a case file attribution consistency rate of 94.1%.
[0031] In terms of processing efficiency, although Method A has simpler rule calculations, manual review accounts for 26.4%, with an average processing time of 91 minutes per thousand pages. Method B and the unimproved model reduce this to 74 minutes and 69 minutes respectively, but still require significant manual correction of errors in file segmentation and attribution. This invention writes archives with cumulative costs exceeding a preset cost threshold or evidence contribution values below a preset contribution threshold into a review pool, concentrating manual review on high-risk archives. This reduces the proportion of manual review to 9.8% and the average processing time per thousand pages to 59 minutes.
[0032] Based on the above implementation results, the beneficial effects of this invention are as follows: By uniformly carrying page content, page order relationships, and credibility markers within archival data frames, the consistency of mixed-source archival input is improved; the contribution of key evidence areas to candidate results is quantified through the improved LayoutLMv3 archival classification model and evidence masking contribution module; the segmentation-classification constraint graph and path residuals enable joint verification of segmentation status, archival category, retention period, and file relationships; the introduction of state cost and state transition cost through the Viterbi dynamic programming algorithm eliminates reliance on the maximum probability judgment of a single page for the final classification result; and updating the evidence area weight, state transition cost, and segmentation-classification constraint graph through review results provides continuous correction capabilities for subsequent batch archival processing, thereby improving the accuracy, consistency, and review efficiency of automatic batch archival classification.
[0033] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A machine learning-based automatic classification method for archival data, characterized in that, Includes the following steps: Step 1: Receive batch data of archives, perform preprocessing on the batch data of archives and write the page sequence relationship to generate archive data frames; Step 2: Read the archive classification configuration data and generate a segmentation-classification constraint diagram based on archive category, file relationship, retention period rules, and segmentation status; Step 3: Input the archive data frame into the improved LayoutLMv3 archive classification model, identify the title, document number, seal and table fields as evidence areas, generate candidate results based on the complete evidence area and the obscured evidence area respectively, and generate evidence contribution value based on the difference between the two candidate results; Step 4: Based on the credibility of the evidence area and the evidence contribution value, align the text features and image layout features with the credibility of the evidence to generate enhanced archival features; Step 5: Generate segmentation-classification candidate states based on the archive enhancement features, match the segmentation-classification candidate states with the segmentation-classification constraint graph, and generate path residuals; Step 6: Configure the candidate states of segmentation-classification for the same batch of archives as the classification states of the Viterbi dynamic programming algorithm. Configure the state cost according to the model confidence, evidence contribution value and path residual. Configure the state transition cost according to the page order relationship, segmentation state, case file continuity relationship and classification rule matching relationship. Solve for the segmentation-classification state sequence with the lowest cumulative cost. Step 7: Write the final classification result according to the segmentation-classification state sequence, write the archive objects whose cumulative cost exceeds the preset cost threshold or whose evidence contribution value is lower than the preset contribution threshold into the review pool, and update the evidence area weight, state transition cost and segmentation-classification constraint graph according to the review result.
2. The automatic classification method for archival data based on machine learning according to claim 1, characterized in that, Step one specifically includes: Receive batch data of archives, read the source identifier, file directory level and file arrangement order, generate batch identifier, and generate original file identifier and original file sequence number for the original file; Perform file type identification on the batch data of archives, write file type tags, and uniformly convert image files, PDF files and electronic document files into page-level archive objects, and generate page identifiers for page-level archive objects; Perform page standardization processing on page-level archive objects to generate standard page images, and perform OCR recognition and layout region recognition on the standard page images to generate page text, text block coordinates, text block recognition confidence, page content index, layout region index and page number candidate region recognition results; Write an identification confidence tag based on the identification confidence level of the text block; Page sequence numbers are generated based on the original file sequence number, page number, table of contents record order, and page number candidate area recognition results, and page sequence relationship markers are written according to the page sequence numbers; when the page number recognition results are missing, duplicated, or inconsistent with the page order in the original file, a low confidence page sequence marker is written. The batch identifier, original file identifier, page identifier, file type marker, standard page image index, page sequence number, page content index, layout area index, text block recognition confidence, recognition confidence marker, page sequence low confidence marker, and page sequence relationship marker are encapsulated into an archive data frame.
3. The automatic classification method for archival data based on machine learning according to claim 1, characterized in that, Step two specifically includes: Read the file classification configuration data to obtain the file category table, file rule table, retention period rule table, and component cutting rule table; Generate file category nodes based on the file category table, generate file relationship nodes based on the file rule table, generate retention period nodes based on the retention period rule table, and generate cut status nodes based on the cut rule table. The segmentation status node includes a continuation of the previous status node, a new status node, and a new case file status node. Establish a time limit constraint edge between the archive category node and the retention period node, establish an attribution constraint edge between the archive category node and the case file relationship node, and establish a switching constraint edge between the cut status node and the case file relationship node. Write the applicable conditions and constraint weights for the deadline constraint edge, the attribution constraint edge, and the switching constraint edge to generate the segment-classification constraint graph.
4. The automatic classification method for archival data based on machine learning according to claim 1, characterized in that, Step three specifically includes: An improved LayoutLMv3 document classification model is constructed, which includes an input embedding module, a joint encoding module, an evidence localization and candidate output module, and an evidence masking contribution module. The input embedding module reads the page content index, standard page image index, text block coordinates, and layout area index from the archive data frame. It performs word segmentation on the page text to generate a text word sequence, performs image block segmentation on the standard page image to generate an image block sequence, generates a two-dimensional layout position code based on the text block coordinates and image block coordinates, and merges the text word sequence, image block sequence, and two-dimensional layout position code into a text-image-layout joint input sequence. The joint encoding module performs unified Transformer encoding on the text-image-layout joint input sequence, enabling text words and image blocks to participate in self-attention calculation in the same sequence, generating text word hiding features, image block hiding features, and page global hiding features; The evidence location and candidate output module generates evidence type probabilities based on text word character hiding features, image block hiding features, and two-dimensional layout position encoding. The evidence type probabilities include title probability, document number probability, seal probability, table column probability, and non-evidence area probability. When the highest probability among the title probability, document number probability, seal probability, and table field probability reaches the preset evidence threshold and is greater than the probability in the non-evidence area, the corresponding text block or image block is identified as the title evidence area, document number evidence area, seal evidence area, or table field evidence area, and the highest probability is written into the evidence area to identify credibility; when the probability in the non-evidence area is the highest, or when the probability of the title, document number, seal, and table field does not reach the preset evidence threshold, the corresponding text block or image block is excluded from the evidence area. A set of evidence areas is generated based on the title evidence area, document number evidence area, seal evidence area and table column evidence area. The first candidate result under the complete evidence area is generated based on the page's global hidden features and the set of evidence areas. The evidence masking contribution module generates a text-image-layout joint input sequence after masking the evidence area based on the evidence area set. The text-image-layout joint input sequence after masking the evidence area is input into the same joint encoding module and the same evidence location and candidate output module to generate a second candidate result. The evidence contribution value is generated based on the difference between the first candidate result and the second candidate result.
5. The automatic classification method for archival data based on machine learning according to claim 4, characterized in that, The evidence masking contribution module generates an evidence contribution value based on the difference between the first candidate result and the second candidate result. The first candidate result includes the first candidate cut status, the first candidate file category, the first candidate retention period, and the first candidate confidence level; The evidence areas for title, document number, seal, and table fields are to be obscured separately, one evidence area at a time, while leaving the other evidence areas unchanged. A masking mask is generated based on the coordinates of the masked evidence area. Text words in the masked evidence area are replaced with preset masking words, and image blocks in the masked evidence area are replaced with preset masking image blocks. The corresponding two-dimensional layout position code is retained, and a text-image-layout joint input sequence is generated after the masked evidence area is masked. The combined input sequence of text, image and layout after the evidence area is covered is input into the same joint encoding module and the same evidence location and candidate output module to generate a second candidate result. The second candidate result includes the second candidate piece status, the second candidate file category, the second candidate retention period and the second candidate confidence level. When the state of the first candidate cut piece is inconsistent with the state of the second candidate cut piece, the difference in cut piece state is recorded as 1; otherwise, it is recorded as 0. When the first candidate file category is different from the second candidate file category, the difference in file category is recorded as 1; otherwise, it is recorded as 0. When the retention period of the first candidate is inconsistent with that of the second candidate, the difference in retention period is recorded as 1; otherwise, it is recorded as 0. The difference between the first candidate confidence level and the second candidate confidence level is taken as the candidate confidence level change, and is set to 0 when the candidate confidence level change is less than 0. The differences in the status of the cut pieces, the differences in the archive categories, the differences in the retention period, and the changes in the candidate confidence level are weighted and summed according to preset weights to generate the evidence contribution value of the corresponding evidence area.
6. The automatic classification method for archival data based on machine learning according to claim 4, characterized in that, Step four specifically includes: Read the evidence area to identify credibility, evidence contribution value, and text block identification credibility in the archive data frame; The initial weight of evidence is obtained by multiplying the credibility of the evidence area identification and the evidence contribution value of the same evidence area. Then, the initial weight of evidence for each evidence area within the same archive object is normalized to obtain the credibility weight of evidence. The hidden features of text terms in the evidence area are used as text features, and the hidden features of image blocks in the evidence area and the corresponding two-dimensional layout position codes are fused into image layout features. The positional overlap between text features and image layout features is determined based on the coordinates of the evidence area, and the text features and image layout features are aligned according to the positional overlap. When the text block recognition confidence is lower than the preset text confidence threshold, the fusion weight of text features is reduced and the fusion weight of image layout features is increased. The text features and image layout features after region alignment are weighted and fused according to the credibility weight of the evidence to generate archive enhancement features.
7. The automatic classification method for archival data based on machine learning according to claim 3, characterized in that, Step five specifically includes: Candidate archive categories are generated based on the similarity between archive category nodes in the archive enhancement feature and segment-classification constraint diagram. Candidate retention periods are generated based on the retention period constraints corresponding to the candidate archive categories. Candidate segmentation status is generated based on the page order relationship, title changes, document number changes and format continuity between the current archive object and adjacent archive objects. Candidate segmentation status, candidate file category, and candidate retention period are combined into segmentation-classification candidate status, and model confidence is generated based on the probability corresponding to the candidate segmentation status, the probability corresponding to the candidate file category, and the probability corresponding to the candidate retention period. The candidate states of the segmentation-classification are matched with the segmentation-classification constraint graph. When the candidate file category cannot match the file category node, the file category node residual is 1, otherwise it is 0. When the candidate retention period cannot match the retention period node, the retention period node residual is 1, otherwise it is 0. When the candidate segmentation state cannot match the segmentation state node, the segmentation state node residual is 1, otherwise it is 0. Perform graph edge matching between the candidate states of segmentation-classification and the constraint graph of segmentation-classification. When there is no corresponding period constraint edge between the candidate file category and the candidate retention period, the period constraint residual is 1; otherwise, it is 0. When the candidate file category and the case file relationship do not satisfy the affiliation constraint edge, the affiliation constraint residual is 1; otherwise, it is 0. When the candidate segmentation status and the page sequence relationship or the case file continuity relationship do not satisfy the switching constraint edge, the switching constraint residual is 1; otherwise, it is 0. Based on the constraint weights of the corresponding edges in the segmentation-classification constraint graph, a weighted summation is performed on the residuals of the archive category node, the residual of the retention period node, the residual of the segmentation status node, the residual of the period constraint, the residual of the attribution constraint, and the residual of the switching constraint to generate the path residual.
8. The automatic classification method for archival data based on machine learning according to claim 1, characterized in that, Step six specifically includes: Sort the candidate status of cut-and-classify within the same batch of archives according to the page order in the archive data frame; Each sorted file object is configured with a segmentation-classification candidate state as a classification state of the Viterbi dynamic programming algorithm, and the classification state corresponds one-to-one with the segmentation-classification candidate state. Configure the state cost for each classification state based on the model confidence, evidence contribution value, and path residual; Configure the state transition cost between adjacent classification states based on page order relationship, candidate segmentation status, case file continuity relationship and classification rule matching relationship; Establish a Viterbi recursive table and a predecessor state table. Calculate the cumulative cost of each category of state based on the state cost and state transition cost, and record the corresponding predecessor state. In the last file object, select the category state with the lowest cumulative cost as the termination state. Then, backtrack from the termination state according to the predecessor state table to obtain the segment-category state sequence with the lowest cumulative cost.
9. The automatic classification method for archival data based on machine learning according to claim 8, characterized in that, The classification status includes candidate segmentation status, candidate file category, candidate retention period, model confidence, evidence contribution value, and path residual; The page order relationship includes the page order continuity relationship between adjacent archive objects and the page order low confidence marker; The continuity relationship of the case files includes whether adjacent classification states satisfy the continuity rule within the same case file or the cross-case file switching rule; The classification rule matching relationship includes the term constraint edge matching relationship, the attribution constraint edge matching relationship, and the switching constraint edge matching relationship; The model confidence score is processed in reverse to generate a reverse model confidence score value, and the evidence contribution value is processed in reverse to generate a reverse evidence contribution value value. The reverse model confidence score value, the reverse evidence contribution value value, and the path residual are weighted and summed according to preset weights to obtain the state cost. Page order relationship cost is generated based on page order relationship. When the page order of adjacent archive objects is continuous and there is no low confidence marker for page order, the page order relationship cost takes a low value; otherwise, it takes a high value. The segmentation state cost is generated based on the candidate segmentation state. When the candidate segmentation state in the next category state is a continuation of the previous state, and the preceding and following category states satisfy the case file continuity relationship, the candidate file category is consistent, and the candidate retention period is consistent, the segmentation state cost takes the lower value. When the candidate cut-off state in the next classification state is a continuation of the previous state, but the previous and next classification states do not satisfy the case file continuity relationship, the candidate file category is consistent, or the candidate retention period is consistent, the cut-off state cost takes the higher value. When the candidate segmentation state in the next classification state is the "Create a New Item" state, and there are changes in the title or document number of the previous and next archive objects, the segmentation state cost takes a low value. When the candidate segmentation state in the next classification state is the new case file state, and the previous and next classification states satisfy the cross-case file switching rule and classification rule matching relationship, the segmentation state cost takes the lower value. When the candidate segmentation state in the next classification state is either the new item state or the new case file state, but does not meet the corresponding low value condition, the segmentation state cost takes the high value. The case file continuity cost is generated based on the case file continuity relationship. When adjacent classification states satisfy the continuity rule within the same case file or the cross-case file switching rule, the case file continuity cost takes a low value; otherwise, it takes a high value. The classification rule matching relationship cost is generated based on the classification rule matching relationship. When the deadline constraint edge matching relationship, the ownership constraint edge matching relationship, and the switching constraint edge matching relationship corresponding to the state transition are all satisfied, the classification rule matching relationship cost takes the lowest value; when any matching relationship is not satisfied, the classification rule matching relationship cost is increased according to the constraint weight of the corresponding graph edge. The page order relationship cost, the segmentation state cost, the case file continuity relationship cost, and the classification rule matching relationship cost are weighted and summed according to preset weights to generate the state transition cost.
10. The automatic classification method for archival data based on machine learning according to claim 1, characterized in that, Step seven specifically includes: The cut-classification status sequence determines the cut status, archive category, and retention period of each archive object, and the item-level affiliation and file affiliation are determined based on the cut status to generate the final classification result; Read the cumulative cost and evidence contribution value corresponding to each category status in the segmentation-classification status sequence. When the cumulative cost exceeds the preset cost threshold or the evidence contribution value is lower than the preset contribution threshold, write the corresponding file object into the review pool. Read the review results and compare them with the final classification results. When the review results are consistent with the final classification results, mark the corresponding archive object as a confirmed sample. When the review results modify the cut status, archive category or retention period, mark the corresponding archive object as a corrected sample. For confirmed samples, the weight of the evidence area corresponding to the evidence area whose evidence contribution value is higher than the preset contribution threshold and is consistent with the final classification result is increased by a preset step size. For corrected samples, the weight of the evidence area corresponding to the evidence area whose evidence contribution value is higher than the preset contribution threshold and corresponds to the classification result before correction is decreased by a preset step size. For the correct adjacent classification state transition relationship in the confirmed sample, the corresponding state transition cost is reduced by a preset step size. For the modified adjacent classification state transition relationship in the corrected sample, the state transition cost before correction is increased by a preset step size, and the state transition cost after correction is reduced by a preset step size. When the review results confirm that the correspondence between the file category and the retention period is correct, the constraint weight of the corresponding period constraint edge in the cut-classification constraint graph is increased. When the review results modify the file category, retention period, or cut status, the file category node, retention period node, cut status node, and corresponding graph edge weight in the cut-classification constraint graph are updated.
Citation Information
Patent Citations
Long text webpage generation method and device, electronic equipment and storage medium
CN113536182A
Engineering archive intelligent collection method and system based on enabling thinking
CN114117171A
A digital management method, system and storage medium based on enterprise archives
CN116663549B