AI intelligent quality detection and examination system oriented to archive digitization

By constructing technologies such as a multimodal document loader and a semantic chunking engine, the problems of low recognition accuracy and rigid rules in the archive quality inspection system have been solved, realizing automated parsing and quality review of archives, improving recognition accuracy and system security, and meeting the localization needs of sensitive archives.

CN121920377APending Publication Date: 2026-04-24JIANGSU YONGSHANQIAO ARCHIVES MANAGEMENT SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU YONGSHANQIAO ARCHIVES MANAGEMENT SERVICE CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing document quality inspection systems suffer from low accuracy in recognizing cross-page forms, handwritten signatures, and audio/video content. They also have rigid quality inspection rules, lack dynamic update capabilities, and have weak cross-language support, failing to meet the localization and security control requirements of sensitive documents.

Method used

The system constructs a multimodal document loader, a semantic chunking engine, a recursive schema feature extractor, a rule hot update platform, and a 3D quality inspection engine. Combined with role-based access control, data encryption, behavior auditing, and backup and recovery mechanisms, it achieves automated file parsing, feature extraction, and quality review, and supports private deployment.

Benefits of technology

It has enabled automated parsing and quality review of archives, improved recognition accuracy and flexibility, ensured data security, met the localization needs of sensitive archives, and improved work efficiency and system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920377A_ABST
    Figure CN121920377A_ABST
Patent Text Reader

Abstract

The invention provides an AI intelligent quality detection and review system oriented to archive digitization. The AI intelligent quality detection and review system comprises a knowledge storage module, a document loading module, a document partitioning engine module, a rule and strategy configuration module, a document key element extraction module, a quality inspection module and a privatization deployment module. The knowledge storage module is used for managing and storing archive knowledge data, the document loading module is responsible for analyzing and preloading according to archive sample types, and the rule and strategy configuration module is used for configuring rules and strategies such as classification, content detection and structural consistency required by an archive quality inspection process. The document key element extraction module is used for automatically identifying key information elements in a file and carrying out formatting extraction and structured storage, the quality inspection module is used for carrying out multi-dimensional comparison and analysis, and the privatization deployment module is used for providing a localization deployment scheme. According to the invention, the accuracy and integrity in the archive digitization process can be improved, and the workload and error rate of manual review are significantly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and archival management, and in particular relates to an AI-powered intelligent quality inspection and review system for archival digitization. Background Technology

[0002] With the accelerated digital transformation of archives, traditional manual review methods suffer from problems such as low efficiency, large errors, high costs, and difficulty in traceability. Existing archive quality inspection systems mostly adopt an "OCR + fixed rules" architecture, which has the following shortcomings:

[0003] (1) Low recognition accuracy for cross-page tables, handwritten signatures, and audio / video content;

[0004] (2) The quality inspection rules are rigid and lack dynamic updates, conflict detection and simulation capabilities;

[0005] (3) The accuracy of key element extraction is low, and cross-language and cross-text support is weak;

[0006] (4) The system is mostly deployed on the public cloud, which cannot meet the requirements of localization and security control of sensitive files.

[0007] Therefore, there is an urgent need for an intelligent archive quality inspection system that integrates multimodal AI recognition, configurable rule engine, and private deployment to achieve full automation, security, and high quality in the entire process of archive digitization. Summary of the Invention

[0008] Purpose of the Invention: This invention proposes an AI-powered intelligent quality inspection and review system for digitized archives, aiming to address issues in existing technologies such as low accuracy in multimodal archive recognition, rigid quality inspection rules, inaccurate element extraction, and insufficient data security. By constructing a multimodal document loader, a semantic chunking engine, a recursive schema element extractor, a rule hot-update platform, and a 3D quality inspection engine, the system achieves automated parsing, element extraction, and quality review of different types of archives. The system supports private deployment and, combined with role-based access control, data encryption, behavior auditing, and backup and recovery mechanisms, ensures the secure operation of archive data in the local environment.

[0009] This invention is applicable to scenarios with high requirements for archive quality and data security, such as government agencies, corporate archives, and financial institutions, and has significant technological advancements and practical value.

[0010] The system includes a knowledge storage module, a document loading module, a document chunking engine module, a rule and strategy configuration module, a document key element extraction module, a quality inspection module, and a private deployment module.

[0011] The knowledge storage module includes an archive access management submodule and an archive source management submodule. The archive access management submodule is used to perform secure uploading, updating, deletion and sharing operations on archive data based on user roles to prevent unauthorized access. The archive source management submodule is used to manage archive knowledge data from different sources and is linked with the enterprise's existing archive management system to achieve centralized management of archive metadata and quality inspection results.

[0012] The document loading module parses and preloads different types of archive samples, performs layout detection, formula detection and recognition, table extraction, OCR processing and other archive parsing tasks (such as signature recognition and key elements), automatically identifies the file format, encoding and layout characteristics of the archives, and outputs structured and standardized data for use in the quality inspection process.

[0013] The document segmentation engine module is used to build a lightweight document segmenter. Through a fusion strategy of recursive character segmentation, semantic vector breakpoint detection and large language model inference, long documents are segmented into semantically complete and contextually coherent text blocks, so as to better understand, process or store these contents.

[0014] The rules and strategies configuration module is used to support users in configuring classification, content detection, and structural consistency rules required for file quality inspection based on industry standards, regulatory requirements, or user-defined templates. It can flexibly adjust the depth and scope of review to achieve targeted quality inspection and review.

[0015] The document key element extraction module automatically identifies the title, date, number, signature and key fields in the archive based on pre-configured quality inspection rules, and performs formatted extraction and structured storage. It supports cross-language and multi-language element recognition, providing accurate input for comparison and judgment.

[0016] The quality inspection module analyzes the content consistency, completeness and accuracy of the extracted key elements based on predefined strategies, automatically generates review results and difference reports, supports the location and batch processing of abnormal data, and outputs traceable review logs to improve the quality assurance of archive digitization.

[0017] The private deployment module provides localized deployment solutions for enterprises or institutions, ensuring that AI quality inspection models and archive data operate in a secure environment. It can flexibly set deployment modes according to the user's IT infrastructure, support offline quality inspection and review in edge environments, prevent the leakage of sensitive information, and meet application scenarios with strict data security requirements.

[0018] The document access control submodule supports the uploading, updating, deletion and sharing of electronic documents and documents, and has access control to ensure information security: The document access control submodule adopts a fine-grained access control mechanism that combines role-based access control and attribute-based access control to achieve document-level access control.

[0019] The system supports hierarchical permission management based on user roles, departments, and job levels, and can dynamically adjust permission policies.

[0020] The file access control submodule employs TLS / SSL encryption during data transmission to ensure data security. For sensitive data stored in the system, a private file cabinet mechanism is used for encrypted storage, encryption algorithms are used to encrypt sensitive data, and key segmentation management is implemented. The system also establishes a data backup and recovery mechanism.

[0021] The file access management submodule has an operation recording function, which permanently records the operations of the entire knowledge management process. The recorded content includes operation time, operation type, operator identity, operation content and operation result. The recorded content is stored in an encrypted manner to ensure that it cannot be tampered with, and is archived and backed up regularly.

[0022] The file access control submodule integrates a behavior monitoring system to monitor user behavior in real time and identify abnormal operations. The system periodically detects issues such as overlapping, missing, and insufficient permissions, and optimizes permission allocation through a permission conflict detection mechanism and periodic permission review. At the same time, the system establishes a security incident response mechanism, providing real-time alarms, emergency response, and post-incident traceability functions.

[0023] The document loading module uses a document loader to preload documents, which includes a PDF Loader, an ImageLoader, a VideoLoader, and an AudioLoader.

[0024] The PDF Loader is responsible for loading document file content and identifies content including: text, titles, images, tables, table titles, formulas, and references.

[0025] The ImageLoader is responsible for recognizing the text content and description contained in the image;

[0026] The VideoLoader is responsible for extracting keyframes from video files using software or hardware acquisition devices, identifying and generating corresponding images and descriptions.

[0027] The AudioLoader is responsible for recognizing audio files, including: text, summary and topic analysis, keywords, entity and custom tag extraction.

[0028] The PDF Loader file loading process includes the following steps:

[0029] Step 1-1: Read the PDF file page by page;

[0030] Steps 1-2 involve detecting and performing OCR (Optical Character Recognition) on the text and text boxes. The OCR recognition is based on a PaddleOCR model that has been trained and fine-tuned to return the text box coordinates and recognition results. The OCR recognition includes the following steps:

[0031] Step 1-2-1: Perform text detection on the document and obtain the coordinates of the text boxes;

[0032] Step 1-2-2: Rotate and crop each text box;

[0033] Steps 1-2-3 involve text recognition of the cropped document;

[0034] Steps 1-2-4: Filter out recognition results with a confidence level lower than the specified threshold (usually set to 0.5);

[0035] Steps 1-3 involve performing layout analysis on the document image to identify different types of regions;

[0036] Steps 1-4 involve parsing the document content and processing the identified regions, specifically including the following steps:

[0037] Step 1-4-1: Use the LORE table structure recognition model to identify elements in the table region. The LORE table structure recognition model combines the spatial and logical positions of table cells for end-to-end modeling and prediction. The LORE table structure recognition model uses a keypoint segmentation network based on a convolutional neural network to extract visual features and uses two regression heads to predict the spatial and logical positions of cells respectively. Finally, the LORE table structure recognition model reconstructs the row and column structure of the entire table and outputs a machine-readable representation.

[0038] Step 1-4-2: Output the cell content and the cell's spatial and logical positions;

[0039] Steps 1-5 integrate the results of layout analysis and OCR recognition, introducing a layout restoration process to convert the document into Markdown format for output. This process includes the following steps:

[0040] Step 1-5-1: Group the tables, titles, and paragraphs output from the layout analysis;

[0041] Step 1-5-2: Merge the text within the group according to spatial and logical positions;

[0042] Step 1-5-3: Convert the merged text into Markdown (a lightweight markup language) according to its type.

[0043] The lightweight document segmenter includes a recursive character text segmenter, a text semantic segmenter, a fixed-size character text segmenter, a document structure segmenter, an LLM (Large Language Model) segmenter, and a custom segmenter. The lightweight document segmenter divides the text into small blocks.

[0044] The recursive character text segmenter pre-segments the text by specifying a block length and a set of delimiters, according to the priority of the delimiters, dividing the text into block documents with a length smaller than the original document, and then merging the block documents that do not exceed the specified block length (usually set to 1024), and recursively splitting the text that exceeds the specified block length until the desired block size is obtained.

[0045] The text semantic segmenter segments text by identifying breakpoints. The breakpoints are determined as follows: if the distance between the embedding vectors of two consecutive paragraphs exceeds a specified threshold (usually set to 0.5), a breakpoint is set between the two consecutive paragraphs to segment the text.

[0046] The fixed-size character text segmenter cuts the text into small segments of the same size according to a preset number of characters, words, or tokens, and preserves the contextual coherence by overlapping.

[0047] The document structure segmenter uses the inherent structure of the document (such as headings, chapters, tables, etc.) to divide it, with each structural unit as a block. When the content of the document chapters is of different sizes and exceeds the block size limit, it needs to be split and merged in conjunction with a recursive character text segmenter. The document structure segmenter maintains structural integrity by aligning with the logical parts of the document.

[0048] The LLM segmenter directly inputs the original document into the Large Language Model (LLM), which intelligently generates semantic blocks. By utilizing the semantic understanding capabilities of the LLM, the text is dynamically divided, ensuring the accuracy of the segmented semantics. However, this segmentation method has the highest requirements for computing power and will also pose challenges to timeliness and performance.

[0049] The custom segmenter dynamically segments the document using two or more character or strategy definitions defined by the system, for example:

[0050] The video segmentation strategy is set to split by second. This strategy will extract keyframes from the video file in units of seconds to complete the video file segmentation.

[0051] The Audio Loader file loading process includes the following steps:

[0052] Step 2-1: Load the speech recognition model and set up the hot word lexicon; the speech recognition model performs the following steps:

[0053] Step 2-1-1, Prediction Module: A predictor based on Continuous integrate-and-fire (CIF) is used to extract the acoustic feature vector corresponding to the target text, which can more accurately predict the number of target texts in the speech. Continuous integrate-and-fire (CIF) is a mechanism that accumulates continuous acoustic features and triggers the output when a threshold is reached. This method can automatically complete the alignment between acoustic features and target text without explicit alignment, and naturally predict the number of target texts by the number of triggers.

[0054] Step 2-1-2: By sampling, the acoustic feature vector and the target text vector are transformed into feature vectors containing semantic information, which are then used in conjunction with a bidirectional decoder to enhance the model's ability to model the context.

[0055] Step 2-1-3: Execute the MWER (a sequence-level loss function) training criterion based on negative sample sampling. In this step, a set of candidate sequences containing misidentified results is obtained by sampling, and the MWER loss is calculated based on the negative samples to optimize the word error rate of the model. The MWER loss is to make the model more inclined to generate whole sentence results with "lower word error rate" rather than just caring about the correctness of each word.

[0056] Step 2-2: Load the audio file, perform text recognition on the audio using a speech recognition model, and display the recognized text results.

[0057] Steps 2-3: Correct the text results;

[0058] Steps 2-4 involve summarizing and analyzing the text results, extracting keywords and entities, using prompt word engineering.

[0059] Steps 2-5: Extract system tag information contained in the text results.

[0060] The rules and policies configuration module includes a rules management unit and a policies management unit. The rules management unit is used to define and maintain the classification rules, content detection rules, and structural consistency detection rules required for archive quality inspection, and supports rule arrangement based on industry standards, laws and regulations, and user-defined templates. The policies management unit is used to set the detection depth and scope according to different business scenarios and quality inspection objectives to achieve targeted quality review, and supports dynamic loading, hot updating, and version rollback functions for rules and policies.

[0061] The rule and strategy configuration module supports the creation of multi-dimensional combination rules, including text content matching rules, text classification rules, format consistency rules, and signature verification rules. The module also supports rule priority and conflict detection mechanisms, automatically identifying and alerting to conflicting rules or strategies to prevent abnormal execution of the quality inspection process. Furthermore, the module features strategy simulation capabilities, enabling pre-simulation of strategy effects without affecting actual archival data, and outputting a preliminary review report for optimization reference.

[0062] The rules and policies configuration module has a visual configuration interface, which includes:

[0063] The rule list display area shows the name, type, scope of application, and activation status of existing rules;

[0064] The strategy editing area is used to modify, add, or delete selected rules, and to set the rules' activation conditions, execution order, and applicable scenarios.

[0065] The template management area is used to create, import, and export quality inspection templates, and supports the rapid conversion of industry standards into a set of rules that the system can execute.

[0066] The strategy preview and analysis area is used to display the expected effects after the rule is executed, the difference comparison results, and potential risk warnings.

[0067] Access control ensures that only authorized users can modify or enable rules and policies.

[0068] The document key element extraction module includes an element recognition unit, a rule parsing unit, and a structured output unit;

[0069] The element identification unit, based on pre-configured quality inspection rules, uses natural language processing and multimodal recognition technology to automatically identify key information elements in archival documents, including title, date, number, signature, key fields, responsible person, security classification, and retention period.

[0070] The rule parsing unit is responsible for parsing the quality inspection rules issued by the rule and strategy configuration module and converting the quality inspection rules into executable element extraction instructions. It supports dynamic loading of rule changes and enables the element extraction strategy to take effect in real time.

[0071] The structured output unit extracts and stores the identified key elements in a formatted and standardized manner, outputting a data object with a unified structure. It supports machine-readable formats such as JSON and XML, and has the ability to identify and output elements across languages ​​and multiple languages, providing accurate and standardized input data for comparison and analysis by the subsequent quality inspection module.

[0072] The feature identification includes the following steps:

[0073] Step 3-1: Load the PDF file using the PDF Loader, and load the image and video files using the ImageLoader and VideoLoader, then generate the corresponding text descriptions.

[0074] Step 3-2: Parse the quality inspection rules issued by the rules and strategy configuration module and convert them into entity schemas and feature extraction instructions to be extracted;

[0075] Step 3-3: Based on the text conversion results of the file, a recursive method using a display diagram guide is used to extract key entities, expressed by the formula:

[0076]

[0077] Among them, C n Let s be a set of tree schemas of depth n, t be the type of span, and x be the input text. The formula means that the input text x and the tree schema C are used to define the tree schema. n Extract a (s,t) sequence of length n. `schema` is a predefined extraction structure or constraint rule that specifies which entities to extract and the hierarchy or dependencies between them. `span` refers to a continuous substring within the text. `p((s,t)` i |(s,t) <i C n (x) represents the input text x and the extraction structure / constraint C. n And the extracted (s,t) result (s,t) is known. <i In the case of i, the span-type pair (s,t) to be extracted. i The probability of occurrence This represents the set of all span-type pairs (s,t) that need to be extracted. Let (s,t) represent the value corresponding to the i-th extraction step.i A set;

[0078] A recursive approach is adopted, given (s,t). <i In the case of t <i In C n The corresponding next-level t set is concatenated in parallel, and the text prompt information (prompt p) corresponding to the entity-type pair is constructed during the extraction at the i-th level. i The input text, along with the input text, has been fed into the encoder. By predicting the connections between tokens, the start and end boundaries of the span are determined, implementing type t. i and entity s i Extraction and matching;

[0079] Steps 3-4 involve using feature extraction commands on the original file. Leveraging the capabilities of the multimodal large model, the corresponding key entities are identified and output: Based on the predefined schema and the current entity type t, a feature extraction command prompt is constructed, including: a description of the key entity type to be identified; and the extracted upstream entities (s,t). <i As contextual conditions; cue constraints related to visual content (such as "integrate with page layout", "integrate with table cells", etc.) will input the prompt along with the text modality and visual modality into the multimodal large model;

[0080] Step 3-5, Entity fusion: The results of step 3-3 and step 3-4 are fused to output the final element recognition result;

[0081] The quality inspection module includes a consistency detection unit, an integrity detection unit, an accuracy comparison unit, and an anomaly location unit; wherein:

[0082] The consistency detection unit is used to determine the consistency of the key elements extracted from the archive in terms of content, format and structure, including cross-document version comparison, matching detection of archive metadata and actual content, and consistency verification of tabular data; it supports two modes: rule-based exact matching and semantic model-based fuzzy matching.

[0083] The integrity detection unit is used to verify whether the key information elements of the archive are missing or incomplete. It supports defining required fields and structural units in multimodal content and refers to the mandatory field constraints provided by the rule and strategy configuration module in the missing detection.

[0084] The accuracy comparison unit uses text similarity calculation, date validity verification, numbering rule verification, signature recognition and comparison algorithms to compare the identified elements with the standard values ​​in the reference data source item by item and output a difference report.

[0085] The anomaly localization unit accurately locates the detected anomalies, including position marking in the original file, coordinate output, and version difference identifier.

[0086] The consistency detection unit performs the following processing steps:

[0087] Step 4-1: Receive the structured data object output by the document key element extraction module;

[0088] Step 4-2: Invoke the content consistency rule set provided by the rules and strategy configuration module. The content consistency rule set includes text content matching rules, format consistency rules, and metadata association rules.

[0089] Step 4-3: Perform field-level alignment and comparison between the archive metadata and the actual document content;

[0090] Step 4-4: For cross-version file comparison, a version difference analysis algorithm is used to match the corresponding historical version data based on the unique identifier (such as the file number), and to detect differences through text hash value and table element hash value.

[0091] Steps 4-5: Store the consistency detection results in the detection result buffer for use by the integrity detection unit and the accuracy comparison unit;

[0092] The integrity detection unit performs the following processing steps:

[0093] Step 5-1: Load the list of required fields and the list of required structural units defined in the rules and strategy configuration module;

[0094] Step 5-2: Scan each data object output by the document key element extraction module to determine whether the required fields are empty or missing.

[0095] Step 5-3: For multimodal files, check if there are missing cells in the table, if the image is missing a corresponding description, and if the video is missing a keyframe description.

[0096] Step 5-4: If a missing item is detected, a list of missing items is generated and marked in the location information of the abnormal location unit;

[0097] The accuracy comparison unit performs the following processing steps:

[0098] Step 6-1: Based on the document content type, call the corresponding verification algorithm: for text types, call the similarity calculation algorithm (cosine similarity, SimHash, etc.); for date types, call the date rule verification algorithm; for number types, call the regular expression and check digit verification; for signature types, call the image matching algorithm and digital signature verification algorithm.

[0099] Step 6-2: Refer to the authoritative data source or reference tag library in the knowledge storage module to match and compare the extracted elements;

[0100] Step 6-3 involves calculating the corresponding numerical deviation rate for the extracted numerical or tabular elements based on reference standard data, and determining whether the deviation rate exceeds the preset allowable error range to identify any anomalies. This includes the following steps:

[0101] Step 6-3-1, for a single numerical feature, set v e v is the value to be verified extracted from the file. r For reference values ​​obtained from authoritative data sources or reference standard libraries, the deviation is calculated using the following formula:

[0102] Δ abs =|v e -v r |,

[0103]

[0104] Where Δ abs Δ rel These represent the absolute deviation rate and the relative deviation rate, respectively.

[0105] Step 6-3-2, for tabular data elements, set Extract the value for the i-th cell. Using the reference value of the i-th cell, the deviation rate of the i-th cell is calculated using the following formula.

[0106]

[0107] Step 6-3-3: For single-value elements, determine whether the deviation is within the allowable error range. Generally, for monetary values, the deviation should be less than or equal to 0.5%, and for quantity values, it should be less than or equal to 1%. For tabular elements, a combined judgment is used.

[0108]

[0109] in, Δ is the average deviation rate. max P represents the maximum deviation rate. out The percentage of cells exceeding the threshold is given by θ1, θ2, and θ3, where k is the number of cells exceeding the threshold, n is the number of cells to be compared, and θ1, θ2, and θ3 are the allowable error parameters.

[0110] Step 6-4: Output the matching score and difference details to generate a difference report;

[0111] The anomaly localization unit performs the following processing steps:

[0112] Step 7-1: Based on the list of anomalies output by the consistency detection unit, integrity detection unit, and accuracy comparison unit, locate the position of the anomaly element in the original file;

[0113] Step 7-2: For text-related exceptions, return the page number, paragraph number, and line number information;

[0114] Step 7-3: For table-type exceptions, return the table number and row and column coordinates;

[0115] Step 7-4: For image or video exceptions, return the file name and its timestamp or spatial coordinates in the multimedia file.

[0116] Step 7-5: Output the location information in a unified format for front-end presentation or batch repair operations;

[0117] The private deployment module supports the localized deployment of enterprise internal data and AI models, ensuring that data does not cross the enterprise's security boundaries. Combined with the knowledge storage module, the system provides a flexible user management mechanism, allowing enterprises to customize access permissions according to their own needs, ensuring that only authorized users can access sensitive information. A logging and auditing mechanism is introduced to monitor and record system and user behavior in real time, enhancing the system's internal control and security. System backup and recovery functions are provided to ensure data integrity and rapid system recovery in case of emergencies.

[0118] The private deployment module refers to deploying physical servers or virtualized clusters in an enterprise's own data center or server room. All components run on this local hardware. This deployment method offers the highest level of physical isolation and provides complete control over the hardware and network. Enterprises can independently define access permissions and user management mechanisms to ensure that only authorized personnel can access sensitive information, thereby improving internal control capabilities and guaranteeing the security, customizability, and maintenance costs of the knowledge management system. Deploying the system within a local network reduces or avoids public network transmission latency, making it particularly suitable for internal applications with extremely high response speed requirements. Private deployment utilizes open-source models and proprietary infrastructure, significantly reducing dependence on a single public cloud or closed-source API provider, resulting in lower migration costs.

[0119] The present invention has the following beneficial effects:

[0120] (1) Multimodal intelligent parsing and segmentation: Intelligent reading of multimodal archive files, achieving high-precision recognition and cross-page table restoration of PDF, image, audio and video archives. Through the embedding vector and LLM fusion strategy, the semantic integrity and contextual coherence of segmentation are guaranteed.

[0121] (2) Full-process automation: Integrates consistency detection, integrity detection, accuracy comparison and anomaly location mechanisms, generates difference reports and traceable review logs, realizes full-process automation of file quality inspection, supports rule priority, conflict detection, strategy hot update and simulation pre-play, and improves the flexibility and maintainability of quality inspection strategy.

[0122] (3) Private security deployment: Users have complete control over hardware and network, can define their own permission management mechanism, ensure that only authorized personnel can access sensitive information, improve internal control capabilities, and guarantee the security, customizability and maintenance cost of the system.

[0123] (4) Improve work efficiency: Overall, this invention significantly improves the efficiency of document knowledge extraction and management through automated and intelligent processing, reduces the need for manual intervention, solves the challenges faced by traditional knowledge management, transforms the static knowledge base into a dynamic intelligent hub, empowers enterprise employees, optimizes processes and improves decision-making quality. Attached Figure Description

[0124] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0125] Figure 1 This is a structural diagram of the system of the present invention.

[0126] Figure 2 This is a flowchart of the system of the present invention. Detailed Implementation

[0127] In this embodiment of the invention, an AI-powered intelligent quality inspection and review system for digitized archives is specifically provided, such as... Figure 1 As shown, it includes a knowledge storage module, a document loading module, a document chunking engine module, a rules and strategy configuration module, a document key element extraction module, a quality inspection module, and a private deployment module.

[0128] The knowledge storage module includes an archive access control submodule and a knowledge source management submodule. The archive access control submodule is used to manage and store knowledge from different sources, and the knowledge source management submodule is used to manage electronic documents and multimedia documents.

[0129] The document loading module preprocesses different types of files. By preloading files and analyzing their formats, it efficiently identifies and extracts high-quality content from complex and diverse documents. It can recognize and process electronic archives of different formats, such as MP4, MP3, PDF, Word, and Markdown, which facilitates subsequent processing.

[0130] The intelligent document chunking utilizes natural language processing technology to automatically break down long texts into smaller, more manageable text chunks by building a lightweight document chunking engine, so as to better understand, process, or store the content of these files.

[0131] The rules and policies configuration module provides rule management and policy management functions, supports setting quality inspection rules and policies according to industry standards, regulations or user-defined templates, and supports visual configuration, conflict detection and policy rehearsal.

[0132] The document key element extraction module automatically identifies key information such as title, date, number, and signature based on quality inspection rules, and outputs it in a unified structure such as JSON and XML, supporting cross-language and multi-language recognition;

[0133] The quality inspection module, based on a predefined strategy, analyzes the content consistency, completeness, and accuracy of the extracted key elements, automatically generates review results and discrepancy reports, supports the location and batch processing of abnormal data, and outputs traceable review logs, thereby improving the quality assurance of digitized archives.

[0134] The private deployment module provides a localized deployment interface for enterprise private data and AI models, ensuring system data security, guaranteeing that only authorized personnel can access sensitive information, improving internal control capabilities, and ensuring the security, customizability, and maintenance costs of the knowledge management system. Private deployment uses open-source models and proprietary infrastructure, significantly reducing reliance on a single public cloud or closed-source model API vendor, resulting in lower migration costs.

[0135] like Figure 2 The diagram shown is a flowchart of the process of this embodiment, which specifically includes the following steps:

[0136] Step 1-1: Define knowledge permissions and start the system. Enter the document loading module. This process ensures that different users' access permissions to knowledge can be distinguished and managed according to permission levels.

[0137] Steps 1-2 involve importing data from knowledge sources. This data can be electronic documents or multimedia documents (such as audio and video). Different document types are processed through different parsing channels to ensure that subsequent processing is optimized for the data type.

[0138] Steps 1-3: Read the electronic files and upload the electronic file samples that need to be inspected. For example, take "2024 Annual Procurement Contract.pdf", "Telemarketing Record.mp3", and "Warehouse Personnel Entry and Exit Record.mp4" as examples. In the text format samples, there are elements of tables, titles, and paragraphs.

[0139] Steps 1-4 involve automatically selecting a loader for preloading based on the document format. For the document "2024 Annual Procurement Contract.pdf", a PDF Loader named "MinerUParseLoader" will be used to load the document, performing pagination reading, OCR text detection and recognition (for scanned documents), and layout analysis. The document will identify elements such as the contract title, information about both parties, contract number, signing date, goods list (table), monetary clauses, and signature / signature areas. It will also identify tables, titles, paragraphs, and images. Finally, the entire PDF document will be parsed and converted into a structured Markdown data stream, which will then be sent to the next module.

[0140] Steps 1-5: Add the hot words that need attention to the hot word database, such as "Qiaochu Archives". For the document "Telemarketing Records.mp3", use AudioLoader to load it, transcribe the audio file into text and perform word correction before sending it to the next module.

[0141] Steps 1-6: For "Warehouse Personnel Entry and Exit Records.mp4", use the VideoLoader to load the video. This loader extracts keyframes from the video in seconds and outputs keyframe image descriptions in time sequence through the prompt word project to generate the main content of this file and send it to the next module.

[0142] Steps 1-7 involve segmenting and dividing the loaded document into chunks. Specifically, this includes the following steps:

[0143] Step 1-7-1: Select a suitable text segmenter according to the configuration. For example, select RecursiveCharacterMarkdownSpliter, which will segment the loaded document according to the Markdown text format.

[0144] Step 1-7-2: Set the relevant parameters for file processing, including the character list, the maximum length of a single document segment, and the overlap length of adjacent blocks. For example, set the character list to ["\n\n","\n","",""], the maximum length of a single document segment to 1024, and the overlap length of adjacent blocks to 512.

[0145] Step 1-7-3: Execute the block logic. The results of processing "2024 Annual Procurement Contract.pdf" and "Telemarketing Record.mp3" are shown in Table 1 below.

[0146] Table 1

[0147]

[0148] As you can see, the current document has been split into two parts. Markdown is used as the output format for document parsing, and the recognition result is consistent with the content of the original document.

[0149] Steps 1-8, Rule and Policy Configuration: In the rule and policy configuration module, the quality control administrator pre-configures the following rules for this type of "purchase contract":

[0150] Step 1-8-1, Integrity Rule: Must include "Contract Number", "Signing Date", "Names of Party A and Party B", "Total Amount", and "Signature / Seal";

[0151] Step 1-8-2, Consistency Rule: The "Total Amount" in the main body of the contract must be consistent with the total sum of the detailed amounts in the table; the "Signing Date" must not be later than the "Effective Date";

[0152] Step 1-8-3, Accuracy Rules: "Contract Number" must conform to the company's coding rules (e.g., `CG-2024-XXXXX`); "Party B's Name" must match the list of qualified suppliers in the knowledge base;

[0153] Step 1-8-4: Configure the above rules and set priorities through the visual interface. After the system performs conflict detection, it generates an executable quality inspection strategy package.

[0154] Steps 1-9: Key Element Extraction. The document key element extraction module receives the parsed Markdown data and quality inspection strategy package. Based on a recursive schema-guided method, it automatically extracts key fields such as "Contract Number: CG-2024-10086", "Signing Date: March 15, 2024", "Party A: XX Co., Ltd.", "Party B: YY Technology Co., Ltd.", and "Total Amount: ¥125,000.00". Simultaneously, it uses a multimodal large model to recognize the signature area image, outputting element information such as "Party A's Official Seal (Valid)" and "Party B's Official Seal (Valid)". All elements are formatted and stored as JSON objects.

[0155] Steps 1-10: Automated quality inspection. The quality inspection module calls the JSON element object to initiate multi-dimensional inspection.

[0156] Step 1-10-1, Integrity check: The check found that all required fields already exist, passing the test.

[0157] Step 1-10-2, Consistency check: The total amount of the detailed amounts in the table is calculated to be ¥125,000.00, which is consistent with the total amount in the text. Pass.

[0158] Step 1-10-3, accuracy check: "Contract Number" format verification passed; "YY Technology Co., Ltd." was searched in the supplier directory of the knowledge storage module and its status was "Under Cooperation", which passed; the signature image was compared with the reserved imprint and the similarity was 98%, which passed.

[0159] If the detection finds that the "signing date" is "February 30, 2024" (an illegal date), the accuracy comparison unit will mark this anomaly, and the anomaly location unit will locate it in the date field at the bottom right corner of page 1 in the original PDF, and highlight it in the difference report.

[0160] Steps 1-11: Result Output and Archiving. The system generates a quality inspection report, displaying "Passed Items" and "Abnormal Items," along with location information and modification suggestions. All process data, extracted elements, quality inspection logs, and original files are archived to the knowledge storage module, and metadata is synchronized with the enterprise's existing file management system.

[0161] In its specific implementation, this application provides a computer storage medium and a corresponding data processing unit. The computer storage medium is capable of storing a computer program, which, when executed by the data processing unit, can run the invention content of the multimodal AI knowledge base construction system for private deployment provided by this invention, as well as some or all of the steps in various embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0162] Those skilled in the art will clearly understand that the technical solutions in the embodiments of the present invention can be implemented using computer programs and their corresponding general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of computer programs, i.e., software products. These computer program software products can be stored in a storage medium and include several instructions to cause a device containing a data processing unit (which may be a personal computer, server, microcontroller, MUU, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.

[0163] This invention provides an AI-powered intelligent quality inspection and review system for digitized archives. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. An AI-powered intelligent quality inspection and review system for digitized archives, characterized in that: It includes a knowledge storage module, a document loading module, a document chunking engine module, a rules and strategy configuration module, a document key element extraction module, a quality inspection module, and a private deployment module; The knowledge storage module includes an archive access control submodule and an archive source control submodule. The archive access control submodule is used to perform secure uploading, updating, deleting and sharing operations on archive data based on user roles to prevent unauthorized access. The document source management submodule is used to manage document knowledge data from different sources and to link with the enterprise's existing document management system to achieve centralized management of document metadata and quality inspection results; The document loading module parses and preloads different types of archive samples, performs layout detection, formula detection and recognition, table extraction, OCR processing and other archive parsing tasks, automatically identifies the file format, encoding and layout characteristics of the archives, and outputs structured and standardized data for use in the quality inspection process. The document segmentation engine module is used to build a lightweight document segmenter. Through a fusion strategy of recursive character segmentation, semantic vector breakpoint detection and large language model inference, long documents are segmented into semantically complete and contextually coherent text blocks. The rules and strategies configuration module is used to support users in configuring classification, content detection and structural consistency rules required for archive quality inspection based on industry standards, regulatory requirements or user-defined templates. The document key element extraction module automatically identifies the title, date, number, signature and key fields in the file based on pre-configured quality inspection rules, and performs formatted extraction and structured storage. The quality inspection module analyzes the content consistency, completeness and accuracy of the extracted key elements based on predefined strategies, automatically generates review results and difference reports, supports the location and batch processing of abnormal data, and outputs traceable review logs. The private deployment module provides localized deployment solutions for enterprises or institutions, ensuring that AI quality inspection models and archive data operate in a secure environment. It can flexibly set deployment modes according to the user's IT infrastructure and supports offline quality inspection and review in edge environments.

2. The system according to claim 1, characterized in that, The document access control submodule supports the uploading, updating, deletion and sharing of electronic documents and documents, and has access control to ensure information security: The document access control submodule adopts a fine-grained access control mechanism that combines role-based access control and attribute-based access control to achieve document-level access control. The system supports hierarchical permission management based on user roles, departments, and job levels, and can dynamically adjust permission policies. The file access control submodule uses TLS / SSL encryption protocol during data transmission; for sensitive data stored in the system, the system uses a private file cabinet mechanism for encrypted storage, uses encryption algorithms to encrypt sensitive data, and implements key segmentation management; the system also establishes a data backup and recovery mechanism. The file access management submodule has an operation recording function, which permanently records the operations of the entire knowledge management process. The recorded content includes operation time, operation type, operator identity, operation content and operation result. The recorded content is stored in an encrypted manner and is archived and backed up regularly. The file access control submodule integrates a behavior monitoring system to monitor user behavior in real time and identify abnormal operations. The system periodically detects issues such as overlapping, missing, and insufficient permissions, and optimizes permission allocation through a permission conflict detection mechanism and periodic permission review. At the same time, the system establishes a security incident response mechanism, providing real-time alarms, emergency response, and post-incident traceability functions.

3. The system according to claim 2, characterized in that, The document loading module uses a document loader to preload documents, which includes a PDF Loader, an ImageLoader, a VideoLoader, and an AudioLoader. The PDF Loader is responsible for loading document file content and identifies content including: text, titles, images, tables, table titles, formulas, and references. The ImageLoader is responsible for recognizing the text content and description contained in the image; The VideoLoader is responsible for extracting keyframes from video files using software or hardware acquisition devices, identifying and generating corresponding images and descriptions. The AudioLoader is responsible for recognizing audio files, including: text, summary and topic analysis, keywords, entity and custom tag extraction.

4. The system according to claim 3, characterized in that, The PDF Loader file loading process includes the following steps: Step 1-1: Read the PDF file page by page; Steps 1-2 involve detecting and performing OCR recognition on the text and text boxes. The OCR recognition is based on a PaddleOCR model that has been trained and fine-tuned to return the text box coordinates and recognition results. The OCR recognition includes the following steps: Step 1-2-1: Perform text detection on the document and obtain the coordinates of the text boxes; Step 1-2-2: Rotate and crop each text box; Steps 1-2-3 involve text recognition of the cropped document; Steps 1-2-4: Filter out recognition results with a confidence level lower than the specified threshold; Steps 1-3 involve performing layout analysis on the document image to identify different types of regions; Steps 1-4 involve parsing the document content and processing the identified regions, specifically including the following steps: Step 1-4-1: Use the LORE table structure recognition model to identify elements in the table region. The LORE table structure recognition model combines the spatial and logical positions of table cells for end-to-end modeling and prediction. The LORE table structure recognition model uses a keypoint segmentation network based on a convolutional neural network to extract visual features and uses two regression heads to predict the spatial and logical positions of cells respectively. Finally, the LORE table structure recognition model reconstructs the row and column structure of the entire table and outputs a machine-readable representation. Step 1-4-2: Output the cell content and the cell's spatial and logical positions; Steps 1-5 integrate the results of layout analysis and OCR recognition, introducing a layout restoration process to convert the document into Markdown format for output. This process includes the following steps: Step 1-5-1: Group the tables, titles, and paragraphs output from the layout analysis; Step 1-5-2: Merge the text within the group according to spatial and logical positions; Step 1-5-3: Convert the merged text to Markdown according to its type.

5. The system according to claim 4, characterized in that, The lightweight document segmenter includes a recursive character text segmenter, a text semantic segmenter, a fixed-size character text segmenter, a document structure segmenter, an LLM segmenter, and a custom segmenter. The lightweight document segmenter divides the text into small blocks. The recursive character text segmenter pre-segments the text by specifying a block length and a set of delimiters, according to the priority order of the delimiters, and divides the text into block documents with a length smaller than the original document. Then, the block documents that do not exceed the specified block length are merged, and the text that exceeds the specified block length is recursively split until the required block size is obtained. The text semantic segmenter segments text by identifying breakpoints. The breakpoint determination method is as follows: if the distance between the embedding vectors of two consecutive paragraphs exceeds a specified threshold, a breakpoint is set at the position between the two consecutive paragraphs to perform text segmentation. The fixed-size character text segmenter cuts the text into small segments of the same size according to a preset number of characters, words, or tokens, and preserves the contextual coherence by overlapping. The document structure segmenter uses the inherent structure of the document to divide it into blocks, with each structural unit being a block. When the content of the document chapters is of different sizes and exceeds the block size limit, it needs to be split and then merged in conjunction with a recursive character text segmenter. The document structure segmenter maintains structural integrity by aligning with the logical parts of the document. The LLM segmenter directly inputs the original document into the Large Language Model (LLM), which then intelligently generates semantic blocks. The custom segmenter dynamically segments the document using two or more character or strategy-based methods defined by the system.

6. The system according to claim 5, characterized in that, The Audio Loader file loading process includes the following steps: Step 2-1: Load the speech recognition model and set up the hot word dictionary; the speech recognition model performs the following steps: Step 2-1-1, Prediction module: Use a predictor based on Continuous integrate-and-fire to extract the acoustic feature vector corresponding to the target text; Step 2-1-2: By sampling, the acoustic feature vector and the target text vector are transformed into feature vectors containing semantic information, which are then used in conjunction with a bidirectional decoder to enhance the model's ability to model the context. Step 2-1-3: Execute the MWER training criterion based on negative sample sampling. A set of candidate sequences containing incorrect identification results is obtained by sampling, and the MWER loss is calculated based on the negative samples to optimize the word error rate of the model. Step 2-2: Load the audio file, perform text recognition on the audio using a speech recognition model, and display the recognized text results. Steps 2-3: Correct the text results; Steps 2-4 involve summarizing and analyzing the text results, extracting keywords and entities, using prompt word engineering. Steps 2-5: Extract system tag information contained in the text results.

7. The system according to claim 6, characterized in that, The rules and policies configuration module includes a rules management unit and a policies management unit. The rules management unit is used to define and maintain the classification rules, content detection rules, and structural consistency detection rules required for archive quality inspection, and supports rule arrangement based on industry standards, laws and regulations, and user-defined templates. The policies management unit is used to set the detection depth and scope according to different business scenarios and quality inspection objectives to achieve targeted quality review, and supports dynamic loading, hot updating, and version rollback functions for rules and policies. The rule and strategy configuration module supports the creation of multi-dimensional combination rules, including text content matching rules, text classification rules, format consistency rules, and signature verification rules; the rule and strategy configuration module supports rule priority and conflict detection mechanisms, and can automatically identify and prompt conflicting rules or strategies; the rule and strategy configuration module has a strategy simulation function.

8. The system according to claim 7, characterized in that, The rules and policies configuration module has a visual configuration interface, which includes: The rule list display area shows the name, type, scope of application, and activation status of existing rules; The strategy editing area is used to modify, add, or delete selected rules, and to set the rules' activation conditions, execution order, and applicable scenarios. The template management area is used to create, import, and export quality inspection templates, and supports the rapid conversion of industry standards into a set of rules that the system can execute. The strategy preview and analysis area is used to display the expected effects after the rule is executed, the difference comparison results, and potential risk warnings. Access control ensures that only authorized users can modify or enable rules and policies.

9. The system according to claim 8, characterized in that, The document key element extraction module includes an element recognition unit, a rule parsing unit, and a structured output unit; The element identification unit, based on pre-configured quality inspection rules, uses natural language processing and multimodal recognition technology to automatically identify key information elements in archival documents, including title, date, number, signature, key fields, responsible person, security classification, and retention period. The rule parsing unit is responsible for parsing the quality inspection rules issued by the rule and strategy configuration module and converting the quality inspection rules into executable element extraction instructions. It supports dynamic loading of rule changes and enables the element extraction strategy to take effect in real time. The structured output unit extracts and stores the identified key elements in a formatted and standardized manner, outputting a data object with a unified structure. It supports machine-readable formats such as JSON and XML, and has the ability to identify and output elements across languages ​​and multiple languages.

10. The system according to claim 8, characterized in that, The feature identification includes the following steps: Step 3-1: Load the PDF file using the PDF Loader, and load the image and video files using the ImageLoader and VideoLoader, then generate the corresponding text descriptions. Step 3-2: Parse the quality inspection rules issued by the rules and strategy configuration module and convert them into entity schemas and feature extraction instructions to be extracted; Step 3-3: Based on the text conversion results of the file, a recursive method using a display diagram guide is used to extract key entities, expressed by the formula: Among them, C n Let s be a set of tree schemas of depth n, t be the type of span, and x be the input text. The formula means that the input text x and the tree schema C are used to define the tree schema. n Extract a (s,t) sequence of length n, where schema is a predefined extraction structure or constraint rule; span refers to a continuous substring in the text, p((s,t)). i |(s,t) <i C n (x) represents the input text x and the extraction structure / constraint C. n And the extracted (s,t) result (s,t) is known. <i In the case of i, the span-type pair (s,t) to be extracted. i The probability of occurrence This represents the set of all span-type pairs (s,t) that need to be extracted. Let (s,t) represent the value corresponding to the i-th extraction step. i A set; A recursive approach is adopted, given (s,t). <i In the case of t <i In C n The corresponding next-level t set is concatenated in parallel, and the text prompt information (prompt p) corresponding to the entity-type pair is constructed during the extraction at the i-th level. i The input text, along with the input text, has been fed into the encoder. By predicting the connections between tokens, the start and end boundaries of the span are determined, implementing type t. i and entity s i Extraction and matching; Steps 3-4 involve using feature extraction commands on the original file. Leveraging a multimodal large model, the corresponding key entities are identified and output: Based on the predefined schema and the current entity type t, a feature extraction command prompt is constructed, including: a description of the key entity type to be identified; and the extracted upstream entities (s,t). <i As a contextual condition; cue constraints related to visual content, the prompt is input into the multimodal large model along with the text modality and the visual modality; Step 3-5, Entity fusion: The results of step 3-3 and step 3-4 are fused to output the final element recognition result; The quality inspection module includes a consistency detection unit, an integrity detection unit, an accuracy comparison unit, and an anomaly location unit; wherein: The consistency detection unit is used to determine the consistency of the key elements extracted from the archive in terms of content, format and structure, including cross-document version comparison, matching detection of archive metadata and actual content, and consistency verification of tabular data; it supports two modes: rule-based exact matching and semantic model-based fuzzy matching. The integrity detection unit is used to verify whether the key information elements of the archive are missing or incomplete. It supports defining required fields and structural units in multimodal content and refers to the mandatory field constraints provided by the rule and strategy configuration module in the missing detection. The accuracy comparison unit uses text similarity calculation, date validity verification, numbering rule verification, signature recognition and comparison algorithms to compare the identified elements with the standard values ​​in the reference data source item by item and output a difference report. The anomaly localization unit accurately locates the detected anomalies, including position marking in the original file, coordinate output, and version difference identifier. The consistency detection unit performs the following processing steps: Step 4-1: Receive the structured data object output by the document key element extraction module; Step 4-2: Invoke the content consistency rule set provided by the rules and strategy configuration module. The content consistency rule set includes text content matching rules, format consistency rules, and metadata association rules. Step 4-3: Perform field-level alignment and comparison between the archive metadata and the actual document content; Step 4-4: For cross-version file comparison, a version difference analysis algorithm is used to match the corresponding historical version data based on the unique identifier, and to detect differences through text hash value and table element hash value. Steps 4-5: Store the consistency detection results in the detection result buffer for use by the integrity detection unit and the accuracy comparison unit; The integrity detection unit performs the following processing steps: Step 5-1: Load the list of required fields and the list of required structural units defined in the rules and strategy configuration module; Step 5-2: Scan each data object output by the document key element extraction module to determine whether the required fields are empty or missing. Step 5-3: For multimodal files, check if there are missing cells in the table, if the image is missing a corresponding description, and if the video is missing a keyframe description. Step 5-4: If a missing item is detected, a list of missing items is generated and marked in the location information of the abnormal location unit; The accuracy comparison unit performs the following processing steps: Step 6-1: Call the corresponding verification algorithm according to the type of document content: for text type, call the similarity calculation algorithm; for date type, call the date rule verification algorithm; for number type, call the regular expression and check digit verification; for signature type, call the image matching algorithm and digital signature verification algorithm. Step 6-2: Refer to the authoritative data source or reference tag library in the knowledge storage module to match and compare the extracted elements; Step 6-3 involves calculating the corresponding numerical deviation rate for the extracted numerical or tabular elements based on reference standard data, and determining whether the deviation rate exceeds the preset allowable error range to identify any anomalies. This includes the following steps: Step 6-3-1, for a single numerical feature, set v e v is the value to be verified extracted from the file. r For reference values ​​obtained from authoritative data sources or reference standard libraries, the deviation is calculated using the following formula: Δ abs =|v e -v r |, Where Δ abs Δ rel These represent the absolute deviation rate and the relative deviation rate, respectively. Step 6-3-2, for tabular data elements, set Extract the value for the i-th cell. Using the reference value of the i-th cell, the deviation rate of the i-th cell is calculated using the following formula. Step 6-3-3: For single-value features, determine whether the deviation is within the allowable error range; for tabular features, use a combined determination: in, Δ is the average deviation rate. max P represents the maximum deviation rate. out The percentage of cells exceeding the threshold is given by θ1, θ2, and θ3, where k is the number of cells exceeding the threshold, n is the number of cells to be compared, and θ1, θ2, and θ3 are the allowable error parameters. Step 6-4: Output the matching score and difference details to generate a difference report; The anomaly localization unit performs the following processing steps: Step 7-1: Based on the list of anomalies output by the consistency detection unit, integrity detection unit, and accuracy comparison unit, locate the position of the anomaly element in the original file; Step 7-2: For text-related exceptions, return the page number, paragraph number, and line number information; Step 7-3: For table-type exceptions, return the table number and row and column coordinates; Step 7-4: For image or video exceptions, return the file name and its timestamp or spatial coordinates in the multimedia file. Step 7-5: Output the location information in a unified format for front-end presentation or for batch repair operations.