An intelligent AI-driven archive digitization full-process processing system

The intelligent AI-driven end-to-end archival digitization system, employing a multimodal fusion architecture and multi-dimensional quality verification, solves the problems of fragmented processes, reliance on manual labor, insufficient recognition accuracy, and poor scenario adaptability in archival digitization, achieving efficient and accurate automated archival processing and value extraction.

CN121330702BActive Publication Date: 2026-04-21LIAONING HONGTU CHUANGZHAN SURVEYING & MAPPING CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LIAONING HONGTU CHUANGZHAN SURVEYING & MAPPING CO
Filing Date
2025-12-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing archival digitization solutions suffer from fragmented processes, high reliance on manual intervention, insufficient accuracy in identification and recording, inefficient quality control, poor adaptability to different scenarios, and inefficient digital resource management, making it difficult to achieve efficient and accurate automated processing throughout the entire process.

Method used

This intelligent AI-driven end-to-end archival digitization system comprises a physical digitization module, an AI agent collaborative network, and a digital resource management module. Through multi-agent collaboration, it automates the process from physical archives to a structured digital knowledge base. The content recognition and cataloging agent employs a visual-language multimodal fusion architecture, the quality verification and repair agent establishes multi-dimensional verification checkpoints, and the digital resource management module constructs a knowledge graph for multi-dimensional retrieval.

Benefits of technology

It achieves highly efficient automated processing without human intervention, improves the consistency and accuracy of the archive processing process, supports scenario adaptation for multiple types of archives, meets the precise query needs in complex scenarios, and enhances the information credibility and value mining capabilities of archives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330702B_ABST
    Figure CN121330702B_ABST
Patent Text Reader

Abstract

This invention relates to the field of digital system technology, specifically disclosing an intelligent AI-driven end-to-end archival digitization system. The system comprises a physical digitization module, an AI agent collaborative network, and a digital resource management module that work sequentially and collaboratively, enabling automated processing from physical archives to a structured digital knowledge base. The physical digitization module is configured to perform high-fidelity scanning or photography of physical archives, generating a set of original digital images. The AI ​​agent collaborative network is a multi-agent system managed and scheduled by a central coordinator. Through the sequential collaboration of the physical digitization module, the AI ​​agent collaborative network, and the digital resource management module, this system constructs an end-to-end closed-loop processing flow from physical archives to a structured digital knowledge base. No manual intervention is required at each stage; manual review is only necessary in a very few complex and abnormal scenarios, significantly reducing labor costs and improving processing efficiency and process consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital system technology, specifically to an intelligent AI-driven end-to-end digital archival processing system. Background Technology

[0002] With the deepening of information technology construction, the digital transformation of physical archives (such as paper documents, engineering drawings, audio and video materials, etc.) has become a core requirement for various industries to improve the efficiency of archive management and explore the value of archives. If physical archives are stored in traditional ways for a long time, they not only occupy a lot of physical space and are easily damaged by environmental factors (such as temperature, humidity, and pests), but also suffer from problems such as low retrieval efficiency and difficulty in sharing. They cannot meet the requirements of modern management for rapid access, accurate matching, and correlation analysis of archives. Therefore, the digitization of archives has become an inevitable trend.

[0003] However, existing archival digitization solutions still suffer from numerous technical bottlenecks, making it difficult to achieve efficient and accurate automated processing throughout the entire process.

[0004] 1. Fragmented processes and high reliance on manual intervention: Existing systems are mostly based on a modular splicing model. For example, the processes of entity scanning, image processing, content recognition, and quality verification often require manual coordination, which not only increases labor costs but also easily reduces processing efficiency and consistency due to human error (such as manual pagination and manual addition of bibliographic entries).

[0005] 2. Insufficient recognition and cataloging accuracy, poor cross-modal information coordination: In traditional digitization solutions, text recognition (OCR) and structured cataloging are often independent of each other, lacking effective integration of visual features and text information. This can easily lead to problems such as mismatch between text recognition and image content, missing cataloging items, or logical contradictions. The processing effect is even worse for archives containing special elements (such as official seals, signatures, and engineering symbols).

[0006] 3. Inefficient quality control and repair mechanisms: Existing quality checks are mostly limited to a single dimension (such as checking only image clarity or text fluency), lacking multi-dimensional collaborative check logic; and the repair of non-conforming items mostly relies on manual judgment, making it difficult to achieve targeted and efficient intelligent repair, resulting in some digital archives being unable to be effectively reused due to substandard quality.

[0007] 4. Weak scenario adaptability and poor compatibility with multiple types of archives: Most systems are designed only for a single type of archive (such as paper documents), and lack dedicated processing strategies for special archives such as engineering drawings and audio and video (such as vectorization of engineering drawings and extraction of key information from audio and video), which cannot meet the digitization needs of multiple scenarios.

[0008] 5. Inefficient digital resource management and insufficient value mining capabilities: Existing systems mostly adopt a simple storage and metadata retrieval model for digitized archives, lacking in-depth analysis of the relationships between archive entities, events, and connections. This makes it difficult to achieve semantic retrieval and correlation analysis in complex scenarios, and fails to fully release the potential value of archives. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides an intelligent AI-driven end-to-end archival digitization system, which solves the problems mentioned in the background technology.

[0010] To achieve the above objectives, the present invention provides the following technical solution: an intelligent AI-driven full-process system for digitizing archives, comprising a physical digitization module, an AI intelligent agent collaborative network, and a digital resource management module that work in sequence to achieve automated processing from physical archives to a structured digital knowledge base;

[0011] The physical digitization module is configured to perform high-fidelity scanning or photography of physical archives to generate an original digital image set.

[0012] The AI ​​agent collaborative network is a multi-agent system managed and scheduled by a central coordinator. It receives the original digital image set and sequentially calls the image preprocessing agent, content recognition and recording agent, and quality verification and repair agent for pipelined processing; wherein:

[0013] The image preprocessing intelligent agent performs automatic skew correction, noise reduction, and page segmentation operations on the original digital image, and outputs a standard page image.

[0014] The content recognition and recording intelligent agent, based on a fusion model of large language model and computer vision, performs full-text OCR recognition, layout analysis, key field extraction and structured recording on standard page images to generate a primary digital object containing text content and metadata.

[0015] The quality verification and repair agent performs multi-dimensional verification of the image quality, text recognition accuracy, and bibliographical integrity of the primary digital object, and initiates automatic or semi-automatic repair processes for recognition errors or missing items, outputting high-quality digital archive objects.

[0016] The digital resource management module receives the high-quality digital archive objects, automatically classifies, indexes, and stores them in the database, and establishes a multi-dimensional retrieval index, ultimately forming a manageable and exploitable digital archive resource database.

[0017] Preferably, the content recognition and transcription agent adopts a Transformer-based vision-language multimodal fusion architecture, and the specific implementation process is as follows:

[0018] The visual encoder employs a 6-10 layer Transformer encoder structure with 4-6 attention heads. The hidden layers of the Feed-Forward network have a dimension of 1024-1536, and the GELU activation function is used. A sliding window is used to extract global layout features and local text region features from the page image, generating a visual feature sequence. ,in It consists of 256-512 dimensional feature vectors, which fuse image texture, text outline, and region location information;

[0019] The text decoder employs a 6-10 layer Transformer decoder structure, based on the aforementioned visual feature sequence. With preset archival description framework prompts, the system distinguishes between full-text transcription and structured description tasks using intent type masks, and generates two parts of output autoregressively:

[0020] Complete transcribed text It covers all text content and formatting marks on the page;

[0021] Structured cataloging information It is a set of key-value pairs. ,in For the preset bibliographic item name, The corresponding content extracted from an image or text;

[0022] Cross-modal fusion mechanism: Achieves bidirectional interaction between visual and textual features through a cross-attention layer. The decoder generates a bibliographic entry for each entry. All of these will call the local features corresponding to the visual encoder. Verification is performed to ensure consistency of information across modalities.

[0023] Preferably, when generating structured bibliographic information, the content recognition and bibliographic agent introduces a field confidence evaluation mechanism for any bibliographic item. Its value confidence level Calculated by the following formula:

[0024] ;

[0025] in, For the bibliographic item The relevant visual context feature vectors have the same dimensions as the visual encoder output, ranging from 256 to 512, and are composed of text region features and surrounding layout features. This is the visual feature weight matrix, with dimensions 1×D, where D is... Dimension These are bias terms, with a dimension of 1×1, and are all trainable parameters optimized during model training; For candidate text fragments, for In full-text transcription The context embedding representation in the image has a dimension of 768. Let cosine similarity function be used to calculate... and The semantic relevance is [0,1], and its value range is [0,1]. The balancing coefficient, with a value range of [0.3, 0.7], is used to adjust the weights of visual and textual evidence. It is dynamically optimized based on the transcription accuracy during the training phase.

[0026] Preferably, the quality verification and repair agent establishes three-level quality verification checkpoints, and the scores of each checkpoint are mapped to the [0,1] interval through normalization processing. The specific process is as follows:

[0027] Image quality control: Based on four indicators—sharpness, contrast, noise density, and edge integrity—a comprehensive readability score is calculated through weighted summation. The formula is:

[0028] ;

[0029] in The weights range from [0.2, 0.3]. For image clarity, For contrast, For noise density, For edge integrity; when At that time, the image quality was deemed unacceptable;

[0030] Text quality control: for transcribed text The system employs a dual validation process: first, format validation based on an archive format rule base; and second, semantic fluency validation based on a language model, resulting in a text credibility score. for:

[0031] ;

[0032] in The format compliance rate is represented by a value of [0,1]. This is the perplexity normalization value, taking values ​​[0,1].

[0033] Bibliographic quality control: combining field confidence levels Verify structured bibliographic information using the bibliographic rule base. Completeness and logical consistency, cataloging completeness score for:

[0034] ;

[0035] in To preset the total number of bibliographic entries, This is a logic check function.

[0036] Preferably, the quality verification and repair agent initiates an intelligent repair loop for items that fail verification. This loop aims to achieve quality compliance with the minimum number of iterations. The specific process is as follows:

[0037] The repair trigger conditions are as follows:

[0038] when or ( )or ( When this occurs, the repair process will be initiated;

[0039] Targeted repair strategy: For low confidence levels, , For missing or incomplete bibliographic entries, a reinforcement prompt request is sent to the content recognition and bibliographic agent. The prompt information includes the bibliographic entry name, context fragments, and visual region coordinates. Based on the prompt, the agent re-extracts features and generates... Recalculate after update ;

[0040] for and On the page, a parameter optimization request is sent to the image preprocessing agent to adjust the denoising intensity and skew correction accuracy. After re-outputting the standard page image, the content recognition and recording agent is triggered for secondary processing.

[0041] For items that still fail to meet the standard after two iterations, or Mark these as items for manual review and generate a review report;

[0042] Experience feedback mechanism: The input data, processing parameters, output quality score and repair results of each iteration of the repair loop are stored in the experience pool of the central coordinator, which is used to optimize the model parameters and decision-making strategies of each agent.

[0043] Preferably, the central coordinator maintains the global mass entropy. The metric, used to dynamically evaluate and optimize the processing flow of single and batch files, is defined as the average quality uncertainty of all file objects in the current processing batch, and the formula is:

[0044] ;

[0045] in: This represents the number of files within the batch, ranging from 10 to 1000. For the first The file is in The scores on the quality dimension have been normalized to [0,1]. The weights for each quality dimension satisfy... The default value is , , This can be adjusted according to the type of file;

[0046] Optimization logic: The central coordinator aims to minimize... With the goal of dynamically adjusting the batch processing order, agent invocation priority, and parallel processing granularity.

[0047] Preferably, each agent in the AI ​​agent collaborative network adopts a centralized training and distributed execution paradigm, specifically implemented as follows:

[0048] Training phase: Construct a labeled multimodal dataset, adopt an end-to-end joint training mode, and optimize the objective function as follows:

[0049] ;

[0050] in, For text recognition, cross-entropy loss, Due to the loss in matching bibliographic information, For quality score regression loss, , The loss balance coefficient is used; the training optimizer employs AdamW with a learning rate of [missing value]. The training sessions consist of 20-30 rounds.

[0051] Execution phase: Each agent is deployed on an independent containerized service node and responds to the central coordinator's scheduling instructions via RESTful API; the central coordinator dynamically allocates tasks based on the task queue length and agent load rate, supporting parallel processing of multiple batches of files, with a single batch processing latency of ≤3 seconds / file.

[0052] Preferably, before deploying the intelligent AI-driven full-process digitization system for archives, the core content recognition and cataloging intelligent agent undergoes domain-adaptive pre-training, the specific process of which is as follows:

[0053] Pre-trained dataset: Contains over 500,000 unlabeled archival image-text pairs and over 100,000 archival data sets with labeled layout elements;

[0054] Pre-training task:

[0055] Archival document masking language modeling: Randomly mask 15% of the token in the text, train the model to predict the mask content based on visual features and context, and use cross-entropy as the loss function;

[0056] Page element relationship prediction: Mark the relevant blocks of titles, paragraphs, tables, and seals in the document, train the model to predict the relative positional relationship and semantic association between different blocks, and use the loss function as binary cross-entropy.

[0057] Cross-modal cataloging alignment: Construct a cataloging item-image region mapping dataset, train a model to learn the mapping relationship from a specific visual region to the corresponding cataloging item, and use contrastive learning loss as the loss function;

[0058] Pre-training parameters: fine-tuned based on an open-source multimodal large model, with 10-15 training epochs, a batch size of 32, and a learning rate of [missing information]. The weight decay coefficient is .

[0059] Preferably, the intelligent AI-driven end-to-end archival digitization system has scenario-based adaptability for different types of archives, and the specific adjustment strategies are as follows:

[0060] Paper document archives: The intelligent agent for content recognition and recording enhances the detection of special visual elements related to official seals, signatures, and red-headed characters. It uses the YOLOv8 model to extract special element features separately, and then merges them with text features to generate recording information.

[0061] Engineering drawing archives: The image preprocessing agent integrates a vectorization process to convert raster images into SVG vector format; the content recognition agent is adapted to recognize chart symbols, adding relevant bibliographic entries for drawing numbers, scales, and dimensions; and the visual encoder is optimized to a feature extraction structure for lines and symbols.

[0062] Photo and audio / video archives: The content recognition and recording intelligent agent has been expanded into a multimedia content analysis model. Photo archives support scene description and intelligent recognition, while audio / video archives support speech-to-text conversion and extraction of key event timelines. The digital resource management module has added a multimedia format adaptation interface, supporting video frame indexing and audio segment retrieval.

[0063] Preferably, the automatic classification and indexing function integrated into the digital resource management module is driven by an independent intelligent agent constructing an archival knowledge graph, specifically implemented as follows:

[0064] Knowledge Extraction: Intelligent Agents Analyze Structured Descriptive Information With full text content Extract archival entities, event entities, and their relationships;

[0065] Knowledge graph construction: Based on the Neo4j graph database, it dynamically links to the preset archive domain ontology, and supports automatic supplementation of entity attributes and relationship expansion;

[0066] Semantic retrieval implementation: Based on knowledge graphs, a semantic index is built, supporting three types of retrieval methods:

[0067] Metadata retrieval: based on exact matching of bibliographic entries;

[0068] Full-text search: based on text similarity matching;

[0069] Related retrieval: based on entity relationship path matching;

[0070] Knowledge update mechanism: The entities and relationships of newly added files are automatically integrated into the knowledge graph, and the graph is regularly cleaned up for redundancy and relationships are strengthened to ensure retrieval accuracy.

[0071] This invention provides an intelligent AI-driven end-to-end archival digitization system, which has the following beneficial effects:

[0072] 1. This system constructs an end-to-end closed-loop processing flow from physical archives to a structured digital knowledge base through the sequential collaboration of the physical digitization module, the AI ​​intelligent agent collaborative network, and the digital resource management module. No manual intervention is required in the connection between each link. Manual review is only required in a very few complex and abnormal scenarios, which significantly reduces labor costs and improves processing efficiency and process consistency.

[0073] 2. In the AI ​​intelligent agent collaborative network, the content recognition and recording intelligent agent adopts a visual-language multimodal fusion architecture. It realizes bidirectional interactive verification of visual features and text features through a cross-attention layer. At the same time, it introduces a field confidence assessment mechanism to ensure the accuracy of full-text transcription and structured recording. Combined with the multi-dimensional verification of the quality verification and repair intelligent agent, it further reduces recognition errors and recording contradictions, and improves the information credibility of digital archives.

[0074] 3. The quality verification and repair intelligent agent establishes three-level quality verification checkpoints for images, text, and bibliographic information to achieve comprehensive quality assessment. For items that do not meet the standards, an intelligent repair loop is activated through targeted repair strategies (such as optimizing preprocessing parameters and strengthening bibliographic information prompts) to reduce manual repair costs. At the same time, repair experience is accumulated through an experience feedback mechanism to continuously optimize the processing effect and ensure the output of high-quality digital archive objects.

[0075] 4. Design customized processing strategies for different types of archives such as paper documents, engineering drawings, and photos / audio / video files; for example, enhance the detection of special visual elements in paper documents, add vectorized process and symbol recognition to engineering drawings, and expand multimedia content analysis functions for audio / video archives, breaking the scene limitations of traditional systems and meeting the digital processing needs of multiple industries and multiple types of archives.

[0076] 5. The digital resource management module builds intelligent agents based on the archival knowledge graph, enabling in-depth extraction and association of archival entities, events, and relationships, and constructing a structured knowledge graph. It supports multiple retrieval methods such as metadata retrieval, full-text retrieval, and association retrieval, which can meet the precise query needs in complex scenarios. At the same time, it continuously improves the graph through a knowledge update mechanism, providing support for the association analysis and value mining of archives, and helping archives to transform from storage carriers into valuable resources.

[0077] 6. The central coordinator dynamically evaluates the batch processing effect through the global quality entropy index, and adjusts the batch order, agent resource allocation and parallel processing granularity in real time to ensure that the system can maintain a stable and efficient processing state when facing archive batches of different sizes and qualities. Moreover, the experience accumulation of each agent and the model optimization mechanism enable the system's processing capacity to continuously improve over time and adapt to long-term digitization needs. Attached Figure Description

[0078] Figure 1 This is a block diagram illustrating the principle of an intelligent AI-driven end-to-end archival digitization system as described in this invention.

[0079] Figure 2 This is a block diagram illustrating the AI ​​agent collaborative network principle of an intelligent AI-driven end-to-end archival digitization system as described in this invention. Detailed Implementation

[0080] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0081] like Figures 1-2 As shown, the present invention provides a technical solution: an intelligent AI-driven full-process processing system for digitizing archives, comprising a physical digitization module, an AI intelligent agent collaborative network, and a digital resource management module that work in sequence to achieve automated processing from physical archives to a structured digital knowledge base;

[0082] The physical digitization module is configured to perform high-fidelity scanning or photography of physical archives to generate an original digital image set.

[0083] The AI ​​agent collaborative network is a multi-agent system managed and scheduled by a central coordinator. It receives the original digital image set and sequentially calls the image preprocessing agent, content recognition and recording agent, and quality verification and repair agent for pipelined processing; wherein:

[0084] The image preprocessing intelligent agent performs automatic skew correction, noise reduction, and page segmentation operations on the original digital image, and outputs a standard page image.

[0085] The content recognition and recording intelligent agent, based on a fusion model of large language model and computer vision, performs full-text OCR recognition, layout analysis, key field extraction and structured recording on standard page images to generate a primary digital object containing text content and metadata.

[0086] The quality verification and repair agent performs multi-dimensional verification of the image quality, text recognition accuracy, and bibliographical integrity of the primary digital object, and initiates automatic or semi-automatic repair processes for recognition errors or missing items, outputting high-quality digital archive objects.

[0087] The digital resource management module receives the high-quality digital archive objects, automatically classifies, indexes, and stores them in the database, and establishes a multi-dimensional retrieval index, ultimately forming a manageable and exploitable digital archive resource database.

[0088] More specifically, the content recognition and transcription agent adopts a Transformer-based vision-language multimodal fusion architecture, and the specific implementation process is as follows:

[0089] The visual encoder employs a 6-10 layer Transformer encoder structure with 4-6 attention heads. The hidden layers of the Feed-Forward network have a dimension of 1024-1536, and the GELU activation function is used. A sliding window is used to extract global layout features and local text region features from the page image, generating a visual feature sequence. ,in It consists of 256-512 dimensional feature vectors, which fuse image texture, text outline, and region location information;

[0090] The text decoder employs a 6-10 layer Transformer decoder structure, based on the aforementioned visual feature sequence. With preset archival description framework prompts, the system distinguishes between full-text transcription and structured description tasks using intent type masks, and generates two parts of output autoregressively:

[0091] Complete transcribed text It covers all text content and formatting marks on the page;

[0092] Structured cataloging information It is a set of key-value pairs. ,in Preset the names of the bibliographic entries (such as the issuing authority, the date of issuance, and the security classification). The corresponding content extracted from an image or text;

[0093] Cross-modal fusion mechanism: Achieves bidirectional interaction between visual and textual features through a cross-attention layer. The decoder generates a bibliographic entry for each entry. All of these will call the local features corresponding to the visual encoder. Verification is performed to ensure consistency of information across modalities.

[0094] The identification and recording agent employs a Transformer-based visual-language multimodal fusion architecture. Through a layered visual encoder and text decoder, it achieves deep synergy between visual information from archival images and semantic information from text. The visual encoder, with its 6-10 layer Transformer structure and sliding window feature extraction method, can simultaneously capture the detailed features of the global layout of the page and local text regions (fusing texture, contour, and position information), generating a visual feature sequence. (in The feature vectors are 256-512 dimensional, providing high-dimensional and multi-dimensional visual evidence for subsequent recognition and cataloging. The text decoder, relying on intent type masking technology, can accurately distinguish between full-text transcription and structured cataloging tasks, and autoregressively generates complete transcribed text covering the full text and format. and structured record information in key-value pair format that conforms to archival management standards. This avoids the problem of "disconnect between recognition and recording" in traditional processing; while the cross-modal fusion mechanism, through the bidirectional interaction of the cross-attention layer and each recording item, Corresponding local features of the visual encoder The verification effectively ensures the consistency between text content and image sources, significantly reduces recognition errors or recording contradictions caused by single-modal information deviations, and ultimately achieves high-precision, automated conversion of archival content from images to text and from unstructured to structured, meeting the core requirements of archival digitization for content integrity and recording standardization.

[0095] In this embodiment, the visual encoder is configured and features are extracted as follows: the visual encoder is set to an 8-layer Transformer encoder structure with 5 attention heads, the hidden layer dimension of the Feed-Forward network is 1280, and the activation function is GELU; a 256×256 pixel sliding window (step size 128 pixels) is used to extract features from the original scanned image (300 DPI resolution) of A4-sized paper documents, sequentially acquiring global page layout features (such as the position and layout of the title area, body text area, and signature area) and local text area features (such as the outline of the "Issuing Authority" text block and surrounding blank area information), generating a visual feature sequence containing 128 feature vectors. , where each feature vector The dimension is set to 384, and the vector is embedded with the texture grayscale value of the corresponding image region, the text edge gradient value, and the coordinate normalization value relative to the top left corner of the page.

[0096] Text decoder configuration and task differentiation: The text decoder is set to an 8-layer Transformer decoder structure, with the default archival cataloging structure prompt as "[Cataloging items: Issuing authority, date of issuance, security classification, document number; Transcription requirements: retain paragraph separators and punctuation marks]"; the decoder input is marked using intent type masks, where "[transcription mask]" corresponds to the full-text transcription task and "[cataloging mask]" corresponds to the structured cataloging task; the decoder is based on visual feature sequences. As indicated above, autoregression generates complete transcribed text. At the same time, it generates structured cataloging information. .

[0097] Cross-modal fusion verification: The decoder generates the bibliographic entry "(Date of writing, May 10, 2024)". At that time, the local features of the corresponding "date of writing" text block in the visual encoder are automatically invoked. (Should (This is a 384-dimensional vector, corresponding to the visual features of the region where "May 10, 2024" is located in the image); calculation is performed through a cross-attention layer. The similarity score between the text feature and "May 10, 2024" was 0.92 (the preset similarity threshold was 0.7), indicating that the two pieces of information are consistent and the bibliographic entry is valid. Similarly, the visual verification of other bibliographic entries was performed, and the final output included... and A basic digital object.

[0098] More specifically, when generating structured bibliographic information, the content recognition and bibliographic agent introduces a field confidence evaluation mechanism for any bibliographic item. Its value confidence level Calculated by the following formula:

[0099] ;

[0100] in, For the bibliographic item The relevant visual context feature vectors have the same dimensions as the visual encoder output, ranging from 256 to 512, and are composed of text region features and surrounding layout features. This is the visual feature weight matrix, with dimensions 1×D, where D is... Dimension These are bias terms, with a dimension of 1×1, and are all trainable parameters optimized during model training; Candidate text fragments (i.e., bibliographic entries) (text content) for In full-text transcription The context embedding representation in the language model has a dimension of 768 (generated based on a pre-trained language model). Let cosine similarity function be used to calculate... and The semantic relevance is [0,1], and its value range is [0,1]. The balancing coefficient, with a value range of [0.3, 0.7], is used to adjust the weights of visual and textual evidence. It is dynamically optimized based on the transcription accuracy during the training phase.

[0101] The field confidence assessment mechanism constructs a two-dimensional quantitative model of visual and textual evidence to assess the confidence of bibliographical items. Accurate reliability assessment: on the one hand, relying on the visual contextual feature vector associated with the bibliographic item. (Integrating text area and surrounding layout information), through a trainable weight matrix With bias After processing, the matching degree index is transformed into a visual level using the Softmax function; on the other hand, the context embeddings generated based on the pre-trained language model are... Using the cosine similarity function Calculate candidate text The semantic relevance to the entire text context is used to obtain text-level credibility; at the same time, a balance coefficient is used... (Dynamic optimization within the [0.3, 0.7] interval) Adjusting the weights of the two types of evidence to ultimately determine the confidence level. Quantized to a value in the [0,1] interval; this mechanism avoids the bias of single-modal (visual only or text only) evaluation and provides a clear "reliability benchmark" for subsequent quality verification; low Items can be prioritized to trigger the repair process, high The items pass the verification directly, significantly improving the accuracy and quality control efficiency of structured bibliographic information.

[0102] In this embodiment, the "document number" entry in the paper document archive is used ( Document number, Taking the document X (Document No. 15 of 2024) as an example, the specific implementation of the field confidence assessment mechanism is as follows:

[0103] Parameter configuration: Balance coefficient =0.6 (obtained during the training phase based on the accuracy of document and archive recording). Adopt a 384-dimensional vector (consistent with the output dimension of the visual encoder, formed by splicing the text area features of "X Government Document

[2024] No. 15" and the layout features of the "Document Number:" label on the left), and the visual feature weight matrix is 1×384-dimensional (example of parameters after training: [0.03, 0.01,..., 0.04]), and the bias term = 0.08.

[0104] Calculation of text evidence: is "X Government Document

[2024] No. 15", is the context embedding representation (768-dimensional vector, generated based on the BERT-base pre-trained model) of the fragment "Document Number: X Government Document

[2024] No. 15, Date of Issue: May 10, 2024" in the full-text transcribed text . Calculate the cosine similarity between the two, and the result is 0.93.

[0105] Calculation of confidence:

[0106] First calculate , add and the result is 0.90. The output of the Softmax function is 0.89 (approximate this value in a single-class scenario); then substitute it into the formula . Finally, the confidence of this "Document Number" bibliographic item , which is higher than the preset low-confidence threshold of 0.7, is determined to be a reliable bibliographic item and there is no need to trigger the repair process.

[0107] More specifically, the quality verification and repair agent sets up three levels of quality verification checkpoints, and the scores of each checkpoint are mapped to the interval [0, 1] through normalization. The specific process is as follows:<00,00375>

[0108] Image quality checkpoint: Based on 4 indicators of clarity, contrast, noise density, and edge integrity, calculate the comprehensive readability score through weighted summation. The formula is:

[0109] ;

[0110] where , the value range of each weight is [0.2, 0.3], is the image clarity (calculated based on the Laplacian variance), is the contrast (calculated based on the gray histogram distribution), is the noise density, is the edge integrity; when ( [[ID=�8]]is the default threshold), it is determined that the image quality is unqualified;

[0111] Text quality control: for transcribed text The system employs a dual validation process: first, format validation based on an archive format rule base (e.g., date format, numbering rules, punctuation standards); and second, semantic fluency validation based on a language model (calculating the normalized result of text perplexity). The resulting text credibility score is then used. for:

[0112] ;

[0113] in The format compliance rate is represented by a value of [0,1]. This is the perplexity normalization value, taking values ​​[0,1].

[0114] Bibliographic quality control: combining field confidence levels Verify structured bibliographic information using the bibliographic rule base. Completeness and logical consistency, cataloging completeness score for:

[0115] ;

[0116] in To preset the total number of bibliographic entries, This is a logical verification function (it returns 1 if the rule is met, otherwise it returns 0, such as "date of writing" must not be later than "date of archiving").

[0117] The quality verification and repair intelligent agent establishes a three-tiered quality verification system. Through a layered verification logic of image foundation assurance, text accuracy control, and cataloging standardization verification, it achieves comprehensive and quantitative quality control of digitized archival results. The image quality checkpoint is based on four core visual indicators, including sharpness and contrast. Through weighted summation (with balanced and controllable weights), image readability is converted into a normalized score. This ensures the quality of visual input for subsequent identification and recording; text quality involves fusion of format rule verification and semantic fluency assessment (balancing compliance and content rationality with a weighted average of 0.4:0.6), through... Quantifying Transcription Text To ensure reliability and avoid formatting errors or semantic contradictions; the quality of bibliographic entries should be considered in conjunction with the confidence levels of the fields obtained earlier. Logical verification function ,by Evaluation of bibliographic information The integrity and logical consistency of the data are ensured, eliminating any missing or contradictory items. Scores at all three levels are normalized to the [0,1] interval, forming a unified quality standard. This not only solves the problem of the crudeness of traditional single-dimensional verification but also provides a clear triggering basis for subsequent repair processes, effectively guaranteeing the high quality of the output digital archive objects.

[0118] In this embodiment, taking a scanned image of an A4-sized paper document (300 DPI resolution, content being a policy notice from a municipal government agency) as an example, the three-level quality verification is implemented as follows:

[0119] Image quality control verification: setting weights , , , (satisfy And each weight is in the range [0.2, 0.3]); the variance is calculated using Laplace. Based on the gray-level histogram distribution Noise density detection Edge integrity analysis yielded Substitute into the formula:

[0120] ;

[0121] because The image quality was deemed acceptable.

[0122] Text quality check: Based on archive format rule base check Only one period was found to be missing. The text perplexity is calculated using a pre-trained language model (BERT-based) and then normalized. Substitute into the formula:

[0123] ;

[0124] The text quality is deemed acceptable.

[0125] Bibliographic quality check: Preset total number of bibliographic items (Issuing authority, date of issuance, security classification, document number), corresponding The values ​​were 0.92, 0.906, 0.88, and 0.91 respectively. Verification (e.g., "Document date May 10, 2024 < Archive date May 15, 2024", logical compliance), 4 items. The result was 1 for all cases;

[0126] Substitute into the formula:

[0127] ;

[0128] The cataloging quality was deemed satisfactory, and the archive ultimately passed the three-level quality check, eliminating the need to initiate a repair process.

[0129] More specifically, the quality verification and repair agent initiates an intelligent repair loop for items that fail verification. This loop aims to achieve quality compliance with the minimum number of iterations. The specific process is as follows:

[0130] The repair trigger conditions are as follows:

[0131] when or ( )or ( When this occurs, the repair process will be initiated;

[0132] Targeted repair strategy: For low confidence levels, , For missing or incomplete bibliographic entries, a reinforcement prompt request is sent to the content recognition and bibliographic agent. The prompt information includes the bibliographic entry name, context fragments, and visual region coordinates. Based on the prompt, the agent re-extracts features and generates... Recalculate after update ;

[0133] for and On the page, a parameter optimization request is sent to the image preprocessing agent to adjust the denoising intensity (dynamically adjust the filter kernel size based on noise density) and the skew correction accuracy (increase the angle detection step size to 0.1°). After re-outputting the standard page image, the content recognition and recording agent is triggered for secondary processing.

[0134] For items that still fail to meet the standard after two iterations, or Mark the item as a manual review item and generate a review report (including the problem type, suspected correct result, and relevant visual screenshots).

[0135] Experience feedback mechanism: The input data, processing parameters, output quality score and repair results of each iteration of the repair loop are stored in the experience pool of the central coordinator, which is used to optimize the model parameters and decision-making strategies of each agent.

[0136] The intelligent repair loop of the quality verification and repair agent, with precise triggering, targeted repair, iterative control, and experience accumulation as its core logic, achieves the transformation from inefficient manual repair to efficient intelligent repair: through clearly quantified triggering conditions ( , , To avoid blindly initiating the repair process, targeted strategies were designed for two scenarios: low confidence / missing bibliographical items and poor image and text quality. The former uses enhanced prompts (including bibliographical items, context, and visual coordinates) to drive the content recognition agent to reconstruct features, while the latter improves the basic image quality by optimizing preprocessing parameters (denoising kernel, correction step size), ensuring targeted repair. Two iterations were set as the upper limit (if the standard is not met, manual intervention is required) to balance the repair effect and resource cost. At the same time, an experience feedback mechanism was used to store the iterative data (input, parameters, scores, and results) into the experience pool, providing data support for the optimization of agent model parameters and decision-making strategies, forming a closed loop of repair, learning, and optimization, which not only ensures the digital archive quality compliance rate but also continuously improves the system's autonomous repair capability.

[0137] In this embodiment, taking the personnel files of a certain public institution in 2024 (A4 paper, scanned image resolution 300 DPI) as an example, the intelligent repair loop is implemented as follows:

[0138] Repair triggered: Level 3 verification in progress. (≥) ,qualified), (≥0.85, qualified), (<0.9, triggering repair), tracing back to the source revealed the "archive number" entry. (< =0.7), and valuei = "RS2024-00" (suspected to be missing the last digit).

[0139] Targeted Repair: The quality verification and repair agent sends an enhancement prompt request to the content recognition and cataloging agent. The prompt information includes "Catalogue item name: file number; Context fragment: file number: RS2024-00, name: XX; Visual region coordinates (image coordinate system x:180, y:250, width:200, height:40)". Based on the prompt, the content recognition agent re-extracts the visual features of the region (focusing on capturing the blurred pixels at the end of the number), generates valuei'="RS2024-008", and recalculates. .

[0140] Iterative verification: Recalculation:

[0141] ;

[0142] ≥0.9, meets the standard, repair complete.

[0143] Feedback from experience: The input data for this repair (original visual features of the "file number", contextual text), processing parameters (enhanced cue context length of 20 characters, visual coordinate precision of 1 pixel), and output quality score ( From 0.62 to 0.91, The value (from 0.87 to 0.912) and the repair results are stored in the central coordinator's experience pool for subsequent optimization of the feature extraction weights of "numbered bibliographic items".

[0144] More specifically, the central coordinator maintains the global quality entropy. The metric, used to dynamically evaluate and optimize the processing flow of single and batch files, is defined as the average quality uncertainty of all file objects in the current processing batch, and the formula is:

[0145] ;

[0146] in: This represents the number of files in the batch, ranging from 10 to 1000 (which can be dynamically adjusted based on hardware resources). For the first The file is in The score on the quality dimension (i.e.) , , ), has been normalized to [0,1]; The weights for each quality dimension satisfy... The default value is , , It can be adjusted according to the type of archive (such as engineering drawing archives). );

[0147] Optimization logic: The central coordinator aims to minimize... To achieve the goal, dynamically adjust the batch processing order (prioritize processing). Lower batch size), agent invocation priority (when quality issues are concentrated, increase the resource allocation ratio of quality verification and repair agents to 40%), and parallel processing granularity (the number of files in a single batch varies). The quantity increases and decreases, with a minimum of 10 samples per batch.

[0148] Global mass entropy By converting the multi-dimensional (image, text, and catalog) quality scores of all archives in a batch into a quantified "quality uncertainty" indicator, the problem of "fragmented quality assessment and experience-based scheduling" in traditional batch processing is solved: Based on the principle of information entropy, it integrates the weights of each quality dimension (which can be dynamically adjusted according to the type of archive) with the normalized quality score of a single archive, and intuitively reflects the stability of the overall quality of the batch archives. The lower the value, the more consistent the quality of the archives within the batch and the less uncertain the quality; conversely, the higher the value, the more quality fluctuations there are.

[0149] And surrounding the concept of "minimization" The optimization logic allows the central coordinator to accurately allocate system resources: prioritizing processing... Lower batch sizes (more consistent quality) reduce repair iterations and improve overall processing efficiency; [This is for...] For high-volume batches (with large quality fluctuations), by increasing the proportion of resources allocated to quality verification and repair agents and reducing the granularity of parallel processing, we can avoid the accumulation of quality problems due to insufficient resources, and ultimately achieve a dynamic balance between system processing efficiency and output quality, ensuring the stability and reliability of batch file processing.

[0150] In this embodiment, taking the batch processing of engineering drawing archives (200 copies, 600 DPI resolution) of a design institute as an example, the global quality entropy is... The calculation and optimization are implemented as follows:

[0151] Initial batch configuration: The central coordinator initially sets the batch size. Because engineering drawings have high requirements for image quality, the weight of the quality dimension is adjusted to... , , (satisfy ).

[0152] initial Calculate: the normalized quality score of 5 randomly selected files from this batch ( The entropy values ​​of the following files are: File 1 (0.78, 0.86, 0.89), File 2 (0.82, 0.83, 0.91), File 3 (0.75, 0.81, 0.87), File 4 (0.80, 0.85, 0.90), and File 5 (0.76, 0.82, 0.88). Calculate the entropy value of each file using the formula (taking the natural logarithm as an example). For instance, the entropy value of File 1... Batch average entropy (After calculating all 50 samples) Higher than the historical average It was determined to be "high quality uncertainty".

[0153] Optimized execution: The central coordinator initiated optimizations: the single batch processing granularity was reduced from 50 to 30 (to reduce the impact of quality fluctuations within a single batch); the CPU resource allocation ratio of the quality verification and repair agent was increased from 25% to 40%.

[0154] Optimized result: Reprocessing new batches ( ), calculated (Quality uncertainty is significantly reduced), and the processing latency per batch is reduced from 4.2 seconds / batch to 3.0 seconds / batch, which reduces quality fluctuations while ensuring processing efficiency, in line with the principle of "minimization". The optimization goal is to achieve this.

[0155] More specifically, each agent in the AI ​​agent collaborative network adopts a centralized training and distributed execution paradigm, as implemented as follows:

[0156] Training phase: Construct an annotated multimodal dataset (containing original images, standard transcribed text, structured bibliographical labels, and quality scores from over 100,000 archives), employing an end-to-end joint training mode, and optimizing the objective function as follows:

[0157] ;

[0158] in, For text recognition, cross-entropy loss, Loss for matching bibliographic information (based on field confidence) calculate), For quality score regression loss, , The loss balance coefficient is used; the training optimizer employs AdamW with a learning rate of [missing value]. The training sessions consist of 20-30 rounds.

[0159] Execution phase: Each agent is deployed on an independent containerized service node and responds to the central coordinator's scheduling instructions via RESTful API. The central coordinator dynamically allocates tasks based on the task queue length and agent load rate (CPU / GPU utilization), supporting parallel processing of multiple batches of files, with a single batch processing latency of ≤3 seconds / file (based on GPU server configuration).

[0160] The AI ​​agent collaborative network adopts a centralized training and distributed execution paradigm, achieving the dual goals of collaborative optimization of model performance and efficient utilization of system resources: centralized training constructs a multimodal labeled dataset to jointly optimize the objective function end-to-end. Deeply integrate text recognition, record matching, and quality regression tasks, leveraging... , The loss balance coefficient avoids overall processing deviation caused by over-optimization of a single task, ensuring that each agent achieves synergistic improvement in recognition accuracy, recording accuracy, and quality assessment reliability. Distributed execution enables each agent to run independently and expand flexibly through containerized deployment. Combined with the dynamic scheduling of the central coordinator based on task queue length and load rate, it can achieve parallel processing of multiple batches of files and control the processing latency of a single batch to ≤3 seconds / file. This effectively solves the problems of isolated training and execution congestion in traditional centralized architectures, balancing model performance and system processing efficiency.

[0161] In this embodiment, the training phase involves constructing a multimodal labeled dataset of 120,000 documents (including 40,000 paper documents, 50,000 engineering drawings, and 30,000 audio and video recordings; each dataset contains the original image and standard full-text transcription). Structured cataloging tags Mark quality score ); an end-to-end joint training mode is adopted to optimize the objective function. Cross-entropy loss for CTC text recognition (calculating predicted transcribed text versus standard text) Character-level differences). Mean squared error loss (calculating the confidence level of the predicted field) With annotation (deviation) The mean absolute error loss is calculated as the difference between the predicted quality score and the labeled quality score. The AdamW optimizer is used for training, with an initial learning rate set to... The training process involved a 10% decay every 6 rounds, for a total of 28 rounds (between 20 and 30 rounds). After training, the text recognition accuracy reached 98.5%, the bibliographic information matching accuracy reached 97.8%, and the quality score prediction error was ≤0.03.

[0162] In this embodiment, the execution phase is implemented as follows: the image preprocessing, content recognition and cataloging, and quality verification and repair agents are deployed on three Docker containers (running on two servers equipped with NVIDIA A10 GPUs, each with 24GB of video memory). Each container receives scheduling instructions from the central coordinator via a RESTful API. When the system receives 150 engineering drawing files for processing, the central coordinator detects that the image preprocessing agent's task queue length is 20 and the load rate is 65%, the content recognition and cataloging agent's queue length is 15 and the load rate is 70%, and the quality verification and repair agent's queue length is 10 and the load rate is 60%. It then dynamically allocates them into three parallel batches (50 files per batch) and schedules them to the corresponding containers according to the pipeline of preprocessing, recognition and cataloging, and quality verification. Finally, the average processing latency of the three batches is 2.6 seconds / file, 2.9 seconds / file, and 2.7 seconds / file, respectively, all ≤3 seconds / file. The average GPU resource utilization is stable at 72%-78%, achieving efficient parallel processing.

[0163] More specifically, before deploying the intelligent AI-driven end-to-end archival digitization system, the core content recognition and cataloging intelligent agent undergoes domain-adaptive pre-training. The specific process is as follows:

[0164] Pre-trained dataset: Contains 500,000+ unlabeled archival image-text pairs and 100,000+ labeled archival data, covering multiple types such as paper documents, engineering drawings, and audio-visual archives;

[0165] Pre-training task:

[0166] Archival document masking language modeling: Randomly mask 15% of the tokens (including text and formatting tags) in the text, train the model to predict the mask content based on visual features and context, and use cross-entropy as the loss function;

[0167] Page element relationship prediction: Label the relevant blocks of titles, paragraphs, tables, and seals in the document, train the model to predict the relative positional relationship (such as "title above paragraph") and semantic relevance between different blocks, and use the loss function as binary cross-entropy.

[0168] Cross-modal bibliographic alignment: Construct a bibliographic item-image region mapping dataset, train the model to learn the mapping relationship from a specific visual region (such as the title area of ​​an official document) to the corresponding bibliographic item (such as "issuing authority"), and use the contrastive learning loss function;

[0169] Pre-training parameters: fine-tuned based on an open-source multimodal large model (such as BLIP-2), with 10-15 training epochs, a batch size of 32, and a learning rate of [missing information]. The weight decay coefficient is .

[0170] Domain-adaptive pre-training for content recognition and cataloging agents fundamentally addresses the issues of weak generalization and poor specificity of general multimodal models in the archival domain. This is achieved through a large-scale pre-training dataset covering various archival types, including paper documents, engineering drawings, and audio-visual materials (500,000+ unlabeled image-text pairs providing the domain data foundation, and 100,000+ labeled layout data providing fine-grained supervision), combined with three targeted pre-training tasks. Archival document masking language modeling enhances the predictive ability of text context and visual features; layout element relationship prediction improves the model's understanding accuracy of archival-specific layouts (such as title areas and signature areas); and cross-modal cataloging alignment establishes a unique mapping logic of "visual region - cataloging item," ultimately enabling the model to pre-master the semantic rules, layout features, and cross-modal association rules of the archival domain. Simultaneously, reasonable parameter settings fine-tuned based on BLIP-2 (10-15 training rounds, 32 batch sizes, 2×10) are employed. -5 The learning rate ensures the stability of model training while avoiding overfitting, laying a high-quality foundation for fine-tuning in subsequent tasks and significantly improving the recognition accuracy, cataloging standardization, and cross-modal consistency of the agent in actual archive processing.

[0171] More specifically, the intelligent AI-driven end-to-end archival digitization system has scenario-based adaptation capabilities for different types of archives, and the specific adjustment strategies are as follows:

[0172] Paper document archives: The content recognition and recording intelligent agent enhances the detection of special visual elements related to official seals, signatures, and red header characters. It uses the YOLOv8 model to extract special element features separately, and then merges them with text features to generate recording information. The recording items have added exclusive fields such as "official seal name", "signer", and "red header mark".

[0173] Engineering drawing archives: The image preprocessing agent integrates a vectorization process (using the Potrace algorithm) to convert raster images into SVG vector format; the content recognition agent is adapted to recognize chart symbols, adding relevant bibliographic entries for drawing numbers, scales, and dimensions; the visual encoder is optimized to a feature extraction structure for lines and symbols.

[0174] Photo and audio / video archives: The content recognition and cataloging agent has been expanded into a multimedia content analysis model. Photo archives support scene description and intelligent recognition (based on the FaceNet model), while audio / video archives support speech-to-text conversion and extraction of key event timelines. The digital resource management module has added a multimedia format adaptation interface, supporting video frame indexing and audio segment retrieval.

[0175] The system's scenario-based adaptability breaks through the limitations of traditional archival digitization systems that "adapt to all types of documents with a single process." By customizing the entire process according to the core characteristics of different archives, it achieves "one exclusive processing solution for each type of archive." For paper documents, it strengthens the detection and exclusive recording of special visual elements such as official seals and signatures, solving the problem of missing key voucher information. For engineering drawings, it ensures the accuracy of technical drawings and the accuracy of technical parameter extraction through vectorization and line symbol optimization. For multimedia archives such as photos and audio-visual materials, it expands the multimedia content analysis and format adaptation interfaces, breaking through the technical bottleneck of non-text archive digitization. This adaptation extends from preprocessing and identification to resource management, ensuring that the core value information of each type of archive is not lost and that the digitization results meet industry usage standards. It significantly improves the applicability of the system in various industry scenarios such as government agencies, design institutes, and media, avoiding "quality deviations caused by using a general process to process special archives."

[0176] In this embodiment, taking a news interview audio and video archive from a TV station in 2024 (a total of 500 files, in formats including MP4 and WAV, with a video resolution of 1080P and an audio sampling rate of 44.1kHz) as an example, the scenario-based adaptation is implemented as follows:

[0177] Content Recognition and Recording Adaptation: The system automatically identifies the file type as "audio-video file," and the content recognition and recording agent switches to a multimedia content analysis model; for video files, the FaceNet model is used to detect the interviewee's face (recognition accuracy of 95.2%), generating an "Interviewee" record item; at the same time, keyframes (1 frame every 30 seconds) are extracted for scene description (such as "meeting room interview scene" or "outdoor on-site reporting scene"); for audio files, the speech-to-text model (based on Whisper-large-v3) is used to convert the audio files into text with an accuracy of 94.8%, and a timeline is extracted using a key event detection algorithm (based on voice emotion fluctuation and keyword extraction) (such as "00:15:30 mentions 'policy implementation time'" or "00:28:10 explains 'project progress'"), adding dedicated record items for "Interviewee," "scene description," and "key event timeline."

[0178] Digital Resource Management Adaptation: The digital resource management module calls the newly added multimedia format adaptation interface, supporting direct import of MP4 / WAV formats into the database, and establishing video frame indexes (associated with key frame scene descriptions) and audio segment indexes (associated with speech-to-text). When users search, they can accurately match the corresponding video by "interviewee = name", or locate relevant segments in the audio by entering the keyword "policy implementation time" (response time ≤ 1 second), meeting the business needs of TV stations for rapid person retrieval and key content location in audio and video archives.

[0179] More specifically, the automatic classification and indexing function integrated into the digital resource management module is driven by an independent intelligent agent constructing an archival knowledge graph, and is implemented as follows:

[0180] Knowledge Extraction: Intelligent Agents Analyze Structured Descriptive Information With full text content Extract archival entities (such as issuing authorities and responsible persons), event entities (such as meetings and approvals), and their relationships (such as "issuing authorities - issuance - archives" and "archives - related - events").

[0181] Knowledge graph construction: Based on the Neo4j graph database, it dynamically links to the preset archival domain ontology (containing 4 core entities: "archives, agencies, personnel, and events"), and supports automatic supplementation of entity attributes and relationship expansion;

[0182] Semantic retrieval implementation: Based on knowledge graphs, a semantic index is built, supporting three types of retrieval methods:

[0183] Metadata retrieval: based on exact matching of bibliographic entries;

[0184] Full-text search: based on text similarity matching;

[0185] Related search: Based on entity relationship path matching (e.g., "search for archives of documents issued by a certain government agency in 2023 that involve a certain meeting");

[0186] Knowledge update mechanism: Newly added entities and relationships are automatically integrated into the knowledge graph. Redundancy is cleaned up and relationships are strengthened regularly (weekly by default) to ensure retrieval accuracy.

[0187] The automatic classification and indexing function of the digital resource management module, relying on the archival knowledge graph to build an intelligent agent, upgrades from "archival storage" to "knowledge association": it accurately captures entities (organizations, personnel), events, and relationships in the archives through knowledge extraction, solving the problem of traditional classification and indexing "relying only on metadata and lacking deep association"; based on the Neo4j graph database and domain ontology, the graph construction organizes fragmented archival information into a structured knowledge network, supports automatic supplementation of entity attributes and relationship expansion, and improves the systematic nature of archival management; among the three semantic retrieval methods, the association retrieval breaks through the limitations of precise matching of metadata and full-text similarity matching, and can realize complex cross-entity query needs (such as the archival association location of "specific organization and specific event"); and the regular knowledge update mechanism (weekly redundancy cleanup and relationship strengthening) ensures the timeliness and accuracy of the knowledge graph and avoids retrieval bias caused by information redundancy.

[0188] In this embodiment, taking the 2023 policy archives of a municipal government agency (a total of 1000 documents, including scanned copies of paper documents and electronic files) as an example, the automatic classification and indexing function is implemented as follows:

[0189] Knowledge Extraction: Intelligent Agent Analyzes Structured Descriptive Information of Each Archive by Constructing an Archival Knowledge Graph (e.g., "Issuing Authority: A Municipal Government Agency, Responsible Person: XY, Date of Issuance: June 15, 2023") and the full text. (Including "This subsidy plan is formulated in accordance with the requirements of the 2023 Livelihood Policy Deployment Meeting"), the entities to be extracted are: the archive entity "2023 Livelihood Subsidy Policy Archive of a certain city", the agency entity "agency of a certain city", the personnel entity "XY", and the event entity "2023 Livelihood Policy Deployment Meeting"; the extraction relationships are: "agency of a certain city - issued document - 2023 Livelihood Subsidy Policy Archive of a certain city", "2023 Livelihood Subsidy Policy Archive of a certain city - involved - 2023 Livelihood Policy Deployment Meeting", "XY - responsible for - 2023 Livelihood Subsidy Policy Archive of a certain city".

[0190] Knowledge graph construction: Based on the Neo4j graph database, the extracted entities and relationships are dynamically linked to the preset "ontology of archives of government agencies" (including 4 core entities of "archives, agencies, personnel, and events" and preset relationships such as "issued documents, involved, and responsible"), and entity attributes are automatically supplemented.

[0191] Semantic retrieval implementation:

[0192] Metadata retrieval: When a user inputs "issuing authority = municipal government agency", the system returns 320 relevant files based on precise matching of bibliographic data.

[0193] Full-text search: When a user inputs "people's livelihood subsidies", the system returns 85 files containing the keyword based on text similarity matching (cosine similarity threshold 0.8);

[0194] Related Search: When a user inputs "archives of documents issued by a municipal government agency in 2023 that involve the deployment of policies related to people's livelihood", the system matches based on the entity relationship path "municipal government agency - issued documents - archives - involved - 2023 policy deployment meeting for people's livelihood", accurately returning 12 target archives with a response time of ≤2 seconds.

[0195] Knowledge Update: Ten new supplementary policy documents for 2023 were added the following month. The intelligent agent automatically extracted their entities and relationships and integrated them into the graph. Every Sunday, the system performs redundancy cleanup (deleting two duplicate "meeting-involved" relationships) and relationship strengthening (updating the time range of the "XY-responsible" relationship to "June-December 2023") to ensure retrieval accuracy.

[0196] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A smart AI-driven end-to-end archival digitization system, characterized in that, It includes a physical digitization module, an AI intelligent agent collaborative network, and a digital resource management module that work in sequence to achieve automated processing from physical archives to a structured digital knowledge base; The physical digitization module is configured to perform high-fidelity scanning or photography of physical archives to generate an original digital image set. The AI ​​agent collaborative network is a multi-agent system managed and scheduled by a central coordinator. It receives the original digital image set and sequentially calls the image preprocessing agent, content recognition and recording agent, and quality verification and repair agent for pipelined processing; wherein: The image preprocessing intelligent agent performs automatic skew correction, noise reduction, and page segmentation operations on the original digital image, and outputs a standard page image. The content recognition and recording intelligent agent, based on a fusion model of large language model and computer vision, performs full-text OCR recognition, layout analysis, key field extraction and structured recording on standard page images to generate a primary digital object containing text content and metadata. The quality verification and repair agent performs multi-dimensional verification of the image quality, text recognition accuracy, and bibliographical integrity of the primary digital object, and initiates automatic or semi-automatic repair processes for recognition errors or missing items, outputting high-quality digital archive objects. The digital resource management module receives the high-quality digital archive objects, automatically classifies, indexes, and stores them, and establishes a multi-dimensional retrieval index, ultimately forming a manageable and exploitable digital archive resource library. The content recognition and transcription agent adopts a Transformer-based vision-language multimodal fusion architecture, and the specific implementation process is as follows: The visual encoder employs a 6-10 layer Transformer encoder structure with 4-6 attention heads. The hidden layers of the Feed-Forward network have a dimension of 1024-1536, and the GELU activation function is used. A sliding window is used to extract global layout features and local text region features from the page image, generating a visual feature sequence. ,in It consists of 256-512 dimensional feature vectors, which fuse image texture, text outline, and region location information; The text decoder employs a 6-10 layer Transformer decoder structure, based on the aforementioned visual feature sequence. With preset archival description framework prompts, the system distinguishes between full-text transcription and structured description tasks using intent type masks, and generates two parts of output autoregressively: Complete transcribed text It covers all text content and formatting marks on the page; Structured cataloging information It is a set of key-value pairs. ,in For the preset bibliographic item name, The corresponding content extracted from an image or text; Cross-modal fusion mechanism: Achieves bidirectional interaction between visual and textual features through a cross-attention layer. The decoder generates a bibliographic entry for each entry. All of these will call the local features corresponding to the visual encoder. Verification was performed to ensure consistency of information across modalities; When generating structured bibliographic information, the content recognition and bibliographic agent introduces a field confidence evaluation mechanism for any bibliographic item. Its value confidence level Calculated by the following formula: ; in, For the bibliographic item The relevant visual context feature vectors have the same dimensions as the visual encoder output, ranging from 256 to 512, and are composed of text region features and surrounding layout features. This is the visual feature weight matrix, with dimensions 1×D, where D is... Dimension These are bias terms, with a dimension of 1×1, and are all trainable parameters optimized during model training; For candidate text fragments, for In full-text transcription The context embedding representation in the image has a dimension of 768. Let cosine similarity function be used to calculate... and The semantic relevance is [0,1], and its value range is [0,1]. The balancing coefficient, with a value range of [0.3, 0.7], is used to adjust the weights of visual and textual evidence. It is dynamically optimized based on the transcription accuracy during the training phase.

2. The intelligent AI-driven end-to-end archival digitization system according to claim 1, characterized in that, The quality verification and repair agent establishes three quality verification checkpoints. The scores of each checkpoint are normalized and mapped to the [0,1] interval. The specific process is as follows: Image quality control: Based on four indicators—sharpness, contrast, noise density, and edge integrity—a comprehensive readability score is calculated through weighted summation. The formula is: ; in The weights range from [0.2, 0.3]. For image clarity, For contrast, For noise density, For edge integrity; when At that time, the image quality was deemed unacceptable; Text quality control: for transcribed text The system employs a dual validation process: first, format validation based on an archive format rule base; and second, semantic fluency validation based on a language model, resulting in a text credibility score. for: ; in The format compliance rate is represented by a value of [0,1]. This is the perplexity normalization value, taking values ​​[0,1]. Bibliographic quality control: combining field confidence levels Verify structured bibliographic information using the bibliographic rule base. Completeness and logical consistency, cataloging completeness score for: ; in To preset the total number of bibliographic entries, This is a logic verification function.

3. The intelligent AI-driven end-to-end archival digitization system according to claim 2, characterized in that, The quality verification and repair agent initiates an intelligent repair loop for items that fail verification. This loop aims to achieve quality compliance with the minimum number of iterations. The specific process is as follows: The repair trigger conditions are as follows: when or ( )or ( When this occurs, the repair process will be initiated; Targeted repair strategy: For low confidence levels, , For missing or incomplete bibliographic entries, a reinforcement prompt request is sent to the content recognition and bibliographic agent. The prompt information includes the bibliographic entry name, context fragments, and visual region coordinates. Based on the prompt, the agent re-extracts features and generates... Recalculate after update ; for and On the page, a parameter optimization request is sent to the image preprocessing agent to adjust the noise reduction intensity and skew correction accuracy. After re-outputting the standard page image, the content recognition and recording agent is triggered for secondary processing. For items that still fail to meet the standard after two iterations, or Mark these as items for manual review and generate a review report; Experience feedback mechanism: The input data, processing parameters, output quality score and repair results of each iteration of the repair loop are stored in the experience pool of the central coordinator, which is used to optimize the model parameters and decision-making strategies of each agent.

4. The intelligent AI-driven end-to-end archival digitization system according to claim 3, characterized in that, The central coordinator maintains the global quality entropy. The metric, used to dynamically evaluate and optimize the processing flow of single and batch files, is defined as the average quality uncertainty of all file objects in the current processing batch, and the formula is: ; in: This represents the number of files within the batch, ranging from 10 to 1000. For the first The file is in The scores on the quality dimension have been normalized to [0,1]. The weights for each quality dimension satisfy... The default value is , , This can be adjusted according to the type of file; Optimization logic: The central coordinator aims to minimize... With the goal of dynamically adjusting the batch processing order, agent invocation priority, and parallel processing granularity.

5. The intelligent AI-driven end-to-end archival digitization system according to claim 4, characterized in that, Each agent in the AI ​​agent collaborative network adopts a centralized training and distributed execution paradigm, specifically implemented as follows: Training phase: Construct a labeled multimodal dataset, adopt an end-to-end joint training mode, and optimize the objective function as follows: ; in, For text recognition, cross-entropy loss, Due to the loss in matching bibliographic information, For quality score regression loss, , The loss balance coefficient is used; the training optimizer employs AdamW with a learning rate of [missing value]. The training sessions consist of 20-30 rounds. Execution phase: Each agent is deployed on an independent containerized service node and responds to the central coordinator's scheduling instructions via RESTful API; the central coordinator dynamically allocates tasks based on the task queue length and agent load rate, supporting parallel processing of multiple batches of files, with a single batch processing latency of ≤3 seconds / file.

6. The intelligent AI-driven end-to-end archival digitization system according to claim 5, characterized in that, Before deploying the intelligent AI-driven end-to-end archival digitization system, the core content recognition and cataloging intelligent agent undergoes domain-adaptive pre-training. The specific process is as follows: Pre-trained dataset: Contains over 500,000 unlabeled archival image-text pairs and over 100,000 archival data sets with labeled layout elements; Pre-training task: Archival document masking language modeling: Randomly mask 15% of the token in the text, train the model to predict the mask content based on visual features and context, and use cross-entropy as the loss function; Page element relationship prediction: Mark the relevant blocks of titles, paragraphs, tables, and seals in the document, train the model to predict the relative positional relationship and semantic association between different blocks, and use the loss function as binary cross-entropy. Cross-modal cataloging alignment: Construct a cataloging item-image region mapping dataset, train a model to learn the mapping relationship from a specific visual region to the corresponding cataloging item, and use contrastive learning loss as the loss function; Pre-training parameters: Fine-tuned based on an open-source multimodal large model, with 10-15 training epochs, a batch size of 32, and a learning rate of [missing information]. The weight decay coefficient is .

7. The intelligent AI-driven end-to-end archival digitization system according to claim 6, characterized in that, The intelligent AI-driven end-to-end archival digitization system has the ability to adapt to different types of archives in various scenarios. The specific adjustment strategies are as follows: Paper document archives: The intelligent agent for content recognition and recording enhances the detection of special visual elements related to official seals, signatures, and red-headed characters. It uses the YOLOv8 model to extract special element features separately, and then merges them with text features to generate recording information. Engineering drawing archives: The image preprocessing agent integrates a vectorization process to convert raster images into SVG vector format; the content recognition agent is adapted to recognize chart symbols, adding relevant bibliographic entries for drawing numbers, scales, and dimensions; and the visual encoder is optimized to a feature extraction structure for lines and symbols. Photo and audio / video archives: The content recognition and recording intelligent agent has been expanded into a multimedia content analysis model. Photo archives support scene description and intelligent recognition, while audio / video archives support speech-to-text conversion and extraction of key event timelines. The digital resource management module has added a multimedia format adaptation interface, supporting video frame indexing and audio segment retrieval.

8. The intelligent AI-driven end-to-end archival digitization system according to claim 7, characterized in that, The automatic classification and indexing functions integrated into the digital resource management module are driven by an independent intelligent agent constructing an archival knowledge graph, specifically implemented as follows: Knowledge Extraction: Intelligent Agents Analyze Structured Descriptive Information With full text content Extract archival entities, event entities, and their relationships; Knowledge graph construction: Based on the Neo4j graph database, it dynamically links to the preset archive domain ontology, and supports automatic supplementation of entity attributes and relationship expansion; Semantic retrieval implementation: Based on knowledge graphs, a semantic index is built, supporting three types of retrieval methods: Metadata retrieval: based on exact matching of bibliographic entries; Full-text search: based on text similarity matching; Related retrieval: based on entity relationship path matching; Knowledge update mechanism: The entities and relationships of newly added files are automatically integrated into the knowledge graph, and the graph is regularly cleaned up for redundancy and relationships are strengthened to ensure retrieval accuracy.

Citation Information

Patent Citations

  • OCR-based document automatic identification intelligent management system

    CN120564202A

  • Intelligent archive description method and device and storage medium

    CN120653819A