Intelligent file processing system and method based on artificial intelligence

By using an AI-based intelligent archival processing system and method, the problems of insufficient digitization quality and low efficiency of intelligent classification and retrieval in archival management systems have been solved. This has enabled efficient knowledge extraction and information utilization, improved user experience, and supported in-depth semantic analysis and intelligent question answering.

CN121579421APending Publication Date: 2026-02-27QIZHI XINLIAN (NANJING) INFORMATION SOFTWARE DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511762563.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing record management systems suffer from problems such as insufficient quality of digitized records, lack of intelligent classification and retrieval, insufficient knowledge extraction capabilities, limited information utilization value mining, and inadequate user experience. In particular, when faced with diverse types of records, they are difficult to identify, have low retrieval efficiency, and cannot achieve in-depth semantic analysis and intelligent question answering.

Method used

An AI-based intelligent archival processing system and method are adopted, including archival scanning and image acquisition, image preprocessing, OCR recognition and text generation, text preprocessing and feature extraction, document classification, semantic retrieval, entity extraction, relation extraction, knowledge graph construction, intelligent question answering and analysis, summary generation and data visualization, and automated processing using deep learning and large language models.

Benefits of technology

It improved the quality and accuracy of digitized archives, enhanced the efficiency of intelligent classification and semantic retrieval, strengthened the ability of knowledge extraction and structured management, improved the user experience, and achieved full-process automation and resource conservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579421A_ABST
    Figure CN121579421A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent archive processing method based on artificial intelligence, which relates to the technical field of archive digitization and comprises the steps of archive scanning and image acquisition, image preprocessing, OCR (Optical Character Recognition) and text generation, result storage and formatting, text preprocessing and feature extraction, document classification model training and application, semantic retrieval implementation and entity extraction. The invention further provides an intelligent archive processing system based on artificial intelligence, and the intelligent archive processing system comprises an archive digitization module, an intelligent classification module, a knowledge extraction and graph construction module, an intelligent retrieval and question and answer module and an intelligent abstract and analysis module. According to the method, the digital quality and recognition precision of the archive can be remarkably improved, the intelligent classification and semantic retrieval efficiency is improved, the knowledge extraction and structured management capability is enhanced, the operation is simple and convenient, the whole process is automatic, resources are saved, and the method is environment-friendly.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of file digitization, and particularly relates to a file intelligent processing system and method based on artificial intelligence. BACKGROUND

[0002] With the continuous advancement of informatization and digitization, various types of files (such as accounting files, personnel files, document files, and business files) are gradually changing from traditional paper forms to electronic storage and management. The existing file management technology mainly relies on manual scanning, OCR recognition, and keyword indexing based on rules, and has the following problems and deficiencies: 1. Insufficient file digitization quality, traditional digitization relies on manual scanning and OCR recognition, but the paper quality of historical files is uneven, and scanning images have problems such as blur, damage, and tilt, which leads to low OCR recognition accuracy. Especially for handwritten files, tabular files, and multi-lingual file recognition, it is difficult to affect data utilization efficiency. 2. Lack of intelligent classification and retrieval, most existing systems rely on manual setting of classification rules or keyword matching, and lack semantic understanding ability. In the face of diversified file types (contracts, appointment and removal notices, meeting minutes, etc.), the system cannot automatically identify document categories or understand user query intentions, and the retrieval efficiency is low. 3. Insufficient knowledge extraction capability, files contain time, person, event, unit, and other key information, but existing systems mainly rely on full-text storage and keyword indexing, making it difficult to extract structured information and build a knowledge graph, and unable to support deep semantic analysis and intelligent question answering. 4. Limited information utilization value mining, files are not only historical records, but also important basis for institutional decision-making. The existing management mode takes archiving as the goal, lacks automatic summary generation, trend analysis, and semantic question answering functions, and the value of files cannot be fully utilized. 5. Insufficient user experience, system operation is complex, users need to manually compare and search, and the efficiency is low. In cross-department and cross-type file retrieval, information is often missing or redundant. Therefore, our company designs and proposes a file intelligent processing system and method based on artificial intelligence. SUMMARY

[0003] The purpose of the present application is to solve the problems of insufficient file digitization quality, lack of intelligent classification and retrieval, insufficient knowledge extraction capability, limited information utilization value mining, and insufficient user experience in the prior art, and to provide a file intelligent processing system and method based on artificial intelligence.

[0004] In order to achieve the above purpose, the present application adopts the following technical scheme: A file intelligent processing method based on artificial intelligence is designed, which comprises the following steps: Step 1, file scanning and image acquisition: obtaining scanning images of paper files through a high-speed scanner, and batch naming and archiving the scanning images; Step 2, Image Preprocessing: Use image processing algorithms for denoising, skew correction, and contrast enhancement. Step 3, OCR Recognition and Text Generation: Call Tesseract or deep learning OCR models for character recognition and post-process the recognized text. Step 4, Result Storage and Formatting: Save the OCR results as JSON or database records. Step 5, Text Preprocessing and Feature Extraction: Tokenize, part-of-speech tag, and named entity recognition the OCR text, and extract keywords, phrases, and document metadata as feature vectors to input into a deep learning classification model. Step 6, Document Classification Model Training and Application: Train a BERT, RoBERTa, or Transformer model for document type classification, write the classification results to the database, and generate a classification report. Step 7, Semantic Retrieval Implementation: Users input query statements in natural language, and the system uses large language models to analyze the query intent, match the knowledge graph and text vectors, and return relevant archives. Step 8, Entity Extraction: Use rule matching and deep learning to perform named entity recognition on OCR text, extracting "time," "location," "person," "unit," and "event" information. Step 9, Relationship Extraction: Extract entity relationships based on syntactic analysis and dependency analysis, linking cross-page and cross-document relationships. Step 10, Knowledge Graph Construction: Store entities and relationships in a graph database, generating nodes and edges. Step 11, Intelligent Question Answering and Analysis Application: Answer user questions by reasoning with the knowledge graph and large models, returning accurate answers and source archives. Step 12, Summary Generation: Use extractive or generative text summarization algorithms to generate concise summaries of archive content. Step 13, Data Visualization: Generate visual reports of knowledge graphs, summaries, and statistical information.

[0005] Further, in Step 2, perform region segmentation on tables and handwritten text, use OpenCV or deep learning models for layout analysis, and distinguish text regions, table regions, and image regions.

[0006] Further, in Step 3, use deep learning handwriting recognition models for handwritten text, outputting structured text.

[0007] Further, in Step 3, post-process the recognized text, including spelling correction, character replacement, and paragraph reconstruction.

[0008] Further, in step 4, the database record includes the original image path, text content, page number information and recognition confidence, and the low-confidence area is automatically labeled for manual review.

[0009] To solve the above technical problems, the present application further provides an artificial intelligence-based archive intelligent processing system, comprising: An archive digitization module for paper archive digitization and OCR processing; An intelligent classification module based on a deep learning model for automatic classification, batch naming and archiving of scanned images; A knowledge extraction and graph construction module for entity extraction and relationship construction; An intelligent retrieval and question answering module supporting natural language semantic retrieval and complex question answering for user experience and query; An intelligent summary and analysis module for generating summaries and statistical reports.

[0010] The artificial intelligence-based archive intelligent processing system and method of the present application have the beneficial effects of significantly improving archive digitization quality and recognition accuracy, improving intelligent classification and semantic retrieval efficiency, enhancing knowledge extraction and structured management capability, improving intelligent summary and information utilization value, and being easy to operate, fully automated, and resource-saving and environmentally friendly. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation on the present application. In the drawings: Fig. 1 is a flow chart of the artificial intelligence-based archive intelligent processing method of the present application; Fig. 2 is a block diagram of the artificial intelligence-based archive intelligent processing system of the present application; Fig. 3 is a schematic diagram of the knowledge graph construction of the present application. DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application; obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0013] The structural features of the present application will be described in detail below with reference to the drawings in the specification.

[0014] Referring to Figs. 1-3The application discloses an artificial intelligence-based archive intelligent processing method, which comprises archive digitization and OCR processing, intelligent classification and semantic retrieval, knowledge extraction and graph construction, and intelligent abstract and visual analysis. S1, archive scanning and image acquisition: paper archives are scanned by a high-speed scanner to obtain scanned images, and the scanned images are batch-named and archived. The resolution of the high-speed scanner is not less than 300 dpi, the color mode is grayscale or color, and the scanned images are saved in PNG or TIFF format. The scanned images are batch-named and archived to ensure unique identification and facilitate subsequent processing.

[0015] S2, image preprocessing: grayscale processing, image denoising, tilt correction and contrast enhancement are performed using image processing algorithms; wherein Gaussian filtering is used to realize image denoising processing, Hough transform is used to detect the text tilt angle and rotate to realize image tilt correction processing, and linear stretching or CLAHE algorithm is used to realize image contrast enhancement processing. Wherein: 1. Grayscale processing: If the input image is a color image , it is first converted into a grayscale image . The conversion formula is as follows: Wherein, , , represent the red, green and blue channel intensities of the pixel point ; and is a grayscale value, usually in the range of .

[0016] 2. Gaussian filter denoising: In order to reduce the interference of noise on text feature extraction, the application adopts a two-dimensional Gaussian convolution kernel to smooth the image. The Gaussian kernel is defined as: The smoothed image after convolution is: Wherein, is the standard deviation of the Gaussian kernel, used to control the smoothing degree; is the kernel radius; the convolution kernel satisfies the normalization condition .

[0017] 3. Text tilt angle detection and correction: In order to realize automatic tilt correction of the image, the application adopts Hough transform to detect the text line direction and calculate the rotation angle. The specific steps are as follows: 1), first, the smoothed image is subjected to Canny edge detection, and the gradient is calculated: where, is the gradient magnitude, is the gradient direction.

[0018] 2), Hough transform is performed on the edge points to calculate the accumulator space: where, represents the perpendicular distance of the straight line to the origin, is the included angle between the normal line of the straight line and the horizontal direction. Accumulation is performed on all edge points to obtain the accumulator .

[0019] 3), the main direction angle of the text is determined by detecting the peak position of . Then the tilt correction rotation angle is: When , the image is rotated clockwise ; when , it is rotated counterclockwise .

[0020] 4), the rotation transformation adopts a two-dimensional affine transformation matrix: For the output pixel point , it is inversely mapped to the original image coordinates: where, is the image center coordinate. The gray value is obtained by bilinear interpolation calculation.

[0021] 4, contrast enhancement: In order to enhance the readability of the enhanced file text, the application adopts two optional schemes: 1), linear contrast stretching: where, is the input gray value, is the output gray value, , respectively, the minimum and maximum values of the image gray value, is the gray level.

[0022] This formula realizes the linear mapping of the global gray distribution, thereby improving the contrast.

[0023] 2), adaptive histogram equalization (CLAHE): (1), divide the image into several sub-blocks, each block has a size of ; (2), calculate the gray histogram of each block , and set the clipping threshold : The total amount of excess parts is: (3), evenly distribute the overflow part to each gray level: (4), calculate the cumulative distribution function: (5), generate enhanced gray value: wherein, is the local window size (recommended 32x32), is the clipping threshold (usually 2-4 times the average histogram value), is the cumulative probability distribution.

[0024] By bilinear interpolation of the results of each sub-block, the edge discontinuity problem can be avoided.

[0025] 5, step summary: The image preprocessing flow of the present application can be expressed as the following order: 1. input file image ; 2. get by grayscale processing 3. get by Gaussian convolution smoothing denoising 4. extract edge using Canny operator and get text tilt angle by Hough transform ; 5. realize image tilt correction according to rotation matrix ; 6. linear or adaptive contrast enhancement is carried out on the image Output the optimized image for OCR and semantic analysis.

[0026] The table and handwritten text are segmented, and OpenCV or deep learning model is used for layout analysis to distinguish text area, table area and image area.

[0027] The present application adopts a region segmentation algorithm based on layout analysis. The method comprises the following steps: 1. image preprocessing: Converting input file image to grayscale image and obtaining a binary image by adaptive thresholding Let the threshold function be where is the local adaptive threshold.

[0028] 2. Morphological feature extraction: Enhance the connectivity of the text and weaken the noise points by dilation and erosion operations. Let the structure element be Then the morphological operation is defined as: 3. Candidate region detection: Use Connected Component Analysis (CCA) to extract candidate regions. For each connected component , calculate its bounding rectangle and geometric features such as aspect ratio , pixel density .

[0029] 4. Region classification: According to the geometric and texture features, use a classification model based on rules or convolutional neural networks (such as LayoutLM or U-Net) to divide the candidate regions into: Text region (continuous line of text features, high density, regular rectangle); Table region (exists regular straight line structure, can be detected by Hough transform); Handwritten region (pen outline irregular, line width change large).

[0030] 5. Result post-processing: Merge adjacent regions of the same type to generate the final layout structure diagram .

[0031] S3. OCR recognition and text generation: call Tesseract or deep learning OCR model for character recognition, and perform post-processing on the recognized text; For handwritten text, use a deep learning handwriting recognition model to output structured text. The deep learning handwriting recognition model uses CRNN or Transformer-based Handwriting Recognition Model. Post-processing of recognized text includes spelling correction, character replacement, and paragraph reconstruction to ensure text readability and retrievability.

[0032] S4. Result Storage and Formatting: Save the OCR results as JSON or database records; The database records include the original image path, text content, page number information, and recognition confidence level. Low-confidence areas are automatically labeled for manual review, improving the overall recognition accuracy.

[0033] S5. Text preprocessing and feature extraction: The OCR text is processed by word segmentation, part-of-speech tagging, and named entity recognition, and keywords, phrases and document meta-information are extracted as feature vectors and input into the deep learning classification model.

[0034] S6. Document Classification Model Training and Application: Train the model using BERT, RoBERTa, or Transformer to classify document types, with labels including "Contracts," "Notices," "Meeting Minutes," and "Financial Documents." The model should be trained with at least 5000 labeled samples, through 10-20 training epochs, with a recommended learning rate of 1e-5 to 5e-5. After classification, the document categories should be written to the database, and a classification report should be generated.

[0035] S7. Semantic Retrieval Implementation: Users input queries in natural language, such as "Query 2024 personnel appointment and removal notices". The system uses a large language model to parse the query intent, matches it with knowledge graphs and text vectors, and returns relevant documents. It supports Boolean search, time range filtering, and multi-condition combinations to improve search accuracy.

[0036] S8. Entity Extraction: Using a combination of rule matching and deep learning, named entity recognition is performed on OCR text to extract information such as "time", "location", "person", "unit", and "event". The combination of rule matching and deep learning improves the ability to recognize complex sentence structures and table information.

[0037] This invention proposes a named entity recognition method based on rule matching and deep learning, used to extract key information such as "time," "location," "person," "organization," and "event" from OCR-recognized archival text. The specific steps are as follows: 1. Text standardization Standardize the OCR output text, including: Unify full-width and half-width characters; Remove redundant spaces and noisy characters; Use regular expressions to identify and standardize date and number formats, such as unifying "May 1, 2023" and "2023.5.1" into "2023-05-01".

[0038] 1. Initial screening based on rule matching Constructing a set of regular rules based on domain prior knowledge Each rule corresponds to one entity type. For example: Time entity: Location entity: r_{loc}=\texttt{([^\s]+(province|city|district|county|town|village))} Unit Entity: r_{org}=\texttt{([^\s]+(company|unit|bureau|department|committee))} Apply each rule to the text This yields a preliminary set of entities: 3. Deep learning entity recognition The text sequence is encoded, and a contextual representation is generated using a pre-trained language model (such as BERT, RoBERTa, or a domain-fine-tuned model). in For the first One word, This is the corresponding semantic vector.

[0039] Then, label dependencies are modeled using a Conditional Random Field (CRF) layer to calculate the optimal entity label sequence: in For class weight vectors, This is the label transition matrix.

[0040] 4. Integration and Conflict Resolution The rule matching results are fused with the model recognition results: If both the rule and the model identify the entity at the same location, the label with the higher confidence level of the model is retained. If a rule is identified but the model does not, then the rule confidence threshold will be applied. Decide whether to retain it; If the model identifies a new entity but there are no supporting rules, then contextual similarity is introduced. check.

[0041] Overall entity results: 5. Structured output The final extracted entities are output as structured JSON according to their categories: { Time: ["2023-05-01"], ["Chaoyang District, Beijing City"], ["Person": "Zhang San"], ["Unit": "Beijing Archives Bureau"], ["Event": "File Archiving Approval"] }.

[0042] S9, Relationship Extraction: Based on syntactic analysis and dependency analysis, extract the relationship between entities, such as "appointment relationship" and "event participation relationship". Link the cross-page and cross-archives relationship to ensure the integrity of the knowledge graph.

[0043] S10, Knowledge Graph Construction: Store entities and relationships in a graph database such as Neo4j or ArangoDB, generate nodes and edges, support visual query, and users can view the archive entity relationship network through the graphical interface.

[0044] Specific steps for generating nodes and edges are as follows: (1) Entity Standardization Perform unified processing on the extracted entity set.

[0045] Including: Remove synonyms (such as "Beijing Archives Bureau" and "Beijing Archives Bureau" unified as the same entity); Use cosine similarity-based text similarity judgment: When , it is determined as the same entity.

[0046] (2) Relationship Identification and Triple Generation Use relationship classification model to identify semantic relationships between entities to form basic semantic unit triples: Where: : Head entity (head entity); : Tail entity (tail entity); : Relationship type (relation type), such as "belongs to", "issue", "occur in", etc.

[0047] The relationship classification model can use the BERT+Softmax architecture, and its calculation formula is: Where, is the semantic vector of the head and tail entities,​ , These are trainable parameters.

[0048] (3) Graph structure generation Each entity Mapped to graph database nodes, each triple Mapped to directed edges: In Neo4j, nodes and edges can be generated using Cypher statements: MERGE(h:Entity{name:"File A"}) MERGE(t:Entity{name:"Archives Bureau"}) MERGE(h)-[:BELONGS_TO]->(t) (4) Graph optimization and indexing Create node type indexes (such as Person, Organization, Document, etc.); High-frequency relationships are aggregated and optimized to reduce duplicate edges; Through relation weights Sort the edges, where: Frequency of the relationship; Model confidence; Weighting coefficient.

[0049] (5) Visualization and Query The final knowledge graph is stored in a graph database (such as Neo4j or ArangoDB), forming a semantic network structure. Users can perform visual queries and explorations through a front-end graphical interface, supporting retrieval based on semantic paths, such as: MATCH(p:Person)-[:WORKS_IN]->(o:Organization) WHEREo.nameCONTAINS'Archives' RETURN p.name, o.name.

[0050] S11. Intelligent Question Answering and Analysis Applications: When users ask questions such as "Who was appointed as a department manager in 2024?", the system uses knowledge graphs and large-scale model reasoning to return accurate answers and source files, and supports trend analysis, such as statistics on the frequency of appointments in a certain unit and the distribution of event occurrence times.

[0051] S12, Abstract generation: generate a concise summary of the file content using an extractive or generative text summarization algorithm. Parameter setting example: summary length accounts for 10%-20% of the original text, key entities must be preserved, and summary generation model can be selected from PEGASUS or BART.

[0052] S13, Data visualization: generate visual reports of knowledge graph, summary and statistical information, support column chart, line chart and relationship diagram. Support export in PDF or HTML format, convenient for file management department to archive and decision reference.

[0053] The embodiment also proposes an artificial intelligence-based file intelligent processing system, comprising: a file digitization module for paper file digitization and OCR processing; an intelligent classification module based on a deep learning model for automatic classification, batch naming and archiving of scanned images; a knowledge extraction and graph construction module for entity extraction and relationship construction; an intelligent retrieval and question answering module supporting natural language semantic retrieval and complex question answering for user experience and query; an intelligent summary and analysis module for generating summaries and statistical reports.

[0054] The system supports batch processing of files, and users only need to upload files. The system completes digitization, classification, knowledge extraction, summary generation and visual analysis, automatically records processing logs and abnormal reports, and ensures that the processing process is traceable and reproducible.

[0055] The artificial intelligence-based file intelligent processing system and method of the present application have simple structure and are easy to operate. On the one hand, the level of the measurement platform is ensured, and the adaptability to the measurement plane is improved. On the other hand, the influence of the flowing gas around the measurement and the indoor humidity on the measured raw and auxiliary materials can be effectively avoided during use, and the work efficiency and the measurement accuracy are improved.

[0056] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent replacements to some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An artificial intelligence-based archive intelligent processing method, characterized by, Comprising the following steps: Step 1, archive scanning and image acquisition: paper archives are scanned by high-speed scanners to obtain scanned images, and batch naming and archiving of scanned images are performed; Step 2, image preprocessing: using image processing algorithms for denoising, tilt correction, and contrast enhancement; Step 3, OCR recognition and text generation: calling Tesseract or deep learning OCR model for character recognition, and post-processing of recognized text; Step 4, result storage and formatting: saving OCR results as JSON or database records; Step 5, text preprocessing and feature extraction: processing OCR text for word segmentation, part-of-speech tagging, and named entity recognition, and extracting keywords, phrases, and document metadata as feature vectors to input deep learning classification model; Step 6, document classification model training and application: using BERT, RoBERTa or Transformer model to train archive type, after classification, write archive category to database, and generate classification report; Step 7, semantic retrieval implementation: users input query sentences through natural language, system uses large language model to analyze query intent, matches knowledge graph and text vector, and returns relevant archives; Step 8, entity extraction: using rule matching and deep learning combination to perform named entity recognition on OCR text, and extract "time", "place", "person", "unit", "event" information; Step 9, relationship extraction: extracting entity relationship based on syntax analysis and dependency analysis, and linking cross-page and cross-archive relationships; Step 10, knowledge graph construction: storing entities and relationships in graph database to generate nodes and edges; Step 11, intelligent question answering and analysis application: answering user questions, system returns accurate answers and source archives through knowledge graph and large model reasoning; Step 12, abstract generation: using extractive or generative text summarization algorithm to generate concise abstract of archive content; Step 13, data visualization: generating visual reports of knowledge graph, abstract and statistical information. 2.The method of claim 1, wherein, In step 2, region segmentation is performed on tables and handwritten text, and OpenCV or deep learning model is used for layout analysis to distinguish text area, table area and image area. 3.The method of claim 1, wherein, In step 3, deep learning handwriting recognition model is used for handwritten text to output structured text. 4.The method of claim 1, wherein, In step 3, post-processing of recognized text includes spelling correction, character replacement and paragraph reconstruction. 5.The method of claim 1, wherein, In step 4, database records include original image path, text content, page information and recognition confidence, and low confidence areas are automatically labeled for manual review.

6. An artificial intelligence-based archive intelligent processing system, characterized by, Comprise: Archive digitization module for paper archive digitization and OCR processing; Intelligent classification module based on deep learning model for automatic classification, batch naming and archiving of scanned images; Knowledge extraction and graph construction module for entity extraction and relationship construction; Intelligent retrieval and question answering module supports natural language semantic retrieval and complex question answering for user experience and query; Intelligent summary and analysis module for generating abstract and statistical report.