Electronic data classification and retrieval system for archives

By identifying and preprocessing electronic and paper archives, combining optical character recognition and image description, generating labels and using deep learning models for classification and retrieval, the problem of low efficiency in traditional archive management is solved, and efficient and accurate archive management and retrieval are achieved.

CN120747979APending Publication Date: 2025-10-03WEIFANG NURSING VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510842429.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional archive management methods are inefficient, manual classification is inconsistent, retrieval results are inaccurate, paper archives take up a lot of space and are easily damaged, and are difficult to share and integrate between different departments or systems, forming information islands.

Method used

The data acquisition module is used to identify and preprocess electronic and paper archives, combined with optical character recognition and image description, and labels are generated through semantic analysis. Deep learning models are used for classification and retrieval, and a manual proofreading mechanism is integrated to improve accuracy.

Benefits of technology

Improve the speed of archive processing, reduce manual intervention, ensure the accuracy and consistency of retrieval, reduce physical storage requirements, improve the flexibility and security of data management, support multimodal retrieval, and optimize user experience.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses an electronic data classification and retrieval system for archives, and belongs to the technical field of archive induction and retrieval. Comprising a data acquisition module, a content identification and conversion module, a data processing module, a classification module and a storage module, wherein the data acquisition module is used for acquiring and preprocessing characters or pictures in an electronic file and characters or pictures in a paper file after identification; through automatic data collection, preprocessing, OCR recognition and label generation, the file processing speed is remarkably increased, the requirement for manual intervention is reduced, meanwhile, based on efficient label indexing and retrieval engines, a user can quickly find needed information in massive files, and the information searching time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of archive classification and retrieval, and in particular relates to an electronic data classification and retrieval system for archives. Background Art

[0002] In the information age, various organizations, including businesses, government agencies, and educational institutions, have accumulated vast quantities of paper and electronic archives. These archives not only contain a wealth of historical data and business information, but also play a key role in daily operations, decision support, and legal compliance. However, traditional archive management methods rely primarily on manual classification and retrieval. With the rapid growth in the number of archives, this approach has become inefficient and unable to meet the needs of rapid retrieval. Manual classification is also prone to subjective factors, leading to inconsistent standards and inaccurate retrieval results. Furthermore, large numbers of paper archives occupy significant physical space and are susceptible to environmental factors, making them susceptible to damage. Archives between different departments or systems are difficult to share and integrate, forming information silos that limit the overall value of the archives.

[0003] In order to meet these challenges, digital and intelligent archive management solutions have become particularly important. By digitizing paper archives and utilizing optical character recognition (OCR) and natural language processing (NLP) technologies, the efficiency and accuracy of archive management can be significantly improved. Intelligent classification and retrieval systems can automatically generate tags based on content, support multi-dimensional precise searches, and meet users' growing demand for information acquisition. In recent years, the rapid development of artificial intelligence (AI) and machine learning (ML) technologies has provided strong technical support for the implementation of intelligent archive management systems. Deep learning technologies such as convolutional neural networks (CNN), recurrent neural networks (RNN), and Transformer models have made significant progress in image processing and text understanding, enabling the system to efficiently and accurately perform image preprocessing, text recognition, and semantic analysis.

[0004] However, for the existing text and image information classification system, the text and image information is extracted and then classified, but it only enters the original text and image information of the file. When searching subsequently, the corresponding text and image information needs to be accurately input. In actual conditions, it is necessary to take into account the replacement of staff and forgetfulness caused by time, and the original entered text cannot be fully described, which leads to inaccurate retrieval (many files with similar semantics will appear). Most importantly, the image information cannot be accurately matched to the meaning expressed by the image through text description, and thus the image cannot be accurately retrieved. Summary of the Invention

[0005] To address the above shortcomings, the present invention provides an electronic data classification and retrieval system for archives, comprising the following modules:

[0006] Data collection module, used to collect and pre-process text or images in electronic files and paper files after recognition;

[0007] The content recognition and conversion module includes an optical character recognition unit and an image description unit. The optical character recognition unit is used to scan paper files and electronic documents for optical character recognition processing and extract editable text. The image description unit generates a text description after analyzing the image using a visual model.

[0008] a data processing module, comprising a cleaning and standardization unit for cleaning the available editable text and the generated text description, a semantic analysis unit for performing semantic recognition on the cleaned text, and a label generation unit for forming labels based on the semantics;

[0009] A classification module is used to aggregate the labels of the image text content generated by the label generation unit into an image data set, and aggregate the original text content labels generated by the label generation unit into a text data set, thereby forming a separated data set;

[0010] The storage module is used to store the text information and identification information generated by the above modules.

[0011] Furthermore, the data acquisition module includes an electronic document import port, a scanning unit for scanning paper files and converting them into electronic data, and an image pre-processing unit for correcting and segmenting images.

[0012] Furthermore, the content recognition and conversion module further includes a description fusion unit for merging the extracted text with the generated image description to form a unified text representation.

[0013] Furthermore, the data processing module also includes a correction unit for language correction and character correction, which uses context information through a language model to correct common recognition errors, and corrects word-level errors through a spelling check algorithm, and evaluates and corrects character-level errors through a Levenshtein distance algorithm.

[0014] Furthermore, the correction unit is integrated with a manual proofreading mechanism.

[0015] Compared with the prior art, the present invention has the following beneficial effects:

[0016] Through automated data collection, preprocessing, optical character recognition, and label generation, the speed of archive processing is significantly improved, reducing the need for manual intervention. By generating text descriptions and combining them with semantic recognition, a processing model consistent with text archives is achieved. Unique identification information ensures accurate distinction between similar images, shortening information search time. Label generation utilizes natural language processing (NLP) technology and deep learning models to ensure consistency and accuracy in archive classification. Semantic understanding capabilities provide more accurate search results through semantic similarity calculation and context analysis, reducing false positives and missed detections.

[0017] Digital storage through storage modules reduces the need for physical storage space, and electronic storage improves the security and durability of archives. The design of data separation and association avoids data confusion and improves the flexibility and scalability of data management.

[0018] It supports multimodal retrieval methods based on text, images, and a combination of images and text to meet the needs of different users, improve user experience, continuously optimize OCR recognition and deep learning models, and enhance the intelligence level and performance of the system. DETAILED DESCRIPTION

[0019] The following will be combined with the contents of the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] Example

[0021] This embodiment provides an electronic data classification and retrieval system for archives, which specifically includes the following modules:

[0022] The data acquisition module is divided into the import of electronic files (PDF, Word, Excel, and other formats) (imported through the electronic document import port) and the scanning of paper files and their conversion into electronic data (scanning is completed by the scanning unit). For the text content in electronic files (referring to the text information printed on paper files), the paper files are scanned using a high-resolution scanner to convert the paper files into PDF files, and then the data acquisition module extracts the text from the PDF files. For the handwritten text content in paper files, the spatial features of the characters are extracted through a convolutional neural network, and the sequence information of the characters is then processed through a recurrent neural network to capture the temporal features of the handwritten text. The Transformer model is then used to capture the global context, improve recognition accuracy, and perform text recognition at the same time, which can enhance the model's understanding of handwritten text.

[0023] The image correction and segmentation processing is achieved through the image preprocessing unit, which specifically includes the following functions:

[0024] Denoising and correction functions: Apply image processing technologies such as denoising (using filtering techniques such as median filtering and Gaussian filtering to remove noise from images and improve image quality), rotation correction (detecting and correcting skew in images to ensure horizontal or vertical alignment of text), perspective transformation (correcting perspective distortion caused by scanning or shooting angles to restore the true form of text), and contrast adjustment (using histogram equalization or adaptive contrast enhancement such as CLAHE to increase image contrast and make characters clearer) to improve subsequent recognition accuracy;

[0025] Resolution enhancement: Utilize deep learning models, such as those in the field of image super-resolution reconstruction (SRCNN, ESRGAN), to convert low-resolution images into high-resolution images, enhance detail information to achieve super-resolution reconstruction, and then use interpolation algorithms, such as bilinear or bicubic interpolation methods, to magnify the image, while combining deep learning methods to improve the effect.

[0026] Segmentation and region detection functions: Identify different regions in an image, such as text areas, image areas, and table areas, so that they can be processed separately;

[0027] The data acquisition module can also be equipped with a mechanism for automatically identifying file types (text or image) to facilitate automated selection for subsequent processing.

[0028] The content recognition and conversion module processes the text information collected by the data acquisition module and includes an optical character recognition unit and an image description unit. The optical character recognition unit scans paper files and electronic documents for optical character recognition (OCR) and extracts editable text. The image description unit analyzes images using a visual model (such as the Vision AI model) to generate detailed text descriptions. The OCR and image description models support multiple languages ​​to meet global needs.

[0029] It should be noted that each file to be recognized needs to be processed by the optical character recognition unit and the image description unit. When processing a single text file, the image description unit does not work because it cannot recognize the image content. When processing a single image file, the optical character recognition unit does not work because it cannot recognize the text content. If the file contains both text and graphics, the description fusion unit will merge the text extracted by OCR with the generated image description to form a unified text representation, which is convenient for subsequent semantic analysis and label generation. It should be noted that in the process of image processing, each different image file is represented by a unique identifier (i.e., UUID);

[0030] The data processing module is used to process the text output by the content recognition and conversion module, specifically including:

[0031] A cleaning and standardization unit that cleans editable text and generated text descriptions;

[0032] The semantic analysis unit is used to perform semantic recognition on the cleaned text. Before analysis, the word description is also processed with word segmentation and part-of-speech tagging, as well as entity recognition (NER) to obtain key entities such as names of people, places, dates, and organizations. Then, deep learning models such as BERT and GPT are used to extract document topics and understand semantic content.

[0033] The correction unit has language correction and character correction functions. The language correction function refers to inputting the text output by OCR into a preset language model (such as BERT, GPT), using contextual information to correct common recognition errors, and correcting word-level errors through spelling checking algorithms. The character correction function is to use the Levenshtein distance algorithm to evaluate and correct character-level errors, thereby improving the recognition accuracy of proper nouns and terms.

[0034] The tag generation unit generates multiple relevant tags for each file based on keyword extraction and deep learning models, specifically:

[0035] Text archive: Generates tags describing the original text content from the text processed by the data processing module;

[0036] Image archive: Generates labels for image text content using the text description processed by the data processing module;

[0037] Finally, semantic similarity calculations (such as Word2Vec and BERT embedding) are used to ensure the accuracy and relevance of the labels.

[0038] The classification module is used to aggregate the labels of the image text content generated by the label generation unit into an image dataset, and aggregate the labels of the original text content generated by the label generation unit into a text dataset, forming separate datasets instead of the traditional integrated dataset, so as to reduce confusion;

[0039] The record structure of the text dataset is: text content, tags and metadata (such as creation date, author, etc.);

[0040] The record structure of the image dataset is: image description text, label, UUID and metadata (such as creation date, author, etc.);

[0041] The storage module is used to store the text information and identification information generated by the above modules.

[0042] It should be noted that the correction unit is also integrated with a manual proofreading mechanism. In order to reduce the workload of manual work, the work that needs to be proofread is only the combination of images and texts processed by the description fusion unit (because the base number of images and texts combined by the description fusion unit is relatively small in terms of the overall number of archives, and the meaning of images and texts may be recorded based on human subjective factors, so manual proofreading can improve the accuracy and facilitate subsequent retrieval). Proofreading is based on the experience of the staff (manual proofreaders can make more appropriate correction suggestions based on the overall meaning and context of the text. Some subtle language and cultural differences are difficult for machines to recognize, which requires human intuition and experience).

[0043] The following describes the classification and retrieval process of text and image archives respectively:

[0044] Classification and retrieval of text archives include:

[0045] S1. Upload text files → directly store them in the text file dataset;

[0046] S2, the label generation module extracts labels from the text content, such as "finance", "report", and "2023";

[0047] S3, tag storage and index construction ensure that tags are associated with text content;

[0048] S4. The user searches for "financial report" on the interface → the search engine matches the tag in the text archive dataset and returns the relevant text archives;

[0049] S5. Correction unit corrects possible errors to ensure the accuracy of the text.

[0050] Classification and retrieval of image archives:

[0051] S1. Upload image files → image preprocessing to improve quality;

[0052] S2, the optical character recognition unit extracts the text in the image, and the image description unit generates a description;

[0053] S3, the tag generation module extracts tags from the description, such as "conference room", "Zhang San", and "financial report";

[0054] S4, tag storage and index construction ensure that tags are associated with image descriptions;

[0055] S5. The user searches for “financial report” on the interface and selects an image dataset. The search engine matches the tags in the image description dataset and returns relevant image archives.

[0056] S6. The correction unit corrects errors in the description to ensure that the image description is accurate.

[0057] It should be noted that index construction refers to establishing a quick search mapping from "label" to "archive identifier containing the label" (i.e., inverted index), so that when users search, the system does not need to scan the complete content of all archives, but only needs to quickly locate the relevant archive identifier through the index, which greatly improves the retrieval efficiency.

[0058] It should be noted that the structure described in the present invention can be implemented in a variety of different forms and is not limited to the described embodiments. Any equivalent transformations made by ordinary technicians in this field using the contents of the present invention specification, or directly or indirectly applied to other related technical fields, such as the loading and unloading of other items, are included in the scope of protection of the present invention.

Claims

1. An electronic data classification and retrieval system for archives, characterized in that: Includes the following modules: Data collection module, used to collect and pre-process text or images in electronic files and paper files after recognition; The content recognition and conversion module includes an optical character recognition unit and an image description unit. The optical character recognition unit is used to scan paper files and electronic documents for optical character recognition processing and extract editable text. The image description unit generates a text description after analyzing the image using a visual model. a data processing module, comprising a cleaning and standardization unit for cleaning the available editable text and the generated text description, a semantic analysis unit for performing semantic recognition on the cleaned text, and a label generation unit for forming labels based on the semantics; A classification module is used to aggregate the labels of the image text content generated by the label generation unit into an image data set, and aggregate the original text content labels generated by the label generation unit into a text data set, thereby forming a separated data set; The storage module is used to store the text information and identification information generated by the above modules.

2. The electronic data classification and retrieval system for archives according to claim 1, characterized in that: The data acquisition module includes an electronic document import port, a scanning unit for scanning paper files and converting them into electronic data, and an image pre-processing unit for correcting and segmenting images.

3. The electronic data classification and retrieval system for archives according to claim 1, wherein: The content recognition and conversion module further includes a description fusion unit for merging the extracted text with the generated image description to form a unified text representation.

4. The electronic data classification and retrieval system for archives according to claim 1, wherein: The data processing module also includes a correction unit for language correction and character correction, which corrects errors using context information through a preset language model, corrects word-level errors through a spelling check algorithm, and evaluates and corrects character-level errors through a Levenshtein distance algorithm.

5. The electronic data classification and retrieval system for archives according to claim 4, characterized in that: The correction unit is integrated with a manual proofreading mechanism to perform manual proofreading on the image and text combination processed by the description fusion unit.