Paper scientific and technological achievement filing method based on OCR and semantic processing and program product
By using OCR and semantic processing methods, we have achieved fully automated management of scientific and technological achievement data, solving the problems of low efficiency, semantic loss, and difficulty in processing heterogeneous data in traditional methods, and improving data utilization efficiency and innovation transformation capabilities.
Patent Information
- Application Number
- CN202511277623.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-05
AI Technical Summary
Traditional data processing methods are inefficient, paper-based archives rely on manual labor and are prone to errors, OCR technology cannot automatically annotate content hierarchy and metadata, and multi-source heterogeneous data processing is difficult, resulting in slow data querying, lack of semantic understanding and serious data silos, which hinders the efficient utilization and innovative transformation of scientific and technological achievements.
The method adopts OCR and semantic processing, segments image regions through object detection model, combines OCR recognition and language model to correct text, extracts metadata, generates knowledge graph, and uses database, graph database and object storage to associate data, so as to realize full-process automated management.
It has achieved fully automated processing of scientific and technological achievement data, improved data retrieval efficiency and semantic understanding, reduced error rate, supported unified management of multi-source heterogeneous data, and promoted knowledge sharing and innovation transformation.
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, in particular to a scientific achievement data batch processing method based on OCR and semantic processing. BACKGROUND
[0002] Under the background of accelerating scientific and technological innovation, scientific achievement data, as the core carrier of knowledge assets, usually exists in the form of paper archives and electronic documents, and its content covers key information such as technical reports, patent documents, and experimental data. These data play an important role in scientific research management, technology transformation, and intellectual property protection. However, traditional data processing methods have significant defects: first, the retrieval efficiency is low, especially for paper archives, which usually rely on manual input and indexing. This not only consumes a lot of manpower and time (e.g., page-by-page scanning and proofreading), but also introduces errors, resulting in slow data query response and insufficient reliability. Second, although existing optical character recognition (OCR) technology can achieve basic text recognition, it cannot automatically annotate content hierarchy (such as chapter division or logical relationship) and metadata (such as title, author, and secret level), resulting in a lack of semantic understanding, which affects the accurate classification and in-depth analysis of data. Finally, the processing of multi-source heterogeneous data is difficult, and scientific achievement data is widely sourced and has various formats. Traditional methods are difficult to process these heterogeneous types uniformly, resulting in serious data island phenomenon, reducing overall management efficiency and the possibility of knowledge sharing. The combination of these problems not only increases the cost and complexity of data processing, but also hinders the efficient use and innovation transformation of scientific achievements, and more advanced solutions are needed to improve the level of intelligence. SUMMARY
[0003] To solve the above problems, the present application provides the following technical solutions
[0004] The present application provides a paper scientific achievement archiving method based on OCR and semantic processing, comprising the following steps:
[0005] S1: Scanning paper scientific achievements to obtain original scanned copies and then obtaining archive images;
[0006] Pretreating the archive images;
[0007] Using a target detection model to segment the archive images, and the segmented archive images include text regions, table regions, and picture regions;
[0008] S2: Converting the archive images into digital documents, the specific steps are:
[0009] 1) Performing OCR recognition on the text regions in the archive images to convert the text regions from image form to text form;
[0010] 2) Extract information from the table areas in the archive images and convert the table areas from image form into text form containing table structure information;
[0011] 3) Use a language model to perform grammatical correction on the text content in the text area and table area;
[0012] S3: Extract corresponding metadata from the digital document; the metadata is information representing a certain attribute of the scientific and technological achievement; the metadata includes keywords;
[0013] For metadata recorded in a fixed format, extract it using rule matching.
[0014] Metadata for records with non-fixed formats is extracted using a pre-trained machine learning model;
[0015] For digital documents that do not record keywords, use TF-IDF or TextRank algorithms to extract keywords;
[0016] S4: Store digital documents and their corresponding metadata as structured data in a database of paper-based scientific and technological achievements;
[0017] Based on digital documents and their corresponding metadata, a knowledge graph is generated, and the paper-based scientific and technological achievements in the form of the knowledge graph are stored in a graph database.
[0018] The original scanned copies of paper-based scientific and technological achievements are stored in the object storage;
[0019] Establish connections between the contents of the same paper-based scientific and technological achievement in databases, graph databases, and object stores.
[0020] Preferably, the preprocessing includes image noise reduction, binarization, and geometric correction.
[0021] Preferably, Gaussian filtering is used for image denoising, adaptive thresholding is used for binarization, and tilt correction based on Hough transform is used for geometric correction.
[0022] Preferably, the target detection model is YOLOv5 or Mask R-CNN.
[0023] Preferably, when performing OCR recognition on text regions, printed text is recognized using PP-OCRv3 or Tesseract 5.0, and handwritten text is recognized using a pre-trained CRNN model; when the confidence level is lower than the threshold, manual review is requested.
[0024] Preferably, when extracting information from a table area, the TableNet or DeepDeSRT model is used for extraction.
[0025] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements the method described above.
[0026] Advantages:
[0027] The application realizes the full-process automation of the data of scientific and technological achievements from collection, processing to storage through multi-modal OCR recognition, dynamic semantic annotation, intelligent data analysis and knowledge fusion technology. Through the complementation of storage forms, the utilization value of paper scientific and technological achievements is maximized, and reliable support is provided for scientific research management and intellectual property protection. DETAILED DESCRIPTION
[0028] In order to make the objectives, characteristics and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the following described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0029] A paper scientific and technological achievement archiving method based on OCR and semantic processing, comprising the following steps:
[0030] S1: scanning the paper scientific and technological achievements to obtain original scanned copies and then obtain archive images;
[0031] The archive images are preprocessed; the preprocessing includes image denoising, binarization and geometric correction;
[0032] In the embodiment, a high-resolution scanner (≥600 dpi) is used to obtain the archive images of paper documents; Gaussian filtering is used for image denoising, adaptive threshold method is used for binarization, and geometric correction based on Hough transform is used for geometric correction;
[0033] A target detection model is used for region segmentation of the archive images, and the archive images are segmented into text regions, table regions and picture regions;
[0034] In the embodiment, the target detection model uses YOLOv5 or Mask R-CNN;
[0035] S2: converting the archive images into digital documents, and the specific steps are:
[0036] 1) performing OCR recognition on the text regions in the archive images to convert the text regions from image form to text form;
[0037] In this embodiment, when performing OCR recognition on the text area, printed characters are recognized using PP-OCRv3 or Tesseract 5.0, and handwritten characters are recognized using a pre-trained CRNN (Convolutional Recurrent Neural Network) model; when the confidence is lower than a threshold, manual review is requested;
[0038] 2) Information extraction is performed on the table area in the archive image to convert the table area from an image form to a text form containing table structure information; wherein the table area in the text form takes HTML or CSV format as a carrier;
[0039] In this embodiment, when performing information extraction on the table area, TableNet or DeepDeSRT model is used for extraction;
[0040] 3) A language model (such as BERT or GPT-3) is used to perform grammar correction on the text content of the text area and the table area;
[0041] For example, "solar energy bubble" is corrected to "solar cell".
[0042] S3: Extracting corresponding metadata from the digital document; the metadata is information representing a certain attribute of the scientific and technological achievement; the metadata includes keywords;
[0043] For metadata recorded in a fixed format, the extraction is performed in this embodiment using a rule matching method, and common rule matching methods include regular expressions; for example:
[0044] The metadata of the secret level information is extracted from "secret level: confidential | secret | internal";
[0045] The metadata of the title is extracted from "title: 《XXX》" or the first line of bold text;
[0046] The metadata of the author is extracted from "author: XXX" or the email suffix (such as "@xxx.edu.cn");
[0047] For metadata recorded in a non-fixed format, a pre-trained machine learning model is used for extraction; in this embodiment, a pre-trained BiLSTM-CRF model is used for extraction;
[0048] For digital documents without recorded keywords, TF-IDF or TextRank algorithm is used to extract keywords;
[0049] S4: Storing the digital document and the corresponding metadata as a paper scientific and technological achievement in the form of structured data into a database (such as MySQL);
[0050] According to the digital document and the corresponding metadata, a knowledge graph is generated, and the paper scientific and technological achievements in the form of a knowledge graph are stored in a graph database (such as Neo4j);
[0051] The original scanned copy of the paper scientific and technological achievements is stored in an object storage (such as MinIO);
[0052] The contents of the same paper scientific and technological achievements in the database, the graph database and the object storage are associated.
[0053] The paper scientific and technological achievements are stored in three associated forms (database, graph database and object storage), and the technical effects of the technical solution are significantly reflected in the improvement of data processing efficiency and knowledge utilization depth. The specific effects can be classified from the following four angles to ensure comprehensive coverage of the complementary advantages of the storage mode:
[0054] 1) The effect of database (such as MySQL) storage: improve the retrieval efficiency and management convenience of structured data
[0055] Fast and accurate retrieval: structured data storage (including digital documents and their metadata) supports efficient SQL queries, users can perform millisecond-level filtering and indexing based on metadata (such as title, author, keyword or secret level), solving the problem of low retrieval efficiency caused by manual entry in the background technology (for example, the time to retrieve "solar cell" related documents is shortened from hours to seconds).
[0056] Error reduction and cost optimization: replace manual proofreading with automated storage to reduce data entry errors (such as OCR recognition followed by syntax correction to ensure data accuracy), while reducing labor costs.
[0057] Support batch operation: facilitate batch import / export of data, adapt to large-scale data update requirements in scientific research management, and improve overall management efficiency.
[0058] 2) The effect of graph database (such as Neo4j) storage: enhance semantic understanding and knowledge association analysis
[0059] Deep semantic query: knowledge graph storage captures the logical relationships between documents (such as "author-invention topic-keyword" associations generated based on metadata), supports complex graph query (for example, query "all patents related to 'new energy' and with a secret level of 'confidential' through Cypher language"), solving the defect that OCR cannot automatically annotate semantics in the background technology.
[0060] Knowledge discovery and innovation assistance: visual graph facilitates the discovery of hidden patterns (such as technology trend analysis or cross-domain association), accelerating the transformation of scientific and technological achievements (for example, identifying the intersection of "solar cell" and "energy storage technology", promoting innovative applications).
[0061] Adaptive expansion: Knowledge graph can be dynamically updated (e.g., automatically expand nodes when new documents are added), suitable for multi-source heterogeneous data integration, reducing data island phenomenon.
[0062] 3) The effect of object storage (such as MinIO) storage: Ensure data integrity and traceability
[0063] Preservation of raw data: Lossless storage of raw scans provides audit evidence, facilitating verification of data consistency before and after OCR processing (e.g., quickly retrieve original archives to verify authenticity in intellectual property disputes), solving the risk of data loss or damage in traditional methods.
[0064] High reliability and scalability: Object storage supports distributed architecture, ensuring stable access and backup of massive data (such as TB-level scans), reducing storage costs.
[0065] Convenient access and sharing: Associating structured data through unique identifiers (such as UUID), realizing one-key original archive retrieval, promoting cross-department knowledge sharing (e.g., researchers can directly download original scans for review).
[0066] 4) The effect of overall associated storage: Realize the advantages of full-process automation and collaboration
[0067] End-to-end automation: The association of the three storage forms (database storage of logical data, graph database storage of semantic association, and object storage of original files) forms a closed loop, supporting seamless processes from collection to analysis, significantly improving processing speed (e.g., archival time is reduced from several days in traditional methods to several hours).
[0068] Multi-dimensional analysis enhancement: Combining structured queries and graph analysis, supporting complex application scenarios (such as research evaluation report generation or technology risk prediction), solving the problem of multi-source heterogeneous data processing in the background technology.
[0069] Innovation transformation promotion: Through efficient data utilization, reduce the complexity of scientific and technological achievement management, and accelerate the commercialization process (e.g., enterprises can quickly identify high-value patents for investment based on knowledge graphs).
[0070] The above technical effects comprehensively reflect the advancement of the invention in solving the problems of retrieval efficiency, semantic loss, and heterogeneous data. Through the complementarity of storage forms, the utilization value of paper scientific and technological achievements is maximized, and reliable support is provided for scientific research management and intellectual property protection.
[0071] The above description is only the preferred embodiment of the present invention. It should be noted that for ordinary skilled persons in the technical field, without departing from the principles of the invention, several improvements and refinements can be made, and these improvements and refinements should be considered within the protection scope of the invention.
Claims
1. A paper technology achievement archiving method based on OCR and semantic processing, characterized in that, The method comprises the following steps: S1: scanning a paper scientific achievement to obtain an original scan and then an archive image; preprocessing the archive image; using a target detection model to perform region segmentation on the archive image, the segmented archive image comprising a text region, a table region, and a picture region; S2: converting the archive image into a digital document, specifically comprising the following steps: 1) performing OCR recognition on the text region in the archive image to convert the text region from an image form into a text form; 2) performing information extraction on the table region in the archive image to convert the table region from an image form into a text form containing table structure information; 3) using a language model to perform grammar correction on the text content of the text region and the table region; S3: extracting corresponding metadata from the digital document; the metadata is information representing a certain attribute of the scientific achievement; the metadata comprises keywords; for metadata recorded in a fixed format, the metadata is extracted using a rule matching method; for metadata recorded in a non-fixed format, the metadata is extracted using a pre-trained machine learning model; for a digital document without recorded keywords, keywords are extracted using a TF-IDF or TextRank algorithm; S4: storing the digital document and the corresponding metadata as a paper scientific achievement in a structured data form in a database; generating a knowledge graph according to the digital document and the corresponding metadata, and storing the paper scientific achievement in the form of the knowledge graph in a graph database; storing the original scan of the paper scientific achievement in an object storage; associating the contents of the same paper scientific achievement in the database, the graph database, and the object storage.
2. The paper scientific achievement archiving method based on OCR and semantic processing according to claim 1, characterized in that, The preprocessing comprises image denoising, binarization, and geometric correction.
3. The paper scientific achievement archiving method based on OCR and semantic processing according to claim 2, characterized in that, Gaussian filtering is used for image denoising, an adaptive threshold method is used for binarization, and a Hough transform-based tilt correction is used for geometric correction.
4. The paper scientific achievement archiving method based on OCR and semantic processing according to claim 1, characterized in that, The target detection model is YOLOv5 or Mask R-CNN.
5. The paper scientific achievement archiving method based on OCR and semantic processing according to claim 1, characterized in that, When performing OCR recognition on the text region, printed text is recognized using PP-OCRv3 or Tesseract 5.0, and handwritten text is recognized using a pre-trained CRNN model; when the confidence is lower than a threshold, manual review is requested.
6. The paper scientific achievement archiving method based on OCR and semantic processing according to claim 1, characterized in that, When performing information extraction on the table region, TableNet or DeepDeSRT model is used for extraction.
7. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the method of any one of claims 1 to 6.
Citation Information
Cited By
Intelligent archive digital processing method and system
CN121502056A