An intelligent management system for automatic document recognition based on OCR

By building an intelligent document management system, utilizing multimodal preprocessing and deep OCR recognition technology, and combining knowledge graphs and blockchain, we have solved the problems of low accuracy and poor security in document recognition, achieved fine-grained control and efficient process management, and improved the intelligence level of document management.

CN120564202BActive Publication Date: 2025-09-26ANHUI LVBEN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511055398.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-26
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Existing technologies in official document management have problems such as low recognition accuracy, poor security, loose permission control, single retrieval function, and inefficient process management, making it difficult to meet intelligent needs.

Method used

An intelligent document management system is constructed by adopting modules such as multimodal document preprocessing, deep fusion OCR recognition engine, semantic understanding and knowledge graph construction, permission control and security audit, classification and metadata annotation, retrieval and recommendation, process automation, quality assessment and optimization, and visual decision support, combined with convolutional neural networks, generative adversarial networks, knowledge graphs, blockchain technology, etc.

Benefits of technology

It improves the accuracy of official document recognition, realizes fine-grained security control, enhances the convenience of information acquisition and process efficiency, supports cross-modal retrieval and personalized recommendations, provides real-time monitoring and optimization capabilities, and improves the intelligence level of official document management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564202B_ABST
    Figure CN120564202B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent management system for automatic document recognition based on OCR, which relates to the field of intelligent document OCR processing technology. The preprocessing module uses adaptive filtering and GAN network to repair low-quality documents and separate seals from text; the recognition engine uses a multi-scale feature pyramid network to achieve end-to-end recognition of tilted text and mixed text; the semantic module constructs a dynamic knowledge graph to support event evolution reasoning and cross-year association analysis; the classification module implements small sample classification based on the prototype network, and the security module implements fine-grained authority control and anomaly detection through blockchain and federated learning. After system processing, the clarity of text edges is improved and real-time reasoning is supported. The present invention improves the intelligent level of document processing through the collaboration of multiple modules. The preprocessing and recognition modules solve complex scene problems, and semantic understanding and knowledge graphs achieve in-depth analysis and association analysis; the security module ensures data security; and comprehensively improves efficiency and scientific decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent official document OCR processing, and in particular to an intelligent management system for automatic official document recognition based on OCR. Background Art

[0002] With the popularization of digital office work, fields such as electricity, government affairs, and enterprises have placed higher demands on the efficiency and accuracy of document management. Traditional manual processing methods are not only inefficient but also prone to errors and omissions, making it difficult to meet the needs of rapid processing of massive amounts of documents. Although existing optical character recognition (OCR) technology can achieve a certain degree of text extraction, it still faces many challenges in actual application. Quality issues such as scratches, stains, and curling on scanned documents, as well as complex situations such as seal coverage and handwritten text, often lead to a significant decrease in recognition accuracy. In addition, the professional nature of the content of official documents and the diversity of formats make semantic understanding and key information extraction difficult, making it impossible to effectively support intelligent management.

[0003] Traditional systems also have significant shortcomings in document security management. Permission control is often based on simple role-based divisions, making it difficult to achieve fine-grained access control for sensitive documents. Security audits primarily rely on manual log checks, failing to promptly detect abnormal operations and potential risks. With the increasing incidence of data breaches, ensuring the security of document storage, transmission, and use has become a pressing issue. Furthermore, cross-departmental and cross-system document transfers require collaboration from multiple parties, and traditional systems lack effective mechanisms for permission coordination and data sharing, which can easily lead to information silos and security vulnerabilities.

[0004] Existing document retrieval and process management systems are limited in functionality and fail to meet intelligent needs. Search functions are primarily based on keyword matching, failing to understand semantic associations, resulting in low relevance in search results. Process approvals rely on fixed templates, lacking flexibility and dynamic optimization capabilities, leading to low efficiency in complex approval scenarios. Furthermore, the system lacks real-time monitoring and automatic optimization mechanisms for document processing quality, making it unable to adapt to dynamic changes in business needs and hindering efficient management of documents throughout their lifecycle. Summary of the Invention

[0005] The present invention proposes an intelligent management system for automatic document recognition based on OCR to solve the problems mentioned in the above-mentioned prior art.

[0006] In order to achieve the above objectives, the present invention adopts the following technical solution: an intelligent management system for automatic document recognition based on OCR, comprising the following modules:

[0007] Multimodal document preprocessing module: This module uses a dynamic threshold adaptive hybrid filtering algorithm, a convolutional neural network to predict noise distribution and generate compensation masks, three-dimensional reconstruction technology to unfold curled documents, a generative adversarial network to enhance document resolution, and a physical model to inversely optimize preprocessing parameters.

[0008] Deep Fusion OCR Recognition Engine: Builds a hierarchical feature extraction network and introduces a formula to improve model inference speed: , To optimize the inference speed, is the initial inference speed, To save computational effort, The original total computational effort; developing seal penetration recognition technology, separating seals and text based on polarized light imaging, and using phase correlation algorithms to compensate for text deformation;

[0009] Semantic Understanding and Knowledge Graph Construction Module: Build a cross-document semantic association framework, use the self-attention mechanism to extract key semantic units, generate a dynamic knowledge graph using temporal embedding technology, develop an event tracing algorithm, generate a policy evolution graph based on citation relationships and semantic similarity, and introduce adversarial training to optimize graph completion;

[0010] Classification and metadata annotation module: Build a multimodal feature fusion network, input document visual layout, text semantics, and format features, and introduce a formula to improve annotation efficiency: , is the optimized annotation efficiency, is the initial annotation efficiency, is the amount of knowledge of effective labeling habits, The total amount of annotation knowledge; develop an interactive annotation assistance system, use reinforcement learning to understand the user's annotation habits, and automatically recommend candidate tags;

[0011] Permission control and security audit module: Build a dynamic access control system based on attribute-based encryption, combine blockchain technology, design a user behavior profiling system, analyze operation sequences through long-short-term memory networks, and introduce abnormal behavior detection formulas: , where R represents recall, TP is true positive, and FN is false negative.

[0012] Furthermore, the multimodal document preprocessing module also includes: addressing the fading problem of thermal paper documents, using spectral reconstruction technology to restore text information, reconstructing the original text outline by analyzing image differences under different wavelengths, developing a document adhesion separation algorithm, using deep learning to predict paper edges, and combining physical simulation algorithms to achieve automatic separation without damaging the document.

[0013] Furthermore, the deep fusion OCR recognition engine also includes: designing an optimization solution for multi-language mixed text recognition, switching recognition parameters in real time through a language detection module, combining character morphology analysis to eliminate character interference between languages, developing handwriting and printed text hybrid recognition technology, using a generative adversarial network to convert handwriting styles into standard fonts, and then recognizing them through a unified network.

[0014] Furthermore, the semantic understanding and knowledge graph construction module also includes: introducing a logical reasoning model, conducting compliance checks on official document content based on a common sense knowledge base and a policy and regulation base, and introducing a knowledge graph update relevance formula: , where U represents the updated relevance of the knowledge graph, is the number of newly added nodes, is the total number of original nodes in the knowledge graph, is the number of newly generated relations, is the total number of original relations, It is the node weight coefficient. A dynamic knowledge graph update mechanism is developed. When a new document is released, the graph neural network is used to automatically update the associated nodes and relationships based on this formula.

[0015] Furthermore, the classification and metadata annotation module also includes: building an active learning-driven annotation optimization system, automatically screening documents to be annotated through uncertainty sampling and density estimation, designing a semantically enhanced metadata extractor, and using an external knowledge base to expand the extraction scope.

[0016] Furthermore, the permission control and security audit module also includes: developing a cross-departmental security collaboration solution based on federated learning, jointly training anomaly detection models without sharing original data, designing dynamic permission adjustment strategies, raising or lowering permissions in real time based on user operation risk scores, and automatically triggering secondary verification when high-risk behavior is detected, shortening the security incident response time to less than 3 seconds.

[0017] Furthermore, it also includes: a retrieval and recommendation module, which adopts a Transformer-based cross-modal retrieval model, supports mixed retrieval of text, images, and voice, and introduces a recommendation hit rate formula: , where H is the recommendation hit rate, is the document semantic similarity, is the user behavior weight, is the scene context matching degree, is the scene weight, is the maximum similarity value, For the total weight value, the module builds a personalized recommendation engine, which combines user historical behavior and document semantic similarity to automatically recommend relevant policy documents and past cases in scenarios such as project approval.

[0018] Furthermore, it also includes: process automation module, design of intelligent routing algorithm based on reinforcement learning, dynamic planning of approval paths according to factors such as the urgency of documents and the load of processors, development of RPA robot process orchestration tools, and support for low-code custom automation tasks.

[0019] Furthermore, it also includes: quality assessment and optimization module, construction of a multi-dimensional quality monitoring system, real-time evaluation from 12 dimensions, development of a model automatic evolution framework, optimization of model structure and parameters through genetic algorithms, and continuous training based on user feedback data.

[0020] Furthermore, it also includes: a visual decision support module, development of a spatiotemporal information visualization platform, displaying the regional distribution of document processing progress through a geographic information system, presenting the evolution of policy implementation effects in combination with a timeline, designing a three-dimensional knowledge graph interactive interface, supporting users to explore document correlations through gestures and voice, and generating visual analysis reports.

[0021] Compared with the existing technology, the beneficial effects of the present invention are:

[0022] During the document preprocessing and recognition phase, multimodal enhancement technology and a deeply integrated OCR engine effectively address challenges such as low-quality documents and stamp coverage, significantly improving recognition accuracy and ensuring the accuracy of information extraction. The semantic understanding and knowledge graph construction module enables in-depth parsing and correlation analysis of document content, automatically analyzing policy evolution and identifying potential conflicts, providing strong support for decision-making.

[0023] Intelligent classification and metadata annotation significantly improve document processing efficiency. Through small-sample learning and multimodal feature fusion, it quickly adapts to new document types. Meanwhile, the interactive annotation assistance system reduces manual work. The permission control and security audit module establishes a dynamic, fine-grained security system, effectively preventing data leakage risks and ensuring the security and controllability of the entire document process.

[0024] In addition, the retrieval and recommendation module supports cross-modal retrieval and personalized recommendations, making information acquisition more convenient and efficient. The process automation module optimizes the approval process through intelligent routing and RPA technology, shortening processing time. The quality assessment and optimization module enables real-time monitoring and automatic evolution of system performance, ensuring long-term stable operation. The visual decision support module intuitively displays the status of document processing, assisting management in making scientific decisions and comprehensively enhancing the intelligence and business value of document management. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a schematic block diagram of an intelligent management system for automatic document recognition based on OCR proposed by the present invention;

[0026] Figure 2This is a comparison chart of recognition accuracy before and after document preprocessing;

[0027] Figure 3 This is a performance combination diagram of different OCR technologies in document recognition;

[0028] Figure 4 A line chart comparing the efficiency of semantic understanding and knowledge graph construction of different methods. DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0031] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.

[0032] Reference Figures 1 to 4 :An intelligent management system for automatic document recognition based on OCR, including the following modules:

[0033] The Multimodal Document Preprocessing Module utilizes a hybrid filtering algorithm with dynamic threshold adaptation. It first uses a convolutional neural network (CNN) to identify noise types in scanned documents, accurately distinguishing eight noise types, including scratches, stains, and noise. For scratch noise, an adaptive anisotropic diffusion filtering algorithm is used, dynamically adjusting the diffusion coefficient based on the scratch direction. This effectively removes scratch interference while preserving text edge information.

[0034] To address the challenge of flattening curled documents, the system integrates 3D reconstruction technology with physical models. Structured light scanning acquires point cloud data from the document surface, and a Poisson surface reconstruction algorithm is used to generate a 3D model. A physical model is then constructed based on the mechanical properties of paper, simulating the natural unfolding process of the paper and calculating the optimal unfolding position for each point.

[0035] For low-quality documents, the system uses a generative adversarial network (GAN) for quality enhancement and designs a document quality transfer network (DQTN). Through adversarial training, it learns the mapping relationship between high-quality and low-quality documents. It can super-resolution low-resolution documents of 100dpi to 600dpi while maintaining the clarity of text edges.

[0036] To address the fading problem of thermal paper documents, the module uses spectral reconstruction technology to analyze image differences at different wavelengths, reconstruct the original text outline, and restore the text information. Furthermore, a document separation algorithm has been developed. This algorithm uses deep learning to predict paper edges and combines it with a physical simulation algorithm to automatically separate documents without damaging them, supporting the batch digitization of historical archives.

[0037] Deep Fusion OCR Recognition Engine: Builds a hierarchical feature extraction network. The first layer utilizes Deformable Convolution, trained on a dataset containing 100,000 official document images. This allows it to learn the irregular shapes of text and adaptively adjust the convolution kernel sampling position. For text with an inclination angle of ≤45°, the recognition accuracy reaches 98.5%. The middle layer utilizes a bidirectional attention mechanism combined with a Transformer architecture. This self-attention mechanism calculates text sequence dependencies, effectively handling cross-paragraph citations such as "See Article X" and "According to the provisions of the preceding paragraph," improving the accuracy of document semantic understanding. The final layer features a dynamic dictionary matching module, building a dynamic dictionary containing 50,000 entries for official document terminology. This module matches vocabulary in real time during recognition, resolving ambiguous characters through dictionary constraints and language model scoring. The module is particularly effective in recognizing complex vocabulary such as legal texts and organizational names.

[0038] In order to improve the model reasoning speed, the formula is introduced ,in To optimize the inference speed, is the initial inference speed, To save computational effort through knowledge distillation and parameter sharing, is the total raw computational effort. Implementing a lightweight model compression solution based on this formula can triple the inference speed when deployed on edge devices.

[0039] In terms of seal penetration recognition, polarized light imaging is used to obtain multi-angle images of the document, and the seal and text information are separated based on the differences in the polarized light reflection characteristics of different materials. The phase correlation algorithm (PhaseCorrelation) is then used to calculate and compensate for the text deformation caused by seal coverage.

[0040] In addition, we designed an optimized solution for multilingual mixed text recognition, using a language detection module to switch recognition parameters in real time and combining character morphology analysis to eliminate cross-linguistic character interference. We also developed hybrid handwriting and printed text recognition technology, using a generative adversarial network to convert handwriting styles into standard fonts. This is then recognized through a unified network, enabling accurate interpretation of cursive and cursive scripts.

[0041] Semantic Understanding and Knowledge Graph Construction Module: This module proposes a cross-document semantic association analysis framework. First, it uses a self-attention mechanism to extract key semantic units from official documents, dividing long texts into sentences or paragraphs with independent semantics. Then, through temporal embedding technology, it incorporates temporal elements such as document release date and expiration date into the semantic representation, constructing a dynamic knowledge graph.

[0042] In the area of ​​event tracing, an event tracing algorithm was developed based on a graph neural network. This algorithm automatically generates a policy evolution diagram by analyzing reference relationships, semantic similarity, and chronological order between documents. To optimize knowledge graph completion, an adversarial training mechanism was introduced, building a generator-discriminator architecture. The generator attempts to generate knowledge graph relationships, while the discriminator verifies their authenticity, effectively identifying implicit relationships such as "policy-implementing department" and "problem-solution."

[0043] The system can analyze multiple official documents simultaneously, automatically identify policy conflicts and generate amendment suggestions. If two policy documents are found to have inconsistent provisions on the same matter, it will infer and propose amendments based on factors such as release time and effectiveness level to improve the efficiency of policy review. A policy effectiveness evaluation model is established, and the analytic hierarchy process (AHP) is used to determine the weight of each factor. The weight of release time is 0.3, the weight of effectiveness level is 0.5 (such as administrative regulations > departmental regulations > normative documents), and the weight of the formulating department is 0.2. For conflicting clauses, the following amendment priority formula is used: , where P is the revision priority, T is the release time score (the newer the better), L is the effectiveness level score, and D is the authority score of the formulating department; =0.3, β=0.5, γ=0.2. The clauses to be retained and revised are determined based on priority. A rule-based text editing method is used for revisions, matching conflicting content with regular expressions. Replacement or supplementation is performed based on relevant laws and superior policies to ensure that the revised policy document is logically consistent and compliant with regulations.

[0044] In addition, the module introduces a logical reasoning model, relying on the common sense knowledge base and the policy and regulation base to check the compliance of the official document content. It also introduces the knowledge graph to update the correlation formula , where U is the knowledge graph update correlation, which is used to measure the degree of knowledge graph update; is the number of newly added nodes, reflecting the addition of knowledge graph entities; is the total number of original nodes in the knowledge graph; is the number of newly generated relationships, reflecting the changes in knowledge graph relationships; is the total number of original relations; is the node weight coefficient (0≤ ≤1), used to adjust the weight of nodes and relationships in the updated relevance calculation. When new documents are released, a graph neural network automatically updates associated nodes and relationships based on this formula, driving the real-time evolution of policy knowledge and enabling rapid response to emerging business scenarios.

[0045] Classification and Metadata Annotation Module: A meta-learning-driven small-sample classifier is designed. Using the MAML (Model-Agnostic Meta-Learning) algorithm, an efficient classification model can be quickly trained using only a small amount of annotated data (typically 5-10). A multimodal feature fusion network is constructed to integrate document visual layout, text semantics, and formatting features. Faster R-CNN is used to detect paragraphs, headers, and footers, extracting visual layout features and analyzing paragraph indentation, font size variations, and other factors. The BERT-large model is used to extract text semantic features. Format features encompass information such as font type, color, and line spacing, significantly improving the classification of complex documents containing attachments and tables.

[0046] Introducing a formula to improve annotation efficiency ,in is the optimized annotation efficiency, is the initial annotation efficiency, is the amount of effective labeling habit knowledge accumulated through reinforcement learning, The total amount of annotation knowledge. We developed an interactive annotation assistance system that uses reinforcement learning to analyze users' historical annotation behavior, build a preference model, and automatically recommend candidate tags. We also built an active learning-driven annotation optimization system that automatically selects the most valuable documents to be annotated through uncertainty sampling and density estimation, optimizing the annotation process.

[0047] We designed a semantically enhanced metadata extractor that leverages external knowledge bases to expand its extraction scope. Using entity linking technology, we linked entities mentioned in official documents (such as company names and project names) with information in external knowledge bases to automatically extract more relevant metadata.

[0048] The Permission Control and Security Audit Module builds a dynamic access control system based on attribute-based encryption (ABE). This system uses user attributes (such as department, position, and security level), document attributes (such as confidentiality level, subject, and expiration date), and environmental attributes (such as access time and location) as encryption and decryption criteria. For example, "confidential" documents can only be decrypted and viewed by users who meet the following criteria: "belong to the security department," "rank above section chief," and "access during working hours." This system supports permission management for millions of users, with a permission change response time of ≤ 200ms. Incorporating blockchain technology, it enables distributed storage and tamper-proof verification of operation logs, ensuring the authenticity and reliability of operation records. Blockchain technology offers decentralization, distributed ledgers, and cryptographic encryption. In this module, operation logs are packaged into data blocks, each containing operation records within a specific timeframe. These blocks are linked using a cryptographic hash algorithm to form an immutable chain structure. Each participating node (which can be a server, client, or other device) maintains a complete copy of the operation log. When a new operation record is generated, it is broadcast to the entire network. Each node verifies and confirms the validity of the operation record through a consensus algorithm (such as proof-of-work, proof-of-stake, etc. This system adopts a specific consensus algorithm that suits its own security and performance requirements) and adds it to the blockchain ledger. Once the operation record is recorded on the blockchain, it cannot be tampered with. Any modification to a single data block will cause the hash value of all subsequent data blocks to change. This change is easily detected by other nodes, thus ensuring the authenticity and reliability of the operation log and providing a solid data foundation for security audits.

[0049] Design a user behavior profiling system, use the long short-term memory network (LSTM) to analyze user operation sequences, record user viewing, downloading, forwarding and other operations, and build a behavior sequence model to learn normal behavior patterns. Introduce the abnormal behavior detection formula , where R represents recall, a measure of the ability to detect anomalous behavior; TP is the true positive rate (i.e., the number of anomalous behaviors correctly detected); and FN is the false negative rate (i.e., the number of actual anomalous behaviors not detected). This formula is used to optimize detection mechanisms and develop privacy-preserving data enhancement solutions. Using differential privacy techniques, we expand audit samples without leaking sensitive information, enabling real-time and accurate detection of anomalous behaviors such as frequent downloading of sensitive documents.

[0050] A cross-departmental security collaboration solution based on federated learning was developed. Security audit systems in different departments train models locally, share model updates through parameter exchange, and jointly train anomaly detection models without sharing raw data. A dynamic permission adjustment strategy was designed. The system assesses the risk level of user operations based on multiple dimensions, such as access frequency, data sensitivity, and operation time, and calculates a risk score. When the score exceeds a threshold, user permissions are automatically reduced or secondary verification is triggered. In actual testing, this strategy has reduced security incident response time from an average of 5 minutes using traditional methods to less than 3 seconds, effectively preventing data breaches.

[0051] The present invention also includes a retrieval and recommendation module, which utilizes a Transformer-based cross-modal retrieval model and the CLIP model to extract cross-modal features. This module maps text, image, and voice queries into a unified semantic space, supporting hybrid retrieval. For example, users can search for official documents by describing image content (e.g., "a document with a red seal") or by voice commands (e.g., "search for recently released environmental protection policies").

[0052] Introducing the recommendation hit rate formula , where H is the recommendation hit rate, which measures the recommendation accuracy; is the document semantic similarity, is the user behavior weight, is the scene context matching degree, is the scene weight, is the maximum similarity value, is the total weight value. Based on this, a personalized recommendation engine is built, combining user historical behavior with document semantic similarity to automatically recommend relevant policy documents and past cases in scenarios such as project approval.

[0053] To improve retrieval efficiency, a multi-level indexing structure was designed. The first level uses an inverted index to quickly locate relevant documents, the second level utilizes semantic vector indexing for precise ranking, and the third level re-ranks documents based on popularity and timeliness. When processing millions of official documents, the average search response time is kept under 500ms. This multi-dimensional ranking algorithm comprehensively considers factors such as document relevance, timeliness, and authority. Ranking weights are trained using a machine learning model and dynamically adjusted based on the scenario.

[0054] The present invention also includes a process automation module, which uses a reinforcement learning-based intelligent routing algorithm to optimize the document approval path. In practice, the intelligent routing model is trained using the Deep Deterministic Policy Gradient (DDPG) algorithm, taking factors such as document urgency, human workload, and historical processing time as state inputs to generate the optimal approval path.

[0055] The RPA robotic process orchestration tool developed supports low-code custom automation tasks. Through a graphical interface, users can drag and drop components to configure automated processes, such as automatically extracting key information from official documents and populating it into other system forms, or automatically sending approval reminder emails. For example, a contract approval form that previously took two hours to complete manually can now be completed by an RPA robot in just 10 minutes.

[0056] A dynamic process adjustment mechanism automatically optimizes approval nodes based on real-time conditions. When a delay is detected in an approval link, the system automatically assesses whether subsequent processes require adjustment. This assessment first analyzes the nature and importance of the delayed link to determine whether it is a critical path node. If it is a critical node, it further considers factors such as the duration of the delay, its impact on the overall approval process, and the potential for adjustment in subsequent steps. For example, if the delay is long and significantly impacts the overall process, and there are subsequent steps that can be processed in parallel without compromising the accuracy of the approval results, the system will automatically trigger the addition of a parallel approval node. When adding parallel nodes, tasks are appropriately assigned to appropriate individuals based on factors such as their workload and expertise. If the assessment indicates that a delay in a non-critical step could trigger a chain reaction that significantly delays the overall process, the system will determine whether the step can be skipped, ensuring that skipping the step does not violate relevant approval rules and business logic. This allows for dynamic optimization of the approval process while ensuring regulatory compliance, ensuring efficient and smooth document approval.

[0057] The present invention also includes a quality assessment and optimization module, which establishes a multi-dimensional quality monitoring system to comprehensively evaluate system performance. In actual implementation, real-time evaluation is conducted across 12 dimensions, including recognition accuracy, semantic understanding correctness, and format compliance. For example, by comparing OCR recognition results with manually annotated standard text, metrics such as character error rate and word error rate are calculated; semantic understanding correctness is assessed by experts; and document format compliance is checked using pre-set formatting rules.

[0058] The developed model automatic evolution framework utilizes a genetic algorithm to optimize model structure and parameters. The system regularly collects user feedback data, generates candidate models using the genetic algorithm, evaluates performance on a validation set, and selects the optimal model for update. The active learning module automatically selects the most valuable documents for annotation through uncertainty sampling and density estimation. The system calculates the prediction uncertainty of each unlabeled example, taking into account the density distribution of the example in the feature space, and selects both uncertain and representative examples for annotation.

[0059] The interactive feedback mechanism allows users to correct and evaluate system output. The system records every user feedback, analyzes the error types and distribution, and optimizes the model accordingly. For example, if users frequently correct recognition errors in a certain type of professional terminology, the system will automatically adjust the dictionary and language model.

[0060] This invention also includes a visual decision support module, which develops a spatiotemporal information visualization platform to achieve a multi-dimensional display of document processing. In practical applications, this module uses a geographic information system (GIS) to display the regional distribution of document processing progress, using different colors and icons to indicate processing status (e.g., completed, in progress, delayed). Combined with a timeline function, users can view the evolution of policy implementation effects over time.

[0061] The designed 3D knowledge graph interface allows users to explore document relationships through gestures and voice. The system displays document entities (such as departments, policies, and events) and relationships (such as publication, execution, and citation) in a 3D graphical format. Users can use gestures like dragging, zooming, and other operations to delve deeper into hidden relationships.

[0062] The predictive analysis function uses a time series prediction model (LSTM) to predict document processing times. The system analyzes historical processing data, learns the relationship between processing time and factors such as document type, complexity, and the person handling it, and generates a predictive model. The automatic report generation module generates standardized analysis reports based on user needs. The system supports multiple report formats (PDF and Excel), and report content and layout can be customized. For example, users can choose to generate a "quarterly document processing analysis report" that includes processing volume statistics, timeliness analysis, and quality assessment.

[0063] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An intelligent management system for automatic document recognition based on OCR, characterized by: Includes the following modules: Multimodal document preprocessing module: This module uses a dynamic threshold adaptive hybrid filtering algorithm, a convolutional neural network to predict noise distribution and generate compensation masks, three-dimensional reconstruction technology to unfold curled documents, a generative adversarial network to enhance document resolution, and a physical model to inversely optimize preprocessing parameters. Deep Fusion OCR Recognition Engine: Builds a hierarchical feature extraction network and introduces a formula to improve model inference speed; develops seal penetration recognition technology, separates seals and text based on polarized light imaging, and uses a phase correlation algorithm to compensate for text deformation; Semantic Understanding and Knowledge Graph Construction Module: Build a cross-document semantic association framework, use the self-attention mechanism to extract key semantic units, generate a dynamic knowledge graph using temporal embedding technology, develop an event tracing algorithm, generate a policy evolution graph based on citation relationships and semantic similarity, and introduce adversarial training to optimize graph completion; Classification and metadata annotation module: Build a multimodal feature fusion network, integrate document visual layout, text semantics, and format features, and introduce a formula to improve annotation efficiency. Develop an interactive annotation assistance system that uses reinforcement learning to understand user annotation habits and automatically recommend candidate tags. Permission control and security audit module: Build a dynamic access control system based on attribute-based encryption, combine blockchain technology, design a user behavior profiling system, and analyze operation sequences through long-short-term memory networks; The deep fusion OCR recognition engine also includes: designing an optimization scheme for multilingual mixed text recognition, using a language detection module to switch recognition parameters in real time, combining character morphology analysis to eliminate character interference between languages, developing a hybrid handwriting and printed text recognition technology, using a generative adversarial network to convert handwriting styles into standard fonts, and then using a unified network for recognition; The formula for improving the model inference speed in the deep fusion OCR recognition engine is: , To optimize the inference speed, is the initial inference speed, To save computational effort, is the original total computational effort; The classification and metadata annotation module also includes: building an active learning-driven annotation optimization system, automatically screening documents to be annotated through uncertainty sampling and density estimation, designing a semantically enhanced metadata extractor, and using an external knowledge base to expand the extraction scope; The formula for improving the annotation efficiency in the classification and metadata annotation module is: , is the optimized annotation efficiency, is the initial annotation efficiency, is the amount of knowledge of effective labeling habits, It is the total amount of labeled knowledge.

2. The intelligent management system for automatic document recognition based on OCR according to claim 1, characterized in that: The multimodal document preprocessing module also includes: addressing the fading problem of thermal paper documents, using spectral reconstruction technology to restore text information, reconstructing the original text outline by analyzing image differences under different wavelengths, developing a document adhesion separation algorithm, using deep learning to predict paper edges, and combining physical simulation algorithms to achieve automatic separation without damaging the document.

3. The intelligent management system for automatic document recognition based on OCR according to claim 1, characterized in that: The semantic understanding and knowledge graph construction module also includes: introducing a logical reasoning model, relying on the common sense knowledge base and the policy and regulation base to conduct compliance checks on the content of official documents, and introducing the knowledge graph update relevance formula: , where U represents the updated relevance of the knowledge graph, is the number of newly added nodes, is the total number of original nodes in the knowledge graph, is the number of newly generated relations, is the total number of original relations, It is the node weight coefficient. A dynamic knowledge graph update mechanism is developed. When a new document is released, the graph neural network is used to automatically update the associated nodes and relationships based on this formula.

4. The intelligent management system for automatic document recognition based on OCR according to claim 1, characterized in that: The permission control and security audit module also includes: developing a cross-departmental security collaboration solution based on federated learning, jointly training anomaly detection models without sharing original data, designing dynamic permission adjustment strategies, raising and lowering permissions in real time based on user operation risk scores, and automatically triggering secondary verification when high-risk behavior is detected, shortening the security incident response time to less than 3 seconds.

5. The intelligent management system for automatic document recognition based on OCR according to claim 1 is characterized in that: Also includes: The retrieval and recommendation module uses a Transformer-based cross-modal retrieval model, supports mixed retrieval of text, images, and voice, and introduces the recommendation hit rate formula: , where H is the recommendation hit rate, is the document semantic similarity, is the user behavior weight, is the scene context matching degree, is the scene weight, is the maximum similarity value, For the total weight value, the module builds a personalized recommendation engine, which combines user historical behavior and document semantic similarity to automatically recommend relevant policy documents and past cases in scenarios such as project approval.

6. The OCR-based document automatic recognition intelligent management system according to claim 1, characterized in that: Also includes: The process automation module designs an intelligent routing algorithm based on reinforcement learning, dynamically plans the approval path based on factors such as the urgency of the document and the load of the person handling it, develops an RPA robot process orchestration tool, and supports low-code custom automation tasks.

7. The intelligent management system for automatic document recognition based on OCR according to claim 1, characterized in that: It also includes: quality assessment and optimization module, building a multi-dimensional quality monitoring system, conducting real-time assessment from 12 dimensions, developing a model automatic evolution framework, optimizing model structure and parameters through genetic algorithms, and conducting continuous training based on user feedback data.

8. The intelligent management system for automatic document recognition based on OCR according to claim 1, characterized in that: Also includes: A visual decision support module develops a spatiotemporal information visualization platform, displays the regional distribution of document processing progress through a geographic information system, presents the evolution of policy implementation effects in combination with a timeline, designs a three-dimensional knowledge graph interactive interface, supports users in exploring document correlations through gestures and voice, and generates visual analysis reports.

Citation Information

Patent Citations

  • High-speed dynamic QR code identification method based on improved SURF composite algorithm

    CN110210584A

  • Bill identification method and system for RPA office process automation system

    CN118072337A