Intelligent government affair service method and terminal based on OCR identification technology
Through degradation identification, multidimensional information set analysis and logical correlation analysis, and mark-based error correction combined with government process information, the problem that existing OCR technology is difficult to deal with complex government documents is solved, and the efficiency and automation of electronic processing of government documents has been improved.
Patent Information
- Application Number
- CN202510475193.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing OCR technology is difficult to effectively process government service documents with high input complexity and strong business logic coupling in government service systems, resulting in frequent manual intervention in government service processes, and limited auxiliary roles.
By obtaining the document image data set, degradation identification processing is performed to determine the degradation area information, analyzing the multi-dimensional information set of document content, logical correlation analysis is performed in combination with government process information, and mark-based error correction is performed, and document processing reports are output to improve the electronic processing capabilities of government documents.
It has improved the comprehensive processing capability of the electronic processing process of government documents, reduced the frequency of intervention of relevant staff, improved the efficiency of government services, and promoted the digitalization process of government services.
Smart Images

Figure CN119992568A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of e-government technology, and in particular to an intelligent government service method and terminal based on OCR recognition technology. Background Art
[0002] With the comprehensive and in-depth development of information technology, the pace of digital transformation of government services is accelerating. OCR (Optical Character Recognition) technology has been widely used in the digitization of paper documents of government services, effectively improving the processing efficiency of government affairs and promoting the informatization development of government services.
[0003] However, when existing OCR technology is applied to government service systems, it can usually only achieve the conversion between images and text. It lacks the comprehensive processing capabilities for documents in the government service field that have high input complexity and strong business logic coupling. As a result, the government service process still requires frequent intervention from relevant staff, and its auxiliary role in the efficient processing of government services is very limited. Summary of the invention
[0004] The present application provides an intelligent government service method and terminal based on OCR recognition technology to solve the above technical problems.
[0005] In a first aspect, the present application provides an intelligent government service method based on OCR recognition technology, the method comprising: obtaining a document image data set, performing degradation recognition processing on each document image in the document image data set, and determining degradation area information; based on the degradation area information, analyzing the document image data set, and determining a multidimensional information set of document content; obtaining government process information, and based on the government process information, performing logical association analysis on the multidimensional information set of document content, and determining an association analysis result; based on the association analysis result, performing marked error correction on the multidimensional information set of document content, and determining and outputting a document processing report.
[0006] Through this solution, the degraded areas in the document image are identified and processed to obtain the degraded area information. On this basis, the document content multidimensional information set containing each document element and the corresponding complementary content is analyzed and obtained. Combined with the government process information, the document content multidimensional information set is logically associated and analyzed. According to the obtained visit analysis results, the document content multidimensional information set is marked and corrected, and the corresponding document processing report is provided to the corresponding government service personnel for their rapid approval and modification. The comprehensive processing capability of the electronic processing process of government documents based on OCR technology for documents in the government service field with high input complexity and strong business logic coupling is improved, the frequency of intervention of relevant staff in the government service process is reduced, the efficiency of government services is improved, and the digitalization of government services is promoted.
[0007] Optionally, performing degradation recognition processing on each document image in the document image data set to determine degradation area information includes: analyzing the document image data set to extract a number of feature degradation areas of each document image; based on an image convolutional neural network model, extracting a number of image feature vectors corresponding to the feature degradation areas, and performing feature degradation type probability distribution analysis on each of the image feature vectors to determine the degradation distribution probability of each feature degradation area under each feature degradation type; the feature degradation types include creases, fades, stains and missing; based on a preset degradation classification minimization loss function, determining the feature degradation type corresponding to each feature degradation area according to the degradation distribution probability of each feature degradation area under each feature degradation type; and constructing the degradation area information according to each feature degradation area and its corresponding feature degradation type.
[0008] Through this solution, the multi-scale feature fusion analysis process is utilized to accurately analyze the feature degradation types corresponding to different feature degradation areas, effectively solve the confusion problem of similar degradation features, and significantly improve the accuracy of degradation type recognition in complex degradation scenarios. It provides scientific data basis for the subsequent targeted completion analysis of fuzzy document content under different feature degradation types, and reduces the risk of error propagation in the subsequent document completion process.
[0009] Optionally, the preset degradation classification minimization loss function is specifically the following formula: ; in, is the cross entropy loss, is the degenerate type index, is the total number of degenerate types, For the The historical frequency of degenerate types, For the A vector of preset type labels for the degenerate types, For the The degradation distribution probability of each degradation type.
[0010] Through this scheme, based on mathematical analysis methods, a preset degradation classification minimization loss function is constructed, and the difference between the degradation distribution probability of the current degraded area under different degradation types analyzed by the image convolutional neural network model and the distribution probability of the historical degradation characteristics of the document is measured. The cross-entropy loss minimization optimization process is used to improve the accuracy of the degradation distribution probability, thereby improving the accuracy of the degradation type analysis process to which the degraded area belongs.
[0011] Optionally, the document image data set is analyzed based on the degraded area information to determine a multidimensional information set of document content, including: analyzing the document image data set, identifying and determining the document layout structure area, the text area and the additional mark area in each document image in the document image data set, and extracting the document content in each document area; based on the degraded area information, determining a number of cross-fuzzy areas and their corresponding document areas according to the intersection area between the feature degraded area and the document layout structure area / the text area / the additional mark area; based on the feature degradation type and the document area to which the cross-fuzzy area belongs, performing inference completion analysis on the cross-fuzzy areas according to the document content corresponding to the document area, and determining the information to be completed corresponding to each cross-fuzzy area; constructing the document content multidimensional information set according to the document content in each document area and the information to be completed in each cross-fuzzy area.
[0012] Through this scheme, the document layout structure area, main text area and additional markup area in the document are accurately divided, laying the foundation for the completion analysis process and improving the document area sensitivity of the completion analysis process. By targeting the document areas to which different feature degradation types and cross-fuzzy areas belong, and according to the document content corresponding to the document areas, targeted reasoning completion analysis strategies are adopted for different cross-fuzzy areas to obtain the information to be completed corresponding to each cross-fuzzy area, thereby improving the regional matching degree and accuracy of the information to be completed.
[0013] Optionally, the document area to which the feature degradation type and the cross-fuzzy area belong is subjected to an inference-based completion analysis according to the document content corresponding to the document area, and the corresponding information to be completed in each cross-fuzzy area is determined, including: if the document area is the document layout structure area, the overall layout structure outline of the document is extracted according to the document content corresponding to the document layout structure area; according to the overall layout structure outline of the document, a preset document template database is retrieved to determine whether there is a corresponding document template; if the document template exists, the differentiated contour edge between the overall layout structure outline of the document and the corresponding document template is used as the information to be completed; if the document template does not exist, based on the graph neural network, the overall layout structure outline of the document is subjected to a line continuity completion analysis according to the feature degradation type, the completion structure lines of the cross-fuzzy area are determined, and the completion structure lines are used as the information to be completed.
[0014] Through this solution, the overall layout structure contour of the document corresponding to the document layout structure area is extracted, and this is used as a retrieval basis to search the preset document template database. If the document template exists, the differentiated contour edge between the overall layout structure contour of the document and the corresponding document template is used as the information to be completed, thereby reducing the occupation of computing resources. If the document template does not exist, the overall layout structure contour of the document is analyzed for line continuity completion based on the feature degradation type, and the corresponding completed structure lines are used as the information to be completed, thereby improving the matching degree between the completed content and the non-standard document, and improving the flexibility and accuracy of the layout structure completion analysis process.
[0015] Optionally, the document area to which the cross-fuzzy area belongs based on the feature degradation type and the document area corresponding to the cross-fuzzy area is subjected to an inference-based completion analysis according to the document content corresponding to the cross-fuzzy area, and the information to be completed corresponding to each cross-fuzzy area is determined, including: if the document area to which the cross-fuzzy area belongs is the main text area, edge feature analysis is performed on the text information in the main text area according to an image edge detection algorithm to determine printed information and handwritten information; based on an image edge detection algorithm, the handwritten information is analyzed to determine a handwriting feature set; based on a preset feature point descriptor, the handwriting feature set is analyzed to determine a handwriting feature vector; based on the handwriting feature vector, a handwriting analogy analysis is performed on the handwriting information to determine the semantic information of the handwritten text; based on the natural language analysis algorithm, contextual semantic continuity analysis is performed on the printed information and the handwritten text semantic information to determine semantic continuity guarantee completion information, and the semantic continuity guarantee completion information is used as the information to be completed.
[0016] Through this solution, the printed information and handwritten information in the main text area are separated, and the semantic information of the handwritten text is determined by analyzing and extracting the note features of the handwritten text. Based on natural language analysis technology, contextual semantic continuity analysis is performed on the printed information and the handwritten text semantic information to determine the semantic continuity guarantee completion information, and the semantic continuity guarantee completion information is used as the information to be completed, so as to avoid error propagation caused by mixed processing of handwritten and printed text, and improve the adaptability of the completion analysis process to diverse documents.
[0017] Optionally, the document area to which the cross-fuzzy area belongs based on the feature degradation type and the cross-fuzzy area is subjected to an inference-based completion analysis according to the document content corresponding to the cross-fuzzy area, and the corresponding information to be completed in each cross-fuzzy area is determined, including: if the document area is the additional mark area, according to the document content corresponding to the additional mark area, dividing and extracting the background semi-transparent watermark content and the foreground dark seal content; analyzing the background semi-transparent watermark content, extracting the watermark unit content, and determining the unit to which the document belongs; analyzing the foreground dark seal content, and determining the document approval unit; sending the document verification information to the corresponding unit according to the document unit and the document approval unit, and receiving the verification feedback information and the target seal image provided by the corresponding unit; if the verification feedback information is that the verification is correct, taking the difference contour between the watermark unit content and the cross-fuzzy area as the information to be completed, and taking the difference contour between the target seal image and the foreground dark seal content as the information to be completed.
[0018] Through this scheme, the watermark content and the seal content in the additional mark area are separated to clarify the unit to which the current document belongs and the approval unit, and the document verification information is sent to the corresponding unit. After determining that the verification feedback information is correct, the difference contour between the watermark unit content and the cross-fuzzy area is used as the information to be completed. At the same time, the difference contour between the target seal image and the foreground dark seal content is used as the information to be completed. While avoiding document tampering and forgery, accurate analysis of the watermark and seal content completion information is achieved.
[0019] Optionally, the government affairs process information includes a government affairs process directed graph, a set of government affairs rule logical expressions, government affairs timeliness information and a government affairs responsibility matrix. Based on the government affairs process information, a logical association analysis is performed on the multidimensional information set of the document content to determine the association analysis result, including: based on the multidimensional information set of the document content, according to the information to be completed in each cross-fuzzy area, the document content in each document area is completed to determine the completed document content; based on a natural language analysis algorithm, the completed document content is analyzed to determine the document government affairs process, document government affairs rules, document government affairs timeliness information and document government affairs responsibility Subject; according to the government affairs process directed graph, perform a topological matching evaluation on the document government affairs process to determine the topological matching degree of the process nodes; according to the set of logical expressions of the government affairs rules, perform rule compliance verification on the document government affairs rules to determine the rule matching degree; according to the government affairs timeliness information, perform timeliness verification on the document government affairs timeliness information to determine the timing matching degree; according to the government affairs responsibility matrix, perform a cosine similarity evaluation on the document government affairs responsible subject to determine the responsibility matching degree; according to the process node topological matching degree, the rule matching degree, the timing matching degree and the responsibility matching degree, construct the association analysis result.
[0020] Through this solution, based on government process information, starting from the four dimensions of process, rules, timeliness and responsibility, a logical correlation analysis is performed on the multidimensional information set of document content. Through the four quantitative indicators of process node topology matching, rule matching, timing matching and responsibility matching, the matching degree between the multidimensional information set of document content and the corresponding government process information in different dimensions is mapped respectively. This is used as the correlation analysis result to improve the scientificity and comprehensiveness of the logical correlation analysis process.
[0021] Optionally, based on the association analysis result, the multidimensional information set of the document content is marked and corrected, including: comparing the process node topology matching degree, the rule matching degree, the timing matching degree and the responsibility matching degree in the association analysis result with the corresponding matching degree ranges respectively; if there is a situation where the process node topology matching degree / the rule matching degree / the timing matching degree / the responsibility matching degree are not within the corresponding matching degree range, extracting the corresponding abnormal matching degree; based on the color-differentiated document annotation strategy and the government process information, according to the abnormal type corresponding to the abnormal matching degree, highlighting the local content of the document corresponding to the abnormal matching degree, and determining the corresponding adjustment suggestion.
[0022] Through this solution, accurate positioning and efficient correction of government document errors are achieved through matching comparison and visual annotation. The color differentiation strategy reduces the complexity of manual review, and the highlighted annotation directly points to the problem area. Combined with targeted adjustment suggestions, the standardization and efficiency of government processing procedures are significantly improved. At the same time, through automated anomaly detection and prompts, compliance risks caused by omissions or misjudgments are reduced, ensuring that the document content is highly consistent with government process requirements.
[0023] In the second aspect, the present application provides an intelligent government service terminal based on OCR recognition technology, and the terminal includes: a degradation analysis module, which is used to obtain a document image data set, perform degradation recognition processing on each document image in the document image data set, and determine the degradation area information; a document analysis module, which is used to analyze the document image data set based on the degradation area information, and determine the document content multidimensional information set; an association analysis module, which is used to obtain government process information, and perform logical association analysis on the document content multidimensional information set based on the government process information, and determine the association analysis result; a document marking module, which is used to perform marking-type error correction on the document content multidimensional information set according to the association analysis result, and determine and output a document processing report. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0025] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application; Figure 2 A flowchart of an intelligent government service method based on OCR recognition technology provided in one embodiment of the present application; Figure 3 A schematic diagram of the structure of an intelligent government service terminal based on OCR recognition technology provided in one embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.
[0028] The embodiments of the present application are further described in detail below in conjunction with the drawings in the specification.
[0029] Based on this, the present application provides an intelligent government service method and terminal based on OCR recognition technology. The degraded area in the document image is identified and processed to obtain the degraded area information. On this basis, a document content multidimensional information set containing each document element and the corresponding complementary content is analyzed and obtained. In combination with the government process information, a logical association analysis is performed on the document content multidimensional information set. According to the obtained visit analysis results, the document content multidimensional information set is marked and corrected, and the corresponding document processing report is provided to the corresponding government service personnel for their rapid approval and modification, thereby improving the comprehensive processing capability of the electronic processing process of government documents based on OCR technology for documents in the government service field with high input complexity and strong business logic coupling, reducing the frequency of intervention of relevant staff in the government service process, improving the efficiency of government services, and promoting the digitalization of government services.
[0030] Figure 1 A schematic diagram of an application scenario provided by this application. In the process of electronic government documents, the method provided by this application is applied to comprehensively process government service documents with high input complexity and strong business logic coupling, thereby improving government service efficiency and reducing the involvement of relevant staff.
[0031] Specifically, the method of the present application is applied to any server, which communicates with a document scanning device, obtains and analyzes a document image data set through the server, identifies and processes the degraded area in the document image, and obtains the degraded area information. On this basis, a multidimensional information set of document content containing each document element and the corresponding complementary content is analyzed and obtained. Combined with the government process information, a logical association analysis is performed on the multidimensional information set of document content. Based on the obtained visit analysis results, the multidimensional information set of document content is marked and corrected, and the corresponding document processing report is provided to the corresponding government service personnel for their rapid approval and modification, thereby improving the comprehensive processing capability of the electronic processing process of government documents based on OCR technology for documents in the government service field with high input complexity and strong business logic coupling, reducing the frequency of intervention of relevant personnel in the government service process, improving the efficiency of government services, and promoting the digitalization of government services.
[0032] For specific implementation methods, please refer to the following embodiments.
[0033] Figure 2 This is a flowchart of an intelligent government service method based on OCR recognition technology provided in an embodiment of the present application. The method of this embodiment can be applied to the server in the above scenario. Figure 2 As shown, the method includes: S201, obtaining a document image data set, performing degradation recognition processing on each document image in the document image data set, and determining degradation region information.
[0034] The document image dataset may be a collection of images corresponding to various document files involved in the government service processing process, and the document image dataset may be collected by a document scanning device. The degradation recognition process may be a recognition process for degradation features that may exist in the document and cause local document information to be blurred, such as folds, stains, etc. The degradation area information may be the area information in the document that is affected by the degradation features.
[0035] Specifically, in the process of government affairs processing, paper documents are affected by factors such as storage time and the owner's storage method, and the clarity of their document content is highly uncertain. In addition, in the process of government affairs processing, affected by historical issues, the document digitization process often needs to face a large number of old paper documents, such as archival documents. When the existing OCR technology processes old documents in the government affairs process, since the content of government documents involves complex and numerous document elements, such as form structure, seals, watermarks, etc., and the old documents are affected by storage conditions and time, their document content has different degrees of degradation. The existing OCR technology is difficult to ensure the accuracy and completeness of the content of old documents in the process of document digitization. Through mathematical analysis, degradation identification processing is performed on different document images, and the areas affected by degradation in different documents and their corresponding degradation types are determined, providing a scientific data basis for the subsequent analysis and completion of document content in the degraded area.
[0036] S202: Analyze the document image data set based on the degraded region information to determine a multi-dimensional information set of the document content.
[0037] The document content multidimensional information set may be an information set containing different types of components of a document, such as document structure, printed parts, handwritten parts, and the like.
[0038] Specifically, government documents are composed of a large number of different types of document elements. Different document elements require different document content completion methods after being affected by degraded areas. For example, after the document structure is affected by the degraded area, since the document structure is mainly composed of different continuous wireframes and does not involve specific text content, the completion of the document structure is mainly achieved by the completion analysis of the continuity of the structural lines. If the body content of the document is affected by the degraded area, it is necessary to use natural language analysis technology, combined with the contextual information around the degraded area, to infer the text information in the degraded area to achieve the completion of the text content. Therefore, based on the degraded area information, the document content in different document images in the document image dataset is structurally divided to determine the different dimensional document elements affected by the degraded area in the document and the corresponding completion information of the document elements, thereby forming a multidimensional information set of document content as a targeted basis in the document completion analysis process.
[0039] S203: Obtain government affairs process information, and based on the government affairs process information, perform logical correlation analysis on the multi-dimensional information set of the document content to determine the correlation analysis result.
[0040] The government affairs process information may be information used to characterize the relationship between nodes in the current government affairs process.
[0041] Logical association analysis can be an analysis process of whether there is a logical conflict between the document content in the current document content multidimensional information set and the government affairs process based on the corresponding government affairs process.
[0042] The association analysis result may be information including different types of logical conflict analysis results.
[0043] Specifically, although the content of old documents can be guaranteed to be smooth in terms of textual semantics after analysis and completion, it does not mean that the completed document content is logically completely matched with the corresponding government process. For example, the registration certificate is very old and there is no corresponding electronic file. After scanning with the existing OCR technology, some fields are blurred. After completion through the above steps, the accuracy of some basic general information can be guaranteed, but some approval process descriptions and division of responsibilities of approval agencies involved in the completed document may have logical conflicts with the government process corresponding to the document, resulting in conflicts in document responsible entities. Therefore, after completing the document, it is necessary to refer to the corresponding government process and conduct logical correlation analysis on the multi-dimensional information set of the document content to avoid logical conflicts in government processes.
[0044] S204: Based on the association analysis results, mark-based error correction is performed on the multi-dimensional information set of the document content, and a document processing report is determined and output.
[0045] Marking error correction can be the process of marking the part of the multidimensional information set of the document content that has the logical conflict problem of the government process and giving error correction suggestions. The document processing report can be a report information containing a series of corresponding change data of the document content in the current government document processing process.
[0046] Specifically, through the results of correlation analysis, the logical conflict parts of the government affairs process in the completed document content are located, and according to the corresponding government affairs process information, the logical conflict parts are marked, and corresponding error correction suggestions are given. Through data visualization technology, the corresponding document processing report is constructed, and the document processing report is provided to the corresponding government service personnel for their quick approval and modification.
[0047] Through this solution, the degraded areas in the document image are identified and processed to obtain the degraded area information. On this basis, the document content multidimensional information set containing each document element and the corresponding complementary content is analyzed and obtained. Combined with the government process information, the document content multidimensional information set is logically associated and analyzed. According to the obtained visit analysis results, the document content multidimensional information set is marked and corrected, and the corresponding document processing report is provided to the corresponding government service personnel for their rapid approval and modification. The comprehensive processing capability of the electronic processing process of government documents based on OCR technology for documents in the government service field with high input complexity and strong business logic coupling is improved, the frequency of intervention of relevant staff in the government service process is reduced, the efficiency of government services is improved, and the digitalization of government services is promoted.
[0048] In some embodiments, a document image data set is analyzed to extract several feature degradation regions of each document image; based on an image convolutional neural network model, image feature vectors corresponding to the several feature degradation regions are extracted, and a feature degradation type probability distribution analysis is performed on each image feature vector to determine the degradation distribution probability of each feature degradation region under each feature degradation type; feature degradation types include creases, fades, stains, and missing; based on a preset degradation classification minimization loss function, the feature degradation type corresponding to each feature degradation region is determined according to the degradation distribution probability of each feature degradation region under each feature degradation type; and degradation region information is constructed according to each feature degradation region and its corresponding feature degradation type.
[0049] A feature degradation region may be a local region in a document image where visual quality has degraded, which may be manifested as a set of continuous pixels with abnormal pixel values, broken textures, or missing content. The feature degradation region may be divided into regions of the document image through an image segmentation algorithm (such as a threshold segmentation method based on edge detection), and the region with poor visual quality may be selected as a feature degradation region.
[0050] The image convolutional neural network model can be a feature extraction network improved by using the ResNet-34 architecture, whose input layer receives the image block of the degraded area whose size is normalized to the corresponding pixels (such as 224×224 pixels), and after extracting the spatial features through several convolution modules (such as 5 convolution modules), the image feature vector of a specified dimension (such as 1024 dimensions) is output through the global average pooling layer. The image feature vector can be the vector information used to characterize the comprehensive image features in the feature degradation area.
[0051] The feature degradation type probability distribution analysis may be a process of analyzing the probability corresponding to the current feature degradation region under each degradation type. The degradation distribution probability may be a probability distribution value of the current feature degradation region under each degradation type.
[0052] The preset degradation classification minimization loss function may be a mathematical function for measuring the difference between the evaluation of the feature degradation type in the degradation region by the image convolutional neural network model and the actual situation. The feature degradation type may be the degradation cause that causes the document data to be blurred in the current feature degradation region.
[0053] Specifically, in the process of digitizing paper documents, especially for old documents, due to differences in document preservation status, there is a common phenomenon of mixed existence of multiple degradation types in the image. The traditional single threshold detection method is difficult to accurately distinguish similar degradation features such as creases and stains, fading and missing. The image convolutional neural network model is used to quantitatively analyze the probability distribution value of each image feature vector under different degradation types. The preset degradation classification is combined to minimize the loss function to minimize the difference between the evaluation result and the actual situation, thereby determining the feature degradation type corresponding to each feature degradation area, and fusing multi-scale features to enhance the ability to distinguish the feature degradation type corresponding to the feature degradation area, providing scientific data basis for the subsequent targeted completion analysis of fuzzy document content under different feature degradation types.
[0054] Through this solution, the multi-scale feature fusion analysis process is utilized to accurately analyze the feature degradation types corresponding to different feature degradation areas, effectively solve the confusion problem of similar degradation features, and significantly improve the accuracy of degradation type recognition in complex degradation scenarios. It provides scientific data basis for the subsequent targeted completion analysis of fuzzy document content under different feature degradation types, and reduces the risk of error propagation in the subsequent document completion process.
[0055] In some embodiments, the degradation classification minimization loss function is preset, specifically the following formula (1): ; in, is the cross entropy loss, is the degenerate type index, is the total number of degenerate types, For the The historical frequency of degenerate types, For the A vector of preset type labels for the degenerate types, For the The degradation distribution probability of each degradation type.
[0056] The cross entropy loss can be a quantitative value used to characterize the difference between the model's predicted probability distribution and the true distribution.
[0057] The historical occurrence frequency may be the occurrence frequency of various degradation phenomena obtained based on the statistics of the government document database. The historical occurrence frequency relationship of each degradation type in the government document generally follows the following rule: fade > crease > stain > missing. The preset type label vector may be vector information used to refer to the degradation type. The preset type label vector may be represented by One-hot encoding, such as crease corresponding to [1,0,0,0], fade corresponding to [0,1,0,0], etc.
[0058] Specifically, the difference between the degradation distribution probability of the current degradation area under different degradation types analyzed by the image convolutional neural network model and the distribution probability of the document's historical degradation characteristics is measured by formula (1). By introducing the historical occurrence frequency of each type as its influence weight and combining the degradation characteristics of government documents, the influence weight of the high-frequency type is reduced, and the model's attention to the low-frequency degradation type is enhanced. The cross-entropy loss minimization optimization process is used to improve the accuracy of the degradation distribution probability, thereby improving the accuracy of the degradation type analysis process to which the degradation area belongs.
[0059] Through this scheme, based on mathematical analysis methods, a preset degradation classification minimization loss function is constructed, and the difference between the degradation distribution probability of the current degraded area under different degradation types analyzed by the image convolutional neural network model and the distribution probability of the historical degradation characteristics of the document is measured. The cross-entropy loss minimization optimization process is used to improve the accuracy of the degradation distribution probability, thereby improving the accuracy of the degradation type analysis process to which the degraded area belongs.
[0060] In some embodiments, a document image data set is analyzed to identify and determine the document layout structure area, the main text area and the additional mark area in each document image in the document image data set, and the document content in each document area is extracted; based on the degraded area information, a number of cross-fuzzy areas and their corresponding document areas are determined according to the intersection area between the feature degraded area and the document layout structure area / the main text area / the additional mark area; based on the feature degradation type and the document area to which the cross-fuzzy area belongs, and according to the document content corresponding to the document area, an inference-based completion analysis is performed on the cross-fuzzy area to determine the corresponding information to be completed in each cross-fuzzy area; based on the document content in each document area and the information to be completed in each cross-fuzzy area, a multi-dimensional information set of the document content is constructed.
[0061] The document layout structure area can be the framework area used to define the overall layout of the document, such as headers, footers, table borders, column lines and other structural document elements. The body area can be the continuous area in the document that carries the core text information, which can include printed or handwritten text content. The additional mark area can be a special identification area with legal effect in the document, such as a unit watermark, approval stamp, signature area, etc.
[0062] The document content may be the document element content corresponding to different document regions. The intersection region may be the overlapped part of the distribution of different document regions and the feature degradation region. The belonging document region may be the document region where each intersection region is located. The inference completion analysis may be the process of reconstructing the degraded content through multimodal information fusion based on the document context semantics, structural features and degradation type.
[0063] The information to be completed may be document information that needs to be completed in the intersection area.
[0064] Specifically, the existing OCR-based document completion technology lacks document area sensitivity, and it is difficult to perform targeted completion processing on the layout area, text area and mark area in government documents, resulting in low completion accuracy. This solution is based on the MaskR-CNN model, with ResNet-101 as the backbone network, outputting the instance segmentation mask of the document image, dividing the corresponding document layout structure area, and using the connected domain analysis method to detect text lines in the non-layout area, dividing the corresponding text area, and further dividing the corresponding additional mark area based on the color space analysis algorithm (such as HSV channel) and the shape matching algorithm. The overlapping area between the feature degradation area and the corresponding document layout structure area / text area / additional mark area is taken as the cross-fuzzy area. Based on the feature degradation type and the document area to which the cross-fuzzy area belongs, according to the document content corresponding to the document area, a targeted reasoning completion analysis strategy is adopted for different cross-fuzzy areas, and the information to be completed corresponding to each cross-fuzzy area is obtained, and then the document content in each document area and the information to be completed in each cross-fuzzy area are integrated to construct a multi-dimensional information set of document content.
[0065] Through this scheme, the document layout structure area, main text area and additional markup area in the document are accurately divided, laying the foundation for the completion analysis process and improving the document area sensitivity of the completion analysis process. By targeting the document areas to which different feature degradation types and cross-fuzzy areas belong, and according to the document content corresponding to the document areas, targeted reasoning completion analysis strategies are adopted for different cross-fuzzy areas to obtain the information to be completed corresponding to each cross-fuzzy area, thereby improving the regional matching degree and accuracy of the information to be completed.
[0066] In some embodiments, if the document area is a document layout structure area, the overall layout structure outline of the document is extracted according to the document content corresponding to the document layout structure area; based on the overall layout structure outline of the document, a preset document template database is retrieved to determine whether there is a corresponding document template; if a document template exists, the differentiated outline edge between the overall layout structure outline of the document and the corresponding document template is used as information to be completed; if the document template does not exist, based on the graph neural network and according to the feature degradation type, the overall layout structure outline of the document is analyzed for line continuity completion, the completed structural lines in the cross-fuzzy area are determined, and the completed structural lines are used as information to be completed.
[0067] The overall document layout structure outline may be an outline of lines constituting the document layout structure.
[0068] The preset document template database may be a preset database for storing various government document templates.
[0069] The corresponding document template may be a document template in a preset document template database that is consistent with the overall layout structure outline of the current document. The differential outline edge may be the difference between the overall layout structure outline of the document and the structure outline in the corresponding document template, that is, the missing layout structure part in the current document. The line continuity completion analysis may be a document completion analysis process that aims to ensure the continuity of the layout structure lines.
[0070] Specifically, in the process of completing the analysis of the document layout structure, since the document layout structure is usually composed of different types of continuous wireframes and does not involve specific deep meanings, the main focus is on the analysis of the document layout structure contour. Through image analysis algorithms such as edge detection algorithms, the geometric features of the document layout structure area (such as table borders, paragraph boundaries, title bar edges) are identified, and the overall document layout structure contour containing lines, rectangular frames and text block positions is extracted. The overall document layout structure contour is matched with various document templates in the preset document template database for similarity. The similarity matching process can be based on the coordinate distribution and topological relationship of contour key points (such as corner points and intersection points). Calculate the structural similarity score. If there is a template whose matching score exceeds the preset threshold, it is determined that there is a corresponding document template, otherwise it is determined that there is no corresponding template. If there is a corresponding template, the overall layout structure outline of the document is superimposed and compared with the standardized outline of the corresponding template, and the differences between the two in edge direction, line length and connection point position are extracted, and the differentiated contour edges are generated as the information to be completed. If there is no corresponding template, a graph structure with layout structure lines as nodes and line connection relationships as edges is constructed, and a pre-trained graph neural network model (such as GCN, GAT) is used to analyze the topological dependency relationship between nodes, predict the direction and connection method of missing lines, and generate completed structure lines as the information to be completed.
[0071] Through this solution, the overall layout structure contour of the document corresponding to the document layout structure area is extracted, and this is used as a retrieval basis to search the preset document template database. If the document template exists, the differentiated contour edge between the overall layout structure contour of the document and the corresponding document template is used as the information to be completed, thereby reducing the occupation of computing resources. If the document template does not exist, the overall layout structure contour of the document is analyzed for line continuity completion based on the feature degradation type, and the corresponding completed structure lines are used as the information to be completed, thereby improving the matching degree between the completed content and the non-standard document, and improving the flexibility and accuracy of the layout structure completion analysis process.
[0072] In some embodiments, if the document area is a text area, edge feature analysis is performed on the text information in the text area according to an image edge detection algorithm to determine the printed information and handwritten information; based on the image edge detection algorithm, the handwritten information is analyzed to determine a handwriting feature set; based on a preset feature point descriptor, the handwriting feature set is analyzed to determine a handwriting feature vector; based on the handwriting feature vector, a handwriting analogy analysis is performed on the handwriting information to determine the handwritten text semantic information; based on a natural language analysis algorithm, contextual semantic continuity analysis is performed on the printed information and the handwritten text semantic information to determine semantic continuity guarantee completion information, and the semantic continuity guarantee completion information is used as the information to be completed.
[0073] Printed information can be standardized text content generated by printing equipment in a document, with regular fonts and spacing. Handwritten information can be manually written text content in a document, with fonts and handwriting that vary from person to person. A handwriting feature set can be unique handwriting attributes extracted from handwriting, such as stroke features, connecting strokes, tilt angles, etc.
[0074] The preset feature point descriptor may be a quantitative model for quantitatively describing key local features of handwriting, and the preset feature point descriptor may adopt SIFT descriptor (Scale-Invariant Feature Transform). The handwriting feature vector may be a numerical vector obtained by transforming the note features using the preset feature point descriptor. The handwritten text semantic information may be the handwritten text content and its semantic meaning restored based on handwriting analysis. Contextual semantic continuity analysis may be a document content completion process that ensures that the completed content is consistent with the original text in terms of grammar, logic and subject. Semantic continuity guarantees that the completed information may be the completed content obtained through contextual reasoning and consistent with the original text semantics.
[0075] Specifically, printed text and handwritten text often coexist in government documents. The processing logic of printed text and handwritten text is significantly different. Printed text has a fixed font shape and can be directly recognized by OCR; while handwritten text requires personalized handwriting analysis. If it is not distinguished, it will lead to completion errors (such as misusing printed text rules for handwritten text). Handwritten text completion depends on individual handwriting features. Image edge detection algorithms, such as the Canny edge detection algorithm, are used to evaluate the regularity of the text contour in the text area. If the contour regularity is higher than the threshold (such as curvature variance <0.1, spacing standard deviation <5 pixels), it is marked as printed text. If the contour is irregular (such as curvature variance ≥0.1), it is marked as The handwriting is recorded as handwriting, and its connected domain is extracted. The following features are extracted for each connected domain of handwriting: the inclination angle of the stroke is detected by Hough transform to extract the stroke direction feature, the writing force is inferred according to the change of the stroke width to extract the pen pressure feature, and the topological relationship between the connecting points of the strokes is analyzed to extract the characteristics of the connected strokes; the above features are encoded into a handwriting feature vector of a specified dimension (such as 128 dimensions) using the preset SIFT feature point descriptor, and the handwriting information in the document body area is converted into handwriting recognition according to the handwriting feature vector to determine the text information corresponding to the handwriting, and the text information is contextually semantically parsed through a natural language analysis model, such as the BERT model (Bidirectional Encoder Representations from Transformers, bidirectional encoding representation based on transformers) to determine the semantic information of the handwriting text, and further through the natural language analysis model, the printed information and the handwriting text semantic information are contextually semantically analyzed to detect the semantic logic fault in the text, and the semantic continuity guarantee completion information is generated according to the contextual semantic information of the location of the semantic logic fault and combined with the domain dictionary, which is used as the information to be completed.
[0076] Through this solution, the printed information and handwritten information in the main text area are separated, and the semantic information of the handwritten text is determined by analyzing and extracting the note features of the handwritten text. Based on natural language analysis technology, contextual semantic continuity analysis is performed on the printed information and the handwritten text semantic information to determine the semantic continuity guarantee completion information, and the semantic continuity guarantee completion information is used as the information to be completed, so as to avoid error propagation caused by mixed processing of handwritten and printed text, and improve the adaptability of the completion analysis process to diverse documents.
[0077] In some embodiments, if the document area is an additional mark area, the background semi-transparent watermark content and the foreground dark seal content are divided and extracted according to the document content corresponding to the additional mark area; the background semi-transparent watermark content is analyzed and the watermark unit content is extracted to determine the unit to which the document belongs; the foreground dark seal content is analyzed to determine the document approval unit; according to the unit to which the document belongs and the document approval unit, the document verification information is sent to the corresponding unit, and the verification feedback information and the target seal image provided by the corresponding unit are received; if the verification feedback information is that the verification is correct, the difference contour between the watermark unit content and the cross-fuzzy area is used as the information to be completed, and the difference contour between the target seal image and the foreground dark seal content is used as the information to be completed.
[0078] The background semi-transparent watermark content can be an identifying graphic or text superimposed on the bottom layer of the document with low transparency, usually associated with the unit to which the document belongs. The foreground dark seal content can be a dark seal or signature covering the surface of the document, used to identify the approval unit or responsible person.
[0079] The watermark unit content can be the smallest unit content that constitutes the document watermark. The document-affiliated unit can be the organization to which the current document belongs, which can be determined by parsing the watermark content. The document approval unit can be the organization that reviews and confirms the document content, which can be identified by the seal content. The verification feedback information can be the watermark and seal verification results fed back by the corresponding unit (such as "verification is correct" or "verification is abnormal"). The target seal image can be a standard seal image provided by the corresponding unit, which is used as a reference image information for seal completion.
[0080] The difference contour may be a shape difference contour between the cross-blur region and the standard watermark / seal image.
[0081] Specifically, government documents usually contain non-text marking elements such as watermarks and seals, which are used to prove the legality and authority of the document and prevent the document content from being tampered with or forged. In the process of analyzing and completing document elements such as watermarks and seals that serve as document markers, due to the different transparency and color depth of watermarks and seals, they need to be processed separately for accurate completion. Semi-transparent watermarks are easily disturbed by the background, and dark seals require high-precision contour extraction. In addition, since watermarks and seals are used as document security identifiers, they cannot be completed directly, otherwise it is easy to create document forgery loopholes. It is necessary to pass the verification of the unit to which the watermark or seal belongs before the mark completion in the electronic process can be carried out. The background and foreground of the document image are separated through image segmentation algorithms (such as the threshold method based on the HSV color space), the low-saturation background area is identified, and the low-saturation background area is binarized. Through OCR Identify and extract the corresponding repeated text as the background semi-transparent watermark content, and extract the watermark unit content from it, and at the same time identify the foreground high contrast area, use the edge tracking algorithm (such as Suzuki contour detection) to close the contour, extract the foreground dark seal content, and retrieve the corresponding institutional units according to the watermark unit content and the foreground dark seal content. Package the corresponding document content as document verification information and send it to the corresponding structural unit, and receive the verification feedback information provided by the corresponding institutional unit. If the verification feedback information is a verification exception, the current document will be marked as a verification exception document and submitted for manual review. If the verification feedback information is that the verification is correct, then through the image edge detection algorithm, the difference contour between the watermark unit content and the cross-fuzzy area is used as the information to be completed, and the difference contour between the target seal image and the foreground dark seal content is used as the information to be completed.
[0082] Through this scheme, the watermark content and the seal content in the additional mark area are separated to clarify the unit to which the current document belongs and the approval unit, and the document verification information is sent to the corresponding unit. After determining that the verification feedback information is correct, the difference contour between the watermark unit content and the cross-fuzzy area is used as the information to be completed. At the same time, the difference contour between the target seal image and the foreground dark seal content is used as the information to be completed. While avoiding document tampering and forgery, accurate analysis of the watermark and seal content completion information is achieved.
[0083] In some embodiments, based on the multidimensional information set of the document content, the document content in each document area is completed according to the information to be completed in each cross-fuzzy area to determine the completed document content; based on the natural language analysis algorithm, the completed document content is analyzed to determine the document government affairs process, document government affairs rules, document government affairs timeliness information and document government affairs responsible subject; according to the government affairs process directed graph, the document government affairs process is topologically matched and evaluated to determine the topological matching degree of the process nodes; according to the government affairs rule logical expression set, the document government affairs rules are verified for rule compliance to determine the rule matching degree; according to the government affairs timeliness information, the document government affairs timeliness information is verified to determine the timing matching degree; according to the government affairs responsibility matrix, the document government affairs responsible subject is evaluated for cosine similarity to determine the responsibility matching degree; according to the process node topological matching degree, rule matching degree, timing matching degree and responsibility matching degree, an association analysis result is constructed.
[0084] The government affairs process information includes a directed graph of government affairs processes, a set of logic expressions of government affairs rules, government affairs timeliness information, and a government affairs responsibility matrix. The directed graph of government affairs processes can be a topological structure diagram that describes the nodes of government affairs processing processes and their execution order. The nodes represent the approval links, and the directed edges represent the direction of the process. The set of logic expressions of government affairs rules can be a set of rules connected by logical operators (AND / OR / NOT) to constrain government affairs processing rules. Government affairs timeliness information can be the time limit of each link in the government affairs process. The government affairs responsibility matrix can be a matrix that records the corresponding relationship between each process node and the responsible subject, which is used to clarify the ownership of government affairs processing rights and responsibilities. The completed document content can be the complete document information obtained by completing the cross-fuzzy area in the document according to the information to be completed. The topological matching evaluation can be an evaluation process for evaluating the matching degree between the process in the completed document and the corresponding government affairs process. The topological matching evaluation can be implemented by using a directed graph topological matching algorithm. The process node topological matching degree can be a quantitative indicator of the structural similarity between the document government affairs process and the directed graph of the government affairs process.
[0085] Rule compliance verification can be an evaluation process used to verify whether the content of a document complies with the rules specified in the government affairs process. Rule compliance verification converts the logical expression of the document's government affairs rules into conjunctive normal form (CNF) and performs a subset inclusion comparison with the set of government affairs rule logical expressions.
[0086] The rule matching degree can be the compliance score between the document government rules and the government rule logical expression.
[0087] Time validity verification can be an evaluation process for evaluating whether the time information in the current completed document complies with the time validity stipulated in the government affairs process. Time validity verification can be achieved by comparing the timestamp information in the document with the time range corresponding to the time validity stipulated in the government affairs process.
[0088] The temporal matching degree can be a time consistency score between the government affairs timeliness information in a document and the government affairs timeliness information.
[0089] Cosine similarity evaluation can be a process of evaluating the matching degree of responsibility division based on the cosine similarity of the vectors between the responsibility information in the current document and the government process responsibility matrix.
[0090] The responsibility matching degree can be the correlation score between the government responsibility subject in the document and the person responsible in the government responsibility matrix.
[0091] Specifically, in the process of conducting logical association analysis on the multidimensional information set of document content based on government process information, the degree of match between the completed document content and the corresponding government process is evaluated mainly from four dimensions: process, rule, timeliness and responsibility; government processes have strict sequence requirements, and process deviations may lead to process violations; rule deviations between government processes and documents will lead to document approval failure; timeliness mismatch will cause contradictions and deviations in the document processing process; deviations in the responsible subject will lead to unclear division of responsibilities; through topological matching evaluation, the topological matching degree of the process node is determined to reflect the degree of match between the process in the current document and the corresponding government process; through rule compliance verification, the rule matching degree is determined to reflect whether the current document content meets the requirements of government process rules; through timeliness verification, the timing matching degree is determined to reflect whether the timestamp information in the document meets the timeliness regulations in the government process; through the government responsibility matrix and the division information of the responsible subject in the document, the cosine similarity evaluation is performed to reflect whether the division of responsibilities conflicts; according to the process node topological matching degree, rule matching degree, timing matching degree and responsibility matching degree, the association analysis results are comprehensively constructed from four dimensions.
[0092] Through this solution, based on government process information, starting from the four dimensions of process, rules, timeliness and responsibility, a logical correlation analysis is performed on the multidimensional information set of document content. Through the four quantitative indicators of process node topology matching, rule matching, timing matching and responsibility matching, the matching degree between the multidimensional information set of document content and the corresponding government process information in different dimensions is mapped respectively. This is used as the correlation analysis result to improve the scientificity and comprehensiveness of the logical correlation analysis process.
[0093] In some embodiments, the process node topology matching, rule matching, timing matching and responsibility matching in the association analysis results are compared with the corresponding matching ranges respectively; if there is a situation where the process node topology matching / rule matching / timing matching / responsibility matching is not within the corresponding matching range, the corresponding abnormal matching is extracted; based on the color-differentiated document annotation strategy and government process information, according to the abnormality type corresponding to the abnormal matching, the local content of the document corresponding to the abnormal matching is highlighted, and the corresponding adjustment suggestions are determined.
[0094] The color-differentiated document annotation strategy may be an annotation strategy that uses different colors to mark different types of content anomalies in a document.
[0095] Adjustment suggestions can be corresponding adjustment suggestions for abnormal content in documents based on government process information.
[0096] Specifically, government documents must strictly follow the preset processes, rules, time limits and responsibility requirements. By comparing the relationship between different matching degrees and the corresponding range thresholds, the different content deviations in the document can be clarified to avoid subjective judgment errors. The process node topology matching degree, rule matching degree, timing matching degree and responsibility matching degree obtained by the analysis of the above embodiments are compared with the corresponding ranges respectively. If a certain matching degree exceeds the threshold range, it is marked as an abnormal matching degree and its type (such as "process exception") is recorded. Predefined colors are selected according to the exception type: process exception (red), rule exception (blue), time limit exception (yellow), responsibility exception (purple). The local content corresponding to the abnormal matching degree (such as the missing node paragraph in the flowchart) is located in the document, and the area is selected with the corresponding color to realize the marking of abnormal content. Further, according to different types of exceptions, the correct reference information in the government process information is extracted as the corresponding adjustment suggestion.
[0097] Through this solution, accurate positioning and efficient correction of government document errors are achieved through matching comparison and visual annotation. The color differentiation strategy reduces the complexity of manual review, and the highlighted annotation directly points to the problem area. Combined with targeted adjustment suggestions, the standardization and efficiency of government processing procedures are significantly improved. At the same time, through automated anomaly detection and prompts, compliance risks caused by omissions or misjudgments are reduced, ensuring that the document content is highly consistent with government process requirements.
[0098] Figure 3 A schematic diagram of the structure of an intelligent government service terminal based on OCR recognition technology provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, an intelligent government service terminal 300 based on OCR recognition technology in this embodiment includes: a degradation analysis module 301, a document analysis module 302, a correlation analysis module 303 and a document marking module 304.
[0099] The degradation analysis module 301 is used to obtain a document image data set, perform degradation identification processing on each document image in the document image data set, and determine the degradation area information; the document analysis module 302 is used to analyze the document image data set based on the degradation area information, and determine the document content multidimensional information set; the association analysis module 303 is used to obtain government process information, perform logical association analysis on the document content multidimensional information set based on the government process information, and determine the association analysis result; the document marking module 304 is used to mark and correct the document content multidimensional information set according to the association analysis result, and determine and output a document processing report.
[0100] Optionally, the degradation analysis module 301 is specifically used to: analyze the document image data set to extract several feature degradation areas of each document image; based on the image convolutional neural network model, extract several image feature vectors corresponding to the feature degradation areas, and perform feature degradation type probability distribution analysis on each of the image feature vectors to determine the degradation distribution probability of each feature degradation area under each feature degradation type; the feature degradation types include creases, fades, stains and missing; based on a preset degradation classification minimization loss function, determine the feature degradation type corresponding to each feature degradation area according to the degradation distribution probability of each feature degradation area under each feature degradation type; and construct the degradation area information according to each feature degradation area and its corresponding feature degradation type.
[0101] Optionally, the preset degradation classification minimization loss function in the degradation analysis module 301 is specifically the following formula: ; in, is the cross entropy loss, is the degenerate type index, is the total number of degenerate types, For the The historical frequency of degenerate types, For the A vector of preset type labels for the degenerate types, For the The degradation distribution probability of each degradation type.
[0102] Optionally, the document analysis module 302 is specifically used to: analyze the document image data set, identify and determine the document layout structure area, the text area and the additional mark area in each document image in the document image data set, and extract the document content in each document area; based on the degraded area information, determine a number of cross-fuzzy areas and their corresponding document areas according to the intersection area between the feature degraded area and the document layout structure area / the text area / the additional mark area; based on the feature degradation type and the document area to which the cross-fuzzy area belongs, perform inference completion analysis on the cross-fuzzy area according to the document content corresponding to the document area, and determine the corresponding information to be completed in each cross-fuzzy area; construct the document content multidimensional information set according to the document content in each document area and the information to be completed in each cross-fuzzy area.
[0103] Optionally, the document analysis module 302 performs an inferential completion analysis on the cross-fuzzy area based on the feature degradation type and the document area to which the cross-fuzzy area belongs, according to the document content corresponding to the document area, and determines the corresponding information to be completed in each cross-fuzzy area, specifically for: if the document area is the document layout structure area, extracting the overall layout structure outline of the document according to the document content corresponding to the document layout structure area; searching a preset document template database based on the overall layout structure outline of the document to determine whether there is a corresponding document template; if the document template exists, taking the differentiated contour edge between the overall layout structure outline of the document and the corresponding document template as the information to be completed; if the document template does not exist, based on the graph neural network, performing a line continuity completion analysis on the overall layout structure outline of the document according to the feature degradation type, determining the completion structure lines of the cross-fuzzy area, and taking the completion structure lines as the information to be completed.
[0104] Optionally, the document analysis module 302 performs an inference-based completion analysis on the cross-fuzzy area based on the feature degradation type and the document area to which the cross-fuzzy area belongs, according to the document content corresponding to the document area, and determines the information to be completed corresponding to each cross-fuzzy area, specifically for: if the document area to which the cross-fuzzy area belongs is the main text area, performing edge feature analysis on the text information in the main text area according to an image edge detection algorithm to determine the printed information and handwritten information; analyzing the handwritten information based on an image edge detection algorithm to determine a handwriting feature set; analyzing the handwriting feature set based on a preset feature point descriptor to determine a handwriting feature vector; performing a handwriting analogy analysis on the handwriting information based on the handwriting feature vector to determine the semantic information of the handwritten text; performing a contextual semantic continuity analysis on the printed information and the handwritten text semantic information based on a natural language analysis algorithm to determine the semantic continuity-guaranteed completion information, and using the semantic continuity-guaranteed completion information as the information to be completed.
[0105] Optionally, the document analysis module 302 performs an inference-based completion analysis on the cross-fuzzy area based on the feature degradation type and the document area to which the cross-fuzzy area belongs, according to the document content corresponding to the document area, to determine the information to be completed corresponding to each cross-fuzzy area, specifically for: if the document area to which the cross-fuzzy area belongs is the additional mark area, according to the document content corresponding to the additional mark area, divide and extract the background semi-transparent watermark content and the foreground dark seal content; analyze the background semi-transparent watermark content, extract the watermark unit content, and thereby determine the unit to which the document belongs; analyze the foreground dark seal content, and determine the document approval unit; according to the document unit to which the document belongs and the document approval unit, send the document verification information to the corresponding unit, and receive the verification feedback information and the target seal image provided by the corresponding unit; if the verification feedback information is that the verification is correct, use the difference contour between the watermark unit content and the cross-fuzzy area as the information to be completed, and use the difference contour between the target seal image and the foreground dark seal content as the information to be completed.
[0106] Optionally, the association analysis module 303 is specifically used to: based on the multidimensional information set of the document content, according to the information to be completed in each of the cross-fuzzy areas, complete the document content in each document area to determine the completed document content; based on the natural language analysis algorithm, analyze the completed document content to determine the document government affairs process, document government affairs rules, document government affairs timeliness information and document government affairs responsible subject; according to the government affairs process directed graph, perform a topological matching evaluation on the document government affairs process to determine the topological matching degree of the process node; according to the government affairs rule logical expression set, perform rule compliance verification on the document government affairs rules to determine the rule matching degree; according to the government affairs timeliness information, perform timeliness verification on the document government affairs timeliness information to determine the timing matching degree; according to the government affairs responsibility matrix, perform a cosine similarity evaluation on the document government affairs responsible subject to determine the responsibility matching degree; construct the association analysis result according to the process node topological matching degree, the rule matching degree, the timing matching degree and the responsibility matching degree.
[0107] Optionally, the document marking module 304 is specifically used to: compare the process node topology matching degree, the rule matching degree, the timing matching degree and the responsibility matching degree in the association analysis result with the corresponding matching degree ranges respectively; if the process node topology matching degree / the rule matching degree / the timing matching degree / the responsibility matching degree is not within the corresponding matching degree range, extract the corresponding abnormal matching degree; based on the color-differentiated document annotation strategy and the government process information, according to the abnormal type corresponding to the abnormal matching degree, highlight the local content of the document corresponding to the abnormal matching degree, and determine the corresponding adjustment suggestions.
[0108] The terminal of this embodiment can be used to execute the method of any of the above embodiments, and its implementation principle and technical effects are similar, which will not be repeated here.
Claims
1. An intelligent government service method based on OCR recognition technology, characterized in that: include: Acquire a document image data set, perform degradation recognition processing on each document image in the document image data set, and determine degradation area information; Based on the degraded region information, analyzing the document image data set to determine a document content multidimensional information set; Acquire government affairs process information, and based on the government affairs process information, perform a logical association analysis on the multi-dimensional information set of the document content to determine an association analysis result; According to the association analysis result, the document content multidimensional information set is subjected to marked error correction, and a document processing report is determined and output.
2. The method according to claim 1, characterized in that The step of performing degradation identification processing on each document image in the document image data set to determine degradation area information includes: Analyzing the document image data set to extract a number of characteristic degradation regions of each document image; Based on the image convolutional neural network model, extract the image feature vectors corresponding to the plurality of feature degradation regions, and perform feature degradation type probability distribution analysis on each of the image feature vectors to determine the degradation distribution probability of each of the feature degradation regions under each feature degradation type; The types of feature degradation include creases, fades, stains, and loss; Based on a preset degradation classification minimization loss function, determining a feature degradation type corresponding to each feature degradation region according to the degradation distribution probability of each feature degradation region under each feature degradation type; The degradation region information is constructed according to each of the feature degradation regions and the corresponding feature degradation type.
3. The method according to claim 2, characterized in that The preset degradation classification minimizes the loss function, which is specifically the following formula: ; in, is the cross entropy loss, is the degenerate type index, is the total number of degenerate types, For the The historical frequency of degenerate types, For the A vector of preset type labels for the degenerate types, For the The degradation distribution probability of each degradation type.
4. The method according to claim 2, characterized in that: The step of analyzing the document image data set based on the degraded region information to determine a document content multidimensional information set includes: Analyzing the document image data set, identifying and determining the document layout structure area, the text area and the additional mark area in each document image in the document image data set, and extracting the document content in each document area; Based on the degenerate region information, a plurality of cross fuzzy regions and their corresponding document regions are determined according to the cross regions between the characteristic degenerate region and the document layout structure region / the text region / the additional mark region; Based on the feature degradation type and the document area to which the cross-fuzzy area belongs, and according to the document content corresponding to the document area, performing an inference-based completion analysis on the cross-fuzzy area to determine the corresponding information to be completed in each cross-fuzzy area; The document content multidimensional information set is constructed according to the document content in each document area and the information to be completed in each cross-fuzzy area.
5. The method according to claim 4, characterized in that The method of performing inference-based completion analysis on the cross-fuzzy regions based on the feature degradation type and the document regions to which the cross-fuzzy regions belong according to the document content corresponding to the cross-fuzzy regions, and determining the corresponding information to be completed in each cross-fuzzy region, includes: If the document region is the document layout structure region, extracting the overall document layout structure outline according to the document content corresponding to the document layout structure region; According to the overall layout structure outline of the document, a preset document template database is searched to determine whether a corresponding document template exists; If the document template exists, the differential contour edge between the overall layout structure contour of the document and the corresponding document template is used as the information to be completed; If the document template does not exist, based on the graph neural network and according to the feature degradation type, a line continuity completion analysis is performed on the overall layout structure contour of the document to determine the completed structural lines of the cross-fuzzy area, and the completed structural lines are used as the information to be completed.
6. The method according to claim 4, characterized in that The method of performing inference-based completion analysis on the cross-fuzzy regions based on the feature degradation type and the document regions to which the cross-fuzzy regions belong according to the document content corresponding to the cross-fuzzy regions, and determining the corresponding information to be completed in each cross-fuzzy region, includes: If the document area is the text area, edge feature analysis is performed on the text information in the text area according to an image edge detection algorithm to determine printed information and handwritten information; Analyze the handwriting information based on an image edge detection algorithm to determine a handwriting feature set; Analyze the handwriting feature set according to the preset feature point descriptor to determine the handwriting feature vector; Performing handwriting analogy analysis on the handwriting information according to the handwriting feature vector to determine the semantic information of the handwriting text; Based on a natural language analysis algorithm, contextual semantic continuity analysis is performed on the printed information and the handwritten text semantic information to determine semantic continuity guarantee supplement information, and the semantic continuity guarantee supplement information is used as the information to be supplemented.
7. The method according to claim 4, characterized in that The method of performing inference-based completion analysis on the cross-fuzzy regions based on the feature degradation type and the document regions to which the cross-fuzzy regions belong according to the document content corresponding to the cross-fuzzy regions, and determining the corresponding information to be completed in each cross-fuzzy region, includes: If the document area is the additional mark area, dividing and extracting the background semi-transparent watermark content and the foreground dark seal content according to the document content corresponding to the additional mark area; Analyze the background semi-transparent watermark content and extract the watermark unit content to determine the unit to which the document belongs; Analyze the content of the foreground dark seal to determine the document approval unit; According to the unit to which the document belongs and the document approval unit, the document verification information is sent to the corresponding unit, and the verification feedback information and the target seal image provided by the corresponding unit are received; If the verification feedback information is that the verification is correct, the difference contour between the watermark unit content and the cross fuzzy area is used as the information to be completed, and the difference contour between the target seal image and the foreground dark seal content is used as the information to be completed.
8. The method according to claim 7, characterized in that The government affairs process information includes a government affairs process directed graph, a government affairs rule logic expression set, government affairs timeliness information, and a government affairs responsibility matrix. Based on the government affairs process information, performing a logical association analysis on the document content multidimensional information set to determine the association analysis result includes: Based on the document content multidimensional information set, and according to the information to be completed in each cross-fuzzy area, the document content in each document area is completed to determine the completed document content; Based on the natural language analysis algorithm, the content of the completed document is analyzed to determine the document government affairs process, document government affairs rules, document government affairs timeliness information and document government affairs responsible subject; According to the government affairs process directed graph, a topological matching evaluation is performed on the document government affairs process to determine the topological matching degree of the process nodes; According to the set of logical expressions of the government affairs rules, the rule compliance of the document government affairs rules is verified to determine the rule matching degree; According to the government affairs timeliness information, verify the timeliness of the government affairs timeliness information of the document to determine the time sequence matching degree; According to the government affairs responsibility matrix, a cosine similarity evaluation is performed on the government affairs responsibility subject of the document to determine the responsibility matching degree; The association analysis result is constructed according to the process node topology matching degree, the rule matching degree, the timing matching degree and the responsibility matching degree.
9. The method according to claim 8, characterized in that The step of performing mark-based error correction on the document content multidimensional information set according to the association analysis result includes: Comparing the process node topology matching degree, the rule matching degree, the timing matching degree and the responsibility matching degree in the association analysis result with the corresponding matching degree ranges respectively; If there is a situation where the process node topology matching degree / the rule matching degree / the timing matching degree / the responsibility matching degree is not within the corresponding matching degree range, the corresponding abnormal matching degree is extracted; Based on the color-differentiated document annotation strategy and the government affairs process information, according to the exception type corresponding to the exception match degree, the local content of the document corresponding to the exception match degree is highlighted, and the corresponding adjustment suggestion is determined.
10. An intelligent government service terminal based on OCR recognition technology, characterized in that: include: A degradation analysis module, used to obtain a document image data set, perform degradation identification processing on each document image in the document image data set, and determine degradation area information; A document analysis module, used to analyze the document image data set based on the degraded region information to determine a document content multi-dimensional information set; A correlation analysis module, used to obtain government affairs process information, and based on the government affairs process information, perform a logical correlation analysis on the multi-dimensional information set of the document content to determine a correlation analysis result; The document marking module is used to perform marking-based error correction on the multi-dimensional information set of the document content according to the association analysis result, and determine and output a document processing report.
Citation Information
Patent Citations
Document layout analysis method and system based on multi-scale training and cascade detection
CN113420669A
Paper draft extraction and correction method and system based on computer vision and deep learning
CN117423111A
License plate number extraction method and system based on image recognition
CN117854052A
System and method for document management
US20090228819A1