Document structuring method and device based on OCR
Through the OCR-based document structure method, image processing and text recognition of the handwritten construction records of engineers is carried out to generate standardized documents, which solves the problem that documents cannot be directly utilized in the prior art and realizes efficient identification and management of documents.
Patent Information
- Application Number
- CN202311660374.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, engineers need to handwritten construction records when inspecting the construction site, which leads to the inability to use the documents directly and needs to return to the office to print for approval. It is time-consuming and difficult to uniformly format and archive, affecting the progress of the project and document management.
Using the OCR-based document structure method, local feature extraction and overall feature extraction are performed on the images of the document to be processed, weighted fusion is performed, text information is identified, and standardized document materials are generated.
It improves the accuracy and efficiency of document recognition, realizes the standardization, onlineization and structure of documents, reduces the intensity of data entry, and improves the efficiency of document management and archiving.
Smart Images

Figure CN120107980A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a document structuring method and device based on OCR. Background Art
[0002] At present, when engineers go to the construction site for inspection, they need to write down the construction records by hand, and then print them out in the office for approval to form a document. OCR (optical character recognition) text recognition refers to the process in which an electronic device (such as a scanner) checks the characters printed on paper and then uses character recognition methods to translate the shapes into computer text; that is, the process of scanning text materials and then analyzing and processing the image files to obtain text and layout information. How to debug or use auxiliary information to improve the recognition accuracy is the most important topic of OCR. The main indicators for measuring the performance of an OCR system are: rejection rate, false recognition rate, recognition speed, product stability, ease of use and feasibility, etc. Summary of the invention
[0003] In order to provide better standardized documents, an embodiment of the present invention provides a document structuring method and device based on OCR.
[0004] In a first aspect, an embodiment of the present invention provides a document structuring method based on OCR, the method comprising:
[0005] Processing the acquired document to be processed to obtain an image to be processed;
[0006] Dividing the image to be processed into a plurality of blocks;
[0007] Inputting the multiple blocks into a local feature extraction module respectively to obtain corresponding local features;
[0008] Inputting the image to be processed into an overall feature extraction module to obtain overall features;
[0009] Weighted fusion of the plurality of local features and the overall feature to obtain a data image feature;
[0010] Performing text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data;
[0011] Generate standardized document data based on the unstructured text data.
[0012] In one or some optional implementations of the embodiment of the present application, dividing the image to be processed into a plurality of blocks includes:
[0013] Clustering all pixels in the image to be processed to obtain a clustering result;
[0014] According to the clustering result, the pixels belonging to the same category in the image to be processed are divided into a block to obtain the multiple blocks.
[0015] In one or some optional implementations of the embodiment of the present application, each of the local feature extraction modules includes a plurality of convolutional layers and a plurality of depooling layers connected in sequence;
[0016] The step of inputting the plurality of blocks into a local feature extraction module to obtain corresponding local features comprises:
[0017] For each block, inputting the block into the plurality of sequentially connected convolutional layers to obtain a dimensionality reduction feature corresponding to the block;
[0018] The dimension reduction features are input into the plurality of sequentially connected depooling layers to obtain local features corresponding to the blocks.
[0019] In one or some optional implementations of the embodiment of the present application, the weighted fusion of the multiple local features with the overall feature to obtain the data image feature includes:
[0020] Matching each element feature in each of the local features with the overall feature to obtain corresponding positions of each element feature of the local feature in the overall feature;
[0021] According to the corresponding position of each element feature of the local feature in the overall feature, each local feature is weightedly fused with the overall feature to obtain the data image feature.
[0022] In one or some optional implementations of the embodiment of the present application, the performing of text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data includes:
[0023] Obtaining a plurality of feature points in the image to be processed according to the position of each feature of the data image features in the image to be processed;
[0024] For each feature point, a circle with a radius equal to a first preset threshold is set with the feature point as the center;
[0025] Determine whether the number of feature points in the circle is greater than a second preset threshold and all feature points in the circle belong to the same block:
[0026] If yes, extracting texture information from the block where the feature point is located, performing text recognition based on the texture information, and obtaining a text recognition result;
[0027] The unstructured text data is obtained according to all text recognition results.
[0028] In one or some optional implementations of the embodiment of the present application, the generating of standardized document data based on the unstructured text data includes:
[0029] Extracting keywords according to the text queue of the unstructured text data, and obtaining a standardized document template corresponding to the keywords;
[0030] Matching the standardized document template with the text queue to obtain an initialization document;
[0031] Based on the initialization document, the standardized document material is obtained.
[0032] In one or some optional implementations of the embodiment of the present application, obtaining the standardized document material based on the initialization document includes:
[0033] Performing semantic analysis on the initialization document to obtain a semantic analysis result;
[0034] Based on a preset industry terminology database, the semantic analysis result is matched to determine whether there is a mismatch between the semantic analysis result and the industry terminology database;
[0035] If yes, a prompt message is returned, wherein the prompt message is used to prompt that the initialization document needs manual inspection; and the standardized document data is generated according to the manual modification result;
[0036] If not, the initialization document is used as standardized document data.
[0037] In one or some optional implementations of the embodiment of the present application, before dividing the image to be processed into a plurality of blocks, the process further includes:
[0038] Extracting texture features of the image to be processed to obtain texture features of the image to be processed.
[0039] In a second aspect, an embodiment of the present invention provides a document structuring device based on OCR, the device comprising:
[0040] A first processing module is used to process the acquired document to be processed to obtain an image to be processed;
[0041] A first division module, used for dividing the image to be processed into a plurality of blocks;
[0042] A first extraction module, used for inputting the plurality of blocks into a local feature extraction module respectively to obtain corresponding local features;
[0043] A second extraction module is used to input the image to be processed into an overall feature extraction module to obtain overall features;
[0044] A first fusion module, used for weighted fusion of the plurality of local features and the overall feature to obtain a data image feature;
[0045] A first recognition module, configured to perform text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data;
[0046] The first generating module is used to generate standardized document data based on the unstructured text data.
[0047] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned OCR-based document structuring method.
[0048] In a fourth aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the OCR-based document structuring method as described above when executing the computer program.
[0049] In a fifth aspect, an embodiment of the present invention provides a computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the OCR-based document structuring method as described above.
[0050] In a sixth aspect, an embodiment of the present invention provides a chip, the chip including a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run a computer program or instruction to implement the OCR-based document structuring method as described above.
[0051] The beneficial effects of the above technical solution provided by the embodiment of the present invention include at least:
[0052] The OCR-based document structuring method provided by the embodiment of the present invention obtains an image to be processed based on the acquired document to be processed, extracts local features and overall features of the image, performs weighted fusion of corresponding positions, and then performs text recognition on the fused features to obtain standardized document data. The local features extracted by this method contain implicit feature information, and combined with the overall features, can fuse information from multiple levels and angles, thereby improving the accuracy of characterization of text information in the image.
[0053] Through OCR technology, the image information of the document to be processed is automatically identified and mined, and the data is analyzed and processed to form an electronic information view of the business application. With the help of image acquisition technology, mining technology and artificial intelligence technology, information data is intelligently identified, and the standardization, onlineization and structuring of documents are completed, and standard document templates are quickly provided for the project to improve document work efficiency. It is convenient to complete the online review and electronic signature of the engineering process documents in the future, ensure the validity and authenticity of the documents, improve the collection and organization efficiency of archive files, and realize the corresponding storage, organization, filing and archiving of files.
[0054] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0055] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0057] Figure 1 A schematic diagram of the steps of the OCR-based document structuring method provided by an embodiment of the present invention;
[0058] Figure 2 A schematic diagram of a local feature extraction module provided in an embodiment of the present invention;
[0059] Figure 3 A schematic diagram of a process for archiving engineering files after standardization provided by an embodiment of the present invention;
[0060] Figure 4 A schematic diagram of a process of a document structuring method based on OCR provided in an embodiment of the present invention;
[0061] Figure 5 A schematic diagram of an example of an OCR template provided in an embodiment of the present invention;
[0062] Figure 6 A schematic diagram of a second OCR template example provided in an embodiment of the present invention;
[0063] Figure 7 A schematic diagram of the structure of an OCR-based document structuring device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0065] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0066] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0067] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0068] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0069] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0070] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0071] In order to illustrate the technical solution of the present application, a specific embodiment is provided below for illustration.
[0072] The inventors have found that in the prior art, engineering personnel often need to write construction records by hand when inspecting the construction site, especially when the site conditions are poor. The handwritten engineering documents cannot be used directly and need to be printed out in the office for approval, which is very time-consuming and labor-intensive. Construction site personnel cannot unify the document format and compile file numbers, and cannot archive them effectively. After uploading to the internal system, the workload and difficulty of archiving are large, which causes great difficulties for later records and document retrieval. There are many cases of post-approval of information and accumulation of information at the construction site, and timeliness cannot be guaranteed, resulting in the technical information of the project handover being out of sync with the construction progress. When the construction site needs the construction unit, general contractor, and supervision unit to approve the documents, the documents go back and forth between the units. Some projects have poor site environments, and construction personnel need to overcome more difficulties to approve the information.
[0073] Based on OCR technology, being able to accurately identify text content under good conditions is the most important step in document structuring. However, in reality, the handwritten paper documents of engineers often have unclear handwriting due to various reasons, such as unclear handwriting due to pencil writing, handwriting degradation due to wet documents, etc. There are many factors that affect the clarity of document text. As a result, when using OCR for scanning and recognition, it cannot be converted into effective and accurate information.
[0074] Based on this, the inventors have made the present invention after further research and development, providing a document structuring method and device based on OCR.
[0075] Embodiment 1
[0076] The embodiment of the present invention provides a document structuring method based on OCR, referring to Figure 1 As shown, the method includes:
[0077] S101: Processing the acquired document to be processed to obtain an image to be processed;
[0078] In an embodiment of the present application, a document to be processed is obtained, and the document to be processed may be a paper document, a PDF document or other forms. If the document to be processed is a paper document, it is necessary to convert the paper document data into an image format. Among them, the method of converting a paper document into an image can be to scan the paper document with a scanner, or to shoot the paper document with a camera or a mobile phone and other devices to obtain an image to be processed; if the document to be processed is a PDF document, it is necessary to convert the PDF document into a picture format to obtain an image to be processed; similarly, documents to be processed in different formats are converted into images to be processed in a corresponding manner. After obtaining the image to be processed, conventional enhancement and denoising processing are performed on the image to be processed to improve the accuracy of subsequent recognition.
[0079] S102: Divide the image to be processed into a plurality of blocks;
[0080] In the implementation of this application, in the above step S102, the image to be processed is divided into multiple blocks, including:
[0081] Clustering all the pixels in the image to be processed to obtain a clustering result;
[0082] According to the clustering result, the pixels belonging to the same category in the image to be processed are divided into a block to obtain multiple blocks.
[0083] In the embodiment of the present application, before dividing the image to be processed into a plurality of blocks, the process may further include extracting texture features of the image to be processed to obtain texture features of the image to be processed. The texture features of the image to be processed will be directly processed later.
[0084] In an embodiment of the present application, by clustering all the pixels in the image to be processed, a clustering result of the image to be processed is obtained. According to the clustering result, the pixels in the image to be processed that are classified into the same category are combined into a block, and the image to be processed is divided into multiple blocks. The pixels in each block are identified as having similar features or attributes during clustering, forming an obvious set. For example, if clustering is performed based on the texture features of the image to be processed, the clustering algorithm will consider the shape and structural information of the text and divide each word in the image into a block. The clustering algorithm can use classic clustering algorithms such as K-means. Dividing the blocks according to the clustering results provides a good foundation for the accuracy and reliability of subsequent text recognition to generate standard documents.
[0085] S103: inputting the plurality of blocks into a local feature extraction module respectively to obtain corresponding local features;
[0086] In the embodiment of the present application, in the above step S103, each local feature extraction module includes a plurality of convolutional layers and a plurality of depooling layers connected in sequence; the plurality of blocks are respectively input into the local feature extraction module to obtain corresponding local features, including:
[0087] For each block, the block is input into multiple convolutional layers connected in sequence to obtain the dimensionality reduction features corresponding to the block;
[0088] The reduced-dimensional features are input into multiple sequentially connected unpooling layers to obtain the local features corresponding to the blocks.
[0089] In an embodiment of the present application, multiple blocks are respectively input into a local feature extraction module to obtain corresponding multiple local features. First, multiple blocks are input into multiple convolutional layers of a local feature extraction module, and each neuron in the convolutional layer corresponds to a convolutional kernel. These convolutional kernels can be set to extract different features respectively, such as some convolutional kernels extracting texture features in blocks, and some convolutional kernels extracting color features in blocks. The edge features, color features, and texture features in the blocks are extracted through the convolutional layer to obtain the dimensionality reduction features corresponding to each block; then, the dimensionality reduction features are input into multiple anti-pooling layers of the local feature extraction module, and the positions of the dimensionality reduction features can be restored through the anti-pooling layer, which helps to retain more accurate position information in the higher-level features, that is, the anti-pooling layer can restore each feature element in the obtained dimensionality reduction features to the original position, and retain the position information of each feature element in the output high-level local features, and obtain the local features corresponding to each block. The anti-pooling layer realizes denoising and image enhancement at the same time, and extracts the implicit features in the blocks.
[0090] The schematic diagram of the local feature extraction module is shown in the following figure: Figure 2 As shown in the figure, the local feature extraction module includes 3 convolutional layers and 3 anti-pooling layers, and the code in the middle represents the dimension reduction feature. At the same time, since each input block is obtained through a clustering algorithm, the size of the block is different. Therefore, in this solution, the network size of the local feature extraction module can be adaptively changed according to the size of the input block. In specific implementation, it can be realized through the function of dynamically adjusting the input size provided by the deep learning framework.
[0091] S104: inputting the image to be processed into an overall feature extraction module to obtain overall features;
[0092] In the embodiment of the present application, the image to be processed is input into the overall feature extraction module to extract the overall feature representation in the image to be processed. The overall feature extraction module can be a large neural network designed for images, including multi-layer convolution, pooling and full connection operations, so that the overall feature extraction module can extract complex structures and other abstract features in the image to be processed, form a global understanding of the image to be processed, and obtain overall features.
[0093] S105: weighted fusion of the plurality of local features and the overall feature to obtain a data image feature;
[0094] In the embodiment of the present application, in the above step S105, multiple local features are weighted and fused with the overall features to obtain the data image features, including:
[0095] Match each element feature in each local feature with the overall feature to obtain the corresponding position of each element feature of the local feature in the overall feature;
[0096] According to the corresponding position of each element feature of the local feature in the overall feature, each local feature is weightedly fused with the overall feature to obtain the data image feature.
[0097] In the embodiment of the present application, the step of weighted fusion of multiple local features and overall features to obtain data image features includes: the first step, matching each element feature in the local features with the overall features to obtain the position information of each element feature of all local features in the overall features, wherein, since the local features contain the position information of the dimensionality reduction feature, it can be matched with the pixel points in the image to be processed, and the overall features are the same, so there is a one-to-one corresponding position relationship between the local features and the element features in the overall features, wherein the element features represent the specific data in the local features and the overall features; the second step, according to the corresponding positions of each element feature of the local features in the overall features, the element features in the local features are weighted fused with the element features at the corresponding positions in the overall features to obtain the fused features of the element feature positions, and the weighted fusion operation is repeated for each element feature in all local features to obtain the data image features. By fusing the data image features of the local features and the overall features, multi-level and multi-angle information is integrated, which can effectively improve the representation accuracy of the text information in the features.
[0098] S106: performing text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data;
[0099] In the embodiment of the present application, in the above step S106, text recognition is performed based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data, including:
[0100] According to the position of each feature of the data image features in the image to be processed, a plurality of feature points in the image to be processed are obtained;
[0101] For each feature point, a circle with a radius equal to a first preset threshold is set with the feature point as the center;
[0102] Determine whether the number of feature points in the circle is greater than a second preset threshold and all feature points in the circle belong to the same block:
[0103] If yes, extract texture information from the block where the feature point is located, perform text recognition based on the texture information, and obtain a text recognition result;
[0104] Based on all text recognition results, unstructured text data is obtained.
[0105] In the embodiment of the present application, the corresponding position of each feature data in the data image feature in the image to be processed is obtained, each feature data in the data image feature is restored to the corresponding position in the image to be processed, and each feature data in the data feature image is used as a feature point to obtain multiple feature points in the image to be processed; for each feature point, a circle with a radius of the first preset threshold is set with the feature point as the center; it is determined whether the feature point data in the circle is greater than the second preset threshold, and all feature points falling in the circle belong to the same block among the multiple blocks obtained in step S102, that is, it is determined whether the circle encircles part or all of a text in the image to be processed: if so, find the corresponding block where the circle is located, obtain the texture information in the block, perform text recognition based on the block and the corresponding texture information, and obtain the text recognition result; if not, continue to process the next feature point; after completing the processing of all feature points, merge all text recognition results to obtain unstructured text data. Among them, the first preset threshold and the second preset threshold can be set according to the size of the text in the image to be processed during specific implementation. Those skilled in the art can implement the above-mentioned text recognition operation according to the detailed description of the prior art, and no specific limitation is made in the embodiment of the present invention.
[0106] S107: Generate standardized document data based on the unstructured text data.
[0107] In the embodiment of the present application, in the above step S107, generating standardized document data based on unstructured text data includes:
[0108] Extract keywords based on the text queue of the unstructured text data, and obtain standardized document templates corresponding to the keywords;
[0109] Match the standardized document template with the text queue to obtain the initialization document;
[0110] Based on the initialization document, standardized documentation is obtained.
[0111] Among them, based on the initialization document, standardized document materials are obtained, including:
[0112] Perform semantic analysis on the initialization document to obtain semantic analysis results;
[0113] Based on the preset industry terminology database, the semantic analysis results are matched to determine whether there is a mismatch between the semantic analysis results and the industry terminology database;
[0114] If yes, a prompt message is returned, which is used to remind that the initialization document needs manual review; based on the manual modification results, standardized document data is generated;
[0115] If not, use the initialization document as the standard documentation.
[0116] In an embodiment of the present application, unstructured text data is arranged (by row) according to its position information in the image to be processed and placed in a text queue. Based on the text in the text queue, keywords of the text in the text queue are extracted, and a standardized document template corresponding to the keyword is extracted; the text in the text queue is processed by row, and the text is placed in the position of the corresponding content in the standardized document template, and data structuring processing is performed to obtain an initialization document; the initialization document is optimized and adjusted to obtain standardized document data.
[0117] In the embodiment of the present application, the operation of optimizing and adjusting the initialization document includes two parts: text data preprocessing and semantic analysis, wherein the text data preprocessing part uses the NLP intelligent model to perform word analysis and named entity recognition on the standardized structured text data, performs data cleaning on the text data to remove invalid text, limits the maximum length of the text input, and uses the slicing input method if there is a part that exceeds the length, so as to obtain standard data with consistent text length. The semantic analysis part performs semantic analysis on the text at each position in the initialization document; matches the result of the semantic analysis with the preset industry terminology database based on the text at the same position in the standardized document template, and determines whether there is a mismatch between the result of the semantic analysis and the industry terminology database: if so, returns a prompt message, and needs to manually check whether the content of the semantic mismatch is correct, and perform manual modification to obtain standardized document data; if not, the initialization document is used as standardized document data. Among them, if it is inconvenient to perform manual inspection, the initialization document is directly used as standardized document data.
[0118] In the embodiment of the present application, the data structured processing is a method of structured construction using the detection frame and character content corresponding registration. The detection frame output by the detection model outputs the relative coordinates of the two diagonal points, and then the content in the corresponding frame is input into the corresponding coordinate position. The detection frame is aligned and adapted by an intelligent algorithm, and the standardized structured character data of the original image is output, providing the standardized structured character data to the back-end NLP model. Those skilled in the art can implement data structured processing based on the detailed description of the prior art, and no specific limitation is made in the embodiment of the present invention.
[0119] In order to facilitate those skilled in the art to understand the present invention, the specific implementation process of the OCR-based document structuring method provided in the embodiment of the present invention is described more clearly and completely below. Figure 3 and Figure 4 As shown, Figure 3 It is a process diagram based on the standardization and archiving of engineering documents, including 4 steps: the first step is that the technical personnel of the construction unit fill in the inspection record data information on the blank paper document; the second step is that the construction unit scans and OCR recognizes the paper document, and conducts a complete check on the identification information, and submits it for online approval; the third step is to complete the online approval process of the handover technical documents, and stamp them with an electronic seal that complies with the legal and group company archive requirements; the fourth step is to store, organize, file and archive the documents accordingly. Among them, this plan is implemented in the second step. Figure 4 This is a flow chart of the OCR-based document structuring method. After logging into the system, the system will determine the standardized document template corresponding to the current document to be processed, that is, match it with the OCR template in the figure, and then convert the document to be processed into an image format, scan it with the scanner in the figure, and obtain the initialized document through OCR text recognition. The technician will check whether there are any errors. If there are any errors, they will be modified. If there are no errors, they will be submitted for approval and archived.
[0120] In a specific embodiment, the system has a built-in OCR template of QSY1476 "Technical Document Management Specifications for Handover of Refining and Chemical Construction Projects". The technical personnel of the construction unit print the standard blank form of QSY1476 "Technical Document Management Specifications for Handover of Refining and Chemical Construction Projects". After manually filling it out offline, the handwritten document is scanned and recognized through OCR technology, realizing machine intelligent recognition and accurate input. The recognition data is automatically filled in to generate a standard document. After checking and modifying and verifying that it is correct, it is submitted for online approval, stamped with an electronic signature, and sorted and archived. Among them, when recognizing handwritten documents, each time new information is added, it will enter the model training to help improve the analysis system, optimize subsequent decisions, and make the model more accurate.
[0121] The system's built-in OCR templates include: Figure 5 , Figure 6 As shown. Collect all relevant requirements and regulations of the corresponding OCR templates in the petrochemical industry to form a PDF document, split the OCR templates in the PDF document, and split each template that needs to be configured into a single file. For example, if there are 100 templates that need to be configured in the PDF document, it needs to be split into 100 files. Convert the split files into Word documents, take out the Nth cell of the template as the template identifier, which needs to be unique, and then manually configure the blank space of the template that needs to be configured as required to complete the expression, and fill it according to the position count of the cell, such as Figure 5 {{6}}, {{8}}, {{10}}, etc. Figure 6{{@name}}, {{@name1_end}}, {{yn}} year {{mn}} month {{dn}} day, etc., and finally import them into the system.
[0122] In a specific embodiment, the current OCR scanner has a paper loading capacity of 100 sheets, the scanning type is sheet-fed scanning, supports automatic paper feeding scanning, has a resolution of 600dpi, and a scanning speed of 45ppm / 90ipm. The OCR method recognition rate implemented based on this method can reach 88%. After data recognition, users are supported to check and modify the data, and after confirmation, they will be approved according to the approval process corresponding to the form.
[0123] The OCR-based document structuring method provided by the embodiment of the present invention obtains an image to be processed based on the acquired document to be processed, extracts local features and overall features of the image, performs weighted fusion of corresponding positions, and then performs text recognition on the fused features to obtain standardized document data. The local features extracted by this method contain implicit feature information, and combined with the overall features, can fuse information from multiple levels and angles, thereby improving the accuracy of characterization of text information in the image.
[0124] A practical method is proposed based on document structuring, which not only takes into account the influencing factors of manual correction errors and system recognition errors, but also requires confirmation by the approver before successful import, thereby improving the accuracy of text extraction, thereby improving the accuracy and reliability of generating standardized document data. Converting image data into computer internal code can greatly reduce the intensity of data entry work and increase the speed of data entry, providing a certain reference for the storage, management and subsequent query and analysis of construction process documents.
[0125] Embodiment 2
[0126] Based on the same inventive concept, the embodiment of the present invention also provides a document structuring device based on OCR, referring to Figure 7 As shown, the device comprises:
[0127] The first processing module 101 is used to process the acquired document to be processed to obtain an image to be processed;
[0128] A first division module 102, used for dividing the image to be processed into a plurality of blocks;
[0129] A first extraction module 103, used to input the plurality of blocks into a local feature extraction module respectively to obtain corresponding local features;
[0130] The second extraction module 104 is used to input the image to be processed into an overall feature extraction module to obtain overall features;
[0131] A first fusion module 105 is used to weightedly fuse the plurality of local features with the overall feature to obtain a data image feature;
[0132] A first recognition module 106, configured to perform text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data;
[0133] The first generating module 107 is used to generate standardized document data based on the unstructured text data.
[0134] Embodiment 3
[0135] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the OCR-based document structuring method described in the above embodiment 1 is implemented.
[0136] Embodiment 4
[0137] Based on the same inventive concept, an embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the OCR-based document structuring method described in the above-mentioned embodiment 1 is implemented.
[0138] Embodiment 5
[0139] Based on the same inventive concept, an embodiment of the present invention further provides a computer program product comprising instructions. When the computer program product is run on a computer device, the computer device executes the OCR-based document structuring method described in the above embodiment 1.
[0140] Embodiment 6
[0141] Based on the same inventive concept, an embodiment of the present invention further provides a chip, the chip includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run a computer program or instruction to implement the OCR-based document structuring method described in the above embodiment one.
[0142] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0143] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0144] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0146] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A document structuring method based on OCR, It is characterized in that include: Processing the acquired document to be processed to obtain an image to be processed; Dividing the image to be processed into a plurality of blocks; Inputting the multiple blocks into a local feature extraction module respectively to obtain corresponding local features; Inputting the image to be processed into an overall feature extraction module to obtain overall features; Weighted fusion of the plurality of local features and the overall feature to obtain a data image feature; Performing text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data; Generate standardized document data based on the unstructured text data.
2. The method according to claim 1, It is characterized in that The step of dividing the image to be processed into a plurality of blocks comprises: Clustering all pixels in the image to be processed to obtain a clustering result; According to the clustering result, the pixels belonging to the same category in the image to be processed are divided into a block to obtain the multiple blocks.
3. The method according to claim 1, It is characterized in that Each of the local feature extraction modules includes a plurality of convolutional layers and a plurality of depooling layers connected in sequence; The step of inputting the plurality of blocks into a local feature extraction module to obtain corresponding local features comprises: For each block, inputting the block into the plurality of sequentially connected convolutional layers to obtain a dimensionality reduction feature corresponding to the block; The dimension reduction features are input into the plurality of sequentially connected depooling layers to obtain local features corresponding to the blocks.
4. The method according to claim 1, It is characterized in that The step of weighting and fusing the plurality of local features with the overall feature to obtain the data image feature comprises: Matching each element feature in each of the local features with the overall feature to obtain corresponding positions of each element feature of the local feature in the overall feature; According to the corresponding position of each element feature of the local feature in the overall feature, each local feature is weightedly fused with the overall feature to obtain the data image feature.
5. The method according to claim 1, It is characterized in that The performing of text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data includes: Obtaining a plurality of feature points in the image to be processed according to the position of each feature of the data image features in the image to be processed; For each feature point, a circle with a radius equal to a first preset threshold is set with the feature point as the center; Determine whether the number of feature points in the circle is greater than a second preset threshold and all feature points in the circle belong to the same block: If yes, extracting texture information from the block where the feature point is located, performing text recognition based on the texture information, and obtaining a text recognition result; The unstructured text data is obtained according to all text recognition results.
6. The method according to claim 1, It is characterized in that The generating of standardized document data based on the unstructured text data includes: Extracting keywords according to the text queue of the unstructured text data, and obtaining a standardized document template corresponding to the keywords; Matching the standardized document template with the text queue to obtain an initialization document; Based on the initialization document, the standardized document material is obtained.
7. The method according to claim 6, It is characterized in that The step of obtaining the standardized document material based on the initialization document includes: Performing semantic analysis on the initialization document to obtain a semantic analysis result; Based on a preset industry terminology database, the semantic analysis result is matched to determine whether there is a mismatch between the semantic analysis result and the industry terminology database; If yes, a prompt message is returned, wherein the prompt message is used to prompt that the initialization document needs manual inspection; and the standardized document data is generated according to the manual modification result; If not, the initialization document is used as standardized document data.
8. The method according to claim 1, It is characterized in that Before dividing the image to be processed into multiple blocks, it also includes: Extracting texture features of the image to be processed to obtain texture features of the image to be processed.
9. A document structuring device based on OCR, It is characterized in that include: A first processing module is used to process the acquired document to be processed to obtain an image to be processed; A first division module, used for dividing the image to be processed into a plurality of blocks; A first extraction module, used for inputting the plurality of blocks into a local feature extraction module respectively to obtain corresponding local features; A second extraction module is used to input the image to be processed into an overall feature extraction module to obtain overall features; A first fusion module, used for weighted fusion of the plurality of local features and the overall feature to obtain a data image feature; A first recognition module, configured to perform text recognition based on the image to be processed, the data image features and the multiple blocks to obtain unstructured text data; The first generating module is used to generate standardized document data based on the unstructured text data.
10. A computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions are executed on a terminal, the terminal executes the OCR-based document structuring method according to any one of claims 1 to 8.
11. A computer device, It is characterized in that The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the OCR-based document structuring method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the OCR-based document structuring method according to any one of claims 1 to 8.
13. A chip, comprising a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run a computer program or instruction to implement the OCR-based document structuring method as described in any one of claims 1 to 8.