Image attachment auditing method and device based on OCR and NER models, and medium
By combining OCR and NER models to review image attachments, the problem of unstructured image attachments being unable to be automatically parsed has been solved, achieving full-process automation and efficient review of financial audits.
Patent Information
- Application Number
- CN202511060491.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-30
AI Technical Summary
The existing financial audit system cannot effectively parse unstructured image attachments, resulting in low approval efficiency and a high risk of errors, and it cannot achieve full-process automation.
By combining OCR and NER models, the textual structure and layout features of image attachments are identified, a pre-set feature extraction template is loaded, and a rule engine is used for automatic review to generate a review report.
It has achieved automated processing of multi-format image attachments, reduced manual intervention, improved review efficiency and reduced error rate, and realized full-process automation of financial review.
Smart Images

Figure CN120954037A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a method, device, and medium for reviewing image attachments based on OCR and NER models. Background Technology
[0002] As enterprises accelerate their digital transformation, financial process automation has become a core requirement for improving operational efficiency. Currently, structured data in financial systems, such as e-invoices and bank statements, can be highly automated for review through Optical Character Recognition (OCR) and rule engines. However, unstructured image attachments such as meeting attendance sheets, travel application forms, and purchase contracts cannot be automatically parsed by existing recognition models. This results in significant manual intervention still required in the financial review process, leading to low approval efficiency and hindering the achievement of full-process automation.
[0003] Existing solutions fall into two categories. One is template-based OCR recognition technology, which requires using predefined invoice / receipt templates to extract data by matching fixed field positions. However, OCR models are mostly designed for fixed-format invoices and cannot adapt to unstructured image attachments. The other approach involves collaboration between humans and a rule engine to recognize unstructured data. After key fields are manually entered, the rule engine verifies the logic. However, human efficiency is low, and errors are prone to occur during the review process, leading to a decrease in the accuracy of the review results. Summary of the Invention
[0004] To address the aforementioned issues, this application proposes an image attachment review method based on OCR and NER models, comprising:
[0005] Determine the file type corresponding to the image attachments uploaded by the user;
[0006] Based on the recognition patterns corresponding to each file type, image recognition is performed on the image attachments to extract the text structure features and layout features from the image attachments;
[0007] The text structure features are preprocessed, and the preprocessed text structure features are fused with the layout features to obtain fused features;
[0008] Load the preset feature extraction template, and extract the audit elements that match the feature extraction template from the fusion features based on the NER model;
[0009] The pre-built rule engine performs rule verification on the review elements to generate a review report corresponding to the image attachment, and then sends the review report back to the user.
[0010] In one implementation of this application, the file type includes structured documents and image files. Image recognition is performed on the image attachments to extract text structure features and layout features from the image attachments, specifically including:
[0011] When the file type is the structured document, the structured document is image-recognized using a PDF parser to obtain the text structure features and layout features corresponding to the structured document;
[0012] When the file type is an image file, the image file is preprocessed. The preprocessed image file is then subjected to OCR recognition using a preset feature recognition model to obtain the text structure features corresponding to the image file. Based on the text coordinates contained in the text structure features, the layout features corresponding to the image file are determined.
[0013] In one implementation of this application, before performing image recognition on the structured document using a PDF parser, the method further includes:
[0014] The python-docx library is used to parse the Word documents in the structured document to convert the Word documents into PDF documents.
[0015] In one implementation of this application, the layout features include at least table lines and heading levels, and the preprocessed text structure features are fused with the layout features, specifically including:
[0016] The preprocessed text structure features are fused with the layout features to generate the positional encoding of the text structure features relative to the layout features;
[0017] If the confidence level corresponding to the text structured feature does not meet the preset confidence level, the text structured feature is verified through the position encoding.
[0018] In one implementation of this application, a pre-built rule engine is used to perform rule validation on the review elements to generate a review report corresponding to the image attachment, specifically including:
[0019] The pre-built rule engine loads the review rules corresponding to the review elements.
[0020] The review elements are validated using the review rules to determine the review results corresponding to the review elements, and a review report corresponding to the image attachments is generated based on the review results corresponding to each review element.
[0021] In one implementation of this application, an audit report corresponding to the image attachment is generated based on the audit results corresponding to each audit element, specifically including:
[0022] Based on the audit results, identify the specific audit elements that failed the audit, and generate corresponding modification suggestions for the specific audit elements;
[0023] Based on the proposed modifications, each review element, and its corresponding review results, a review report corresponding to the image attachments is generated.
[0024] The review results will be fed back to the user, specifically including:
[0025] The review results and modification suggestions corresponding to the specified review elements are highlighted.
[0026] In one implementation of this application, the preprocessing of the text structure features specifically includes:
[0027] Based on a preset regular expression, invalid characters in the text structure features are filtered out; wherein, the invalid characters include at least one or more of the following: newline character, page break character, and control character;
[0028] Determine whether there are broken texts in the text structure features; if so, perform splicing on the broken texts.
[0029] In one implementation of this application, after extracting the audit elements that match the element extraction template from the fusion features based on the NER model, the method further includes:
[0030] Based on the review elements and the actual review elements contained in the image attachments, determine the accuracy and recall corresponding to the NER model;
[0031] Calculate the F1 score corresponding to the NER model based on the accuracy and the recall.
[0032] If the F1 value does not meet the preset value, the NER model is optimized.
[0033] This application provides an image attachment review device based on OCR and NER models, the device comprising:
[0034] At least one processor;
[0035] And, a memory communicatively connected to the at least one processor;
[0036] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform an image attachment review method based on an OCR and NER model as described above.
[0037] This application provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:
[0038] An image attachment review method based on OCR and NER models, as described in any of the preceding items.
[0039] The image attachment review method based on OCR and NER models proposed in this application can bring the following benefits:
[0040] By leveraging OCR technology and NER models, unified processing of multi-format image attachments is supported, breaking through fixed format limitations. The NER model dynamically extracts non-fixed review elements from unstructured image attachments, eliminating the need for manually predefined templates or input of key fields. Furthermore, through automatic review and verification of extracted review elements and a rule engine, significant manual intervention is reduced, improving review efficiency and lowering the error rate caused by manual operations, thus achieving full automation of the financial review process. Attached Figure Description
[0041] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0042] Figure 1 A flowchart illustrating an image attachment review method based on OCR and NER models provided in this application embodiment;
[0043] Figure 2 This is a schematic diagram of the structure of an image attachment review device based on OCR and NER models, provided for an embodiment of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0045] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0046] like Figure 1 As shown in the embodiment of this application, an image attachment review method based on OCR and NER models is provided, including:
[0047] S101: Determine the file type corresponding to the image attachment uploaded by the user.
[0048] Image attachments refer to image materials uploaded by users in business scenarios to supplement or prove relevant business activities. After users fill out forms on the front end, they need to upload image attachments in formats such as PDF, Word, PNG, and JPG to the system file server. After receiving the uploaded image attachments, the server will generate a unique identifier for each attachment. This unique identifier is used for quick file location and management. Simultaneously, the server will write the unique identifier back to the document database to associate it with the corresponding business document. The server supports batch uploading of image attachments and automatically checks the file format validity after uploading. This can be done by checking features such as file extensions and header information to determine if the file belongs to the system's supported format range. If it is invalid, the user is prompted to re-upload a compliant file. After confirming the file's validity, the server needs to determine the file type of the image attachment. The file type is used to distinguish whether the image attachment is structured data, including structured documents and image files. Structured documents include Word documents and PDF documents, while image files can be in formats such as PNG and JPG.
[0049] S102: Based on the recognition patterns corresponding to each file type, perform image recognition on the image attachments to extract the text structure features and layout features from the image attachments.
[0050] Based on the determined file type, the server will employ different recognition modes to perform image recognition on the image attachments, thereby extracting the text structure features and layout features from the attachments. Text structure features refer to textual characteristics, such as text content, text coordinates, page numbers, and confidence levels. Layout features represent the text page layout characteristics, including table lines and heading levels. Image recognition of image attachments to extract text structure features and layout features is a crucial step in achieving intelligent review. Extracting text structure features allows us to obtain the textual information within the attachments, while extracting layout features helps to reconstruct the original layout logic of the image attachments, more accurately grasp the context of the information, provide strong support for extracting key review elements, and also assist in judging the authenticity and completeness of review elements, ensuring the accuracy and reliability of the review process.
[0051] In one embodiment, for structured documents, before image recognition, the python-docx library is used to parse the Word documents within the structured document to convert them into PDF documents. Converting structured documents uniformly to PDF format effectively preserves key information such as text content and formatting. After uniformly formatting the structured documents, the PDF parser pdfplumber is used to perform image recognition, reading the binary data of the PDF document, parsing its text content and layout features, extracting text content according to the text's layout order, and forming continuous text structure features. Simultaneously, table areas are parsed to determine the table's row and column structure and the text content within cells, extracting the heading hierarchy and clarifying the position and chapter of each heading within the document.
[0052] For image files, preprocessing is performed first. Image processing libraries such as OpenCV are used to denoise the images, employing algorithms such as nonlocal mean to remove noise interference and make the images clearer and cleaner. Image tilt correction is then performed, using techniques such as Hough transform to correct image tilt caused during shooting, ensuring that elements such as text lines remain horizontal or vertical. Finally, image brightness equalization is performed, using algorithms such as CLAHE to adjust the brightness distribution of the image, enhancing the contrast between text and background and highlighting the text.
[0053] After the above preprocessing operations, the preprocessed image file is input into a preset feature recognition model for OCR recognition. In this embodiment, the feature recognition model can be a fusion model of CRNN and CTC. After scanning the image file, the feature recognition model can identify the text content and output it as text blocks with coordinates. Each text block contains the text content and its corresponding position coordinates in the image. The text content and its coordinate information together constitute the text structured features. At the same time, based on the text coordinates, the server can also determine the arrangement position of the text in the image, such as the area where the text is located, line spacing, column spacing, and other layout information, thereby obtaining the layout features corresponding to the image file. Both the text structured features and the layout features are in JSON format.
[0054] S103: Preprocess the text structure features and fuse the preprocessed text structure features with the layout features to obtain the fused features.
[0055] After extracting the text structure features and layout features from the image attachments, the text structure features need to be preprocessed to remove interfering information and improve text quality. The preprocessed text structure features then need to be further fused with the layout features to form fused features. Feature fusion combines text content with layout information, compensating for the shortcomings of relying solely on text or layout information. This allows for a more complete restoration of the document's true content and semantics, leading to a more accurate and in-depth understanding of the image attachments, thereby improving the accuracy of key element extraction and the reliability of the review results.
[0056] In one embodiment, the preprocessing of text structure features mainly includes two parts: filtering invalid characters and splicing broken text. Using a preset regular expression, the extracted text structure features are scanned to identify and remove invalid characters. Invalid characters include at least one or more of the following: line breaks, page breaks, and control characters. Filtering invalid characters in the text structure features makes the text cleaner and more concise. Furthermore, it is determined whether there is broken text in the text structure features due to formatting or recognition errors. If so, the broken text needs to be spliced. For example, when continuous text such as dates or amounts is incorrectly divided into multiple parts, by analyzing the semantic coherence and formatting characteristics of the text content, these broken text fragments are reassembled into complete text items, restoring their original integrity.
[0057] In one embodiment, during the feature fusion process of structured features and layout features, in addition to fusing the features as a whole, it is also necessary to combine the preprocessed text structured features with the positional information in the layout features to generate a positional code relative to the layout features for each text structured feature. That is, determining the row and column number of a text paragraph on the page, and which cell of which table it belongs to, etc., and using this positional information as supplementary features of the text content, so that it can reflect the specific position of the text on the page. For example, "Field A is located in the 3rd row and 2nd column of the table" is the positional code for field A.
[0058] Location coding can pinpoint the structural features of text and correct whether the text content conforms to the expected layout and logical relationships. When recognizing image attachments, the recognition model or algorithm provides the confidence score corresponding to the text's structural features, reflecting the reliability of the recognition result. Through the generated location code, the specific position of the text's structural features within the document can be quickly located. Combined with layout features, this allows for checking whether the text content conforms to the expected layout and logical relationships. For example, it checks whether the data in a table corresponds to the table header, and whether the title is in the appropriate position. If inconsistencies or contradictions are found between the text's structural features and layout features, further analysis of the reasons is needed, followed by correction or re-recognition.
[0059] S104: Load the preset feature extraction template, and extract the audit elements that match the feature extraction template from the fused features based on the NER model.
[0060] The fused features obtained through the above steps need to be input into the NER model (a fine-tuned BERT model) for identification of review elements. A pre-set feature extraction template is loaded, which defines the necessary review elements for different types of image attachments. For example, a "business trip application form" must include review elements such as "traveler, reason, and date." Based on this feature extraction template, the NER model performs semantic recognition on the fused features to extract review elements that match the template. The data input into the NER model is a fusion of text structure features and layout features. The final output of the NER model is the review elements in key-value pair format, such as {"traveler":"Zhang San","date":"2024-01-01"}.
[0061] In one embodiment, if the recognition accuracy of the aforementioned image attachments is low, further iterative training of the NER model is needed to improve its recognition accuracy. Furthermore, if the enterprise modifies or adds existing unstructured text formats, the model can also be further optimized and trained.
[0062] In determining model accuracy, the F1 score can be used to reflect it. Based on the review elements identified by the NER model and the actual review elements contained in the image attachments, the precision and recall of the NER model are determined. Precision is the ratio of the number of review elements correctly extracted by the model to the sum of all extracted results. Recall is the ratio of the number of review elements correctly extracted by the model to the total number of review elements actually present in the image attachments (including correctly extracted and missed elements). After calculating the precision and recall, the F1 score of the NER model is calculated using the following formula:
[0063]
[0064] Precision is the accuracy rate, and Recall is the recall rate.
[0065] If the F1 value is not lower than the preset value, it indicates that the NER model has poor recognition accuracy for this type of image attachment. In this case, it is necessary to continue to optimize the model based on this type of image attachment in order to improve the NER model's recognition accuracy for this type of image attachment.
[0066] During training, this application employs the cross-entropy loss function to measure the difference between the model's predicted values and the true values. The cross-entropy loss function is expressed as:
[0067]
[0068] Where y represents the true label of the sample. The label represents the model's prediction.
[0069] The loss function is optimized using the Adam optimizer, which is represented as follows:
[0070]
[0071] Where, θ t-1 It is the parameter value from the previous step; It is the correction value for the first moment, representing the average direction of the gradient; v t α is the correction value for the second moment, representing the variance of the gradient; α is the learning rate; ε is a very small constant (e.g., 10). -8 This is used to prevent division by zero during calculation updates.
[0072] When the cross-entropy loss is minimized, it indicates that the model has been optimized.
[0073] S105: Through a pre-built rule engine, the review elements are validated to generate a review report corresponding to the image attachments, and the review report is then sent back to the user.
[0074] The review elements need to be validated using a pre-built rule engine to confirm the compliance of all content in the image attachments. After the rule validation is completed, an overall review report for the image attachments is generated based on the validation results of each review element. The review report is then sent to the front end for user confirmation.
[0075] In one embodiment, a pre-built rule engine loads the audit rules corresponding to the audit elements. The rule engine selects the appropriate audit rules from a pre-built rule library based on the type of the audit element and business requirements. For example, for the "number of business trip days" audit element in a business trip application, the rule "business trip days ≤ 3 days require departmental approval" will be loaded. The extracted audit elements are input into the rule engine, which then matches and verifies each element according to the loaded audit rules. For example, the rule engine checks if the "number of business trip days" exceeds 3 days; if so, the audit element is deemed non-compliant. Based on the rule verification results, the audit result for each audit element is determined, including statuses such as "compliant," "non-compliant," and "requires further review." Based on the audit results for each audit element, an audit report corresponding to the image attachments is generated. The audit report details the audit results for each audit element and provides an overall audit conclusion.
[0076] In one embodiment, when generating an audit report, the first step is to filter out the specific audit elements that failed the audit based on the audit results. Then, based on these specific audit elements, corresponding modification suggestions are generated to provide users with clearer and more actionable guidance, helping them quickly resolve issues. For example, modification suggestions could include "Please resubmit the signed meeting minutes" or "The taxi ride was not completed within the specified time."
[0077] The modification suggestions, each review element, and their corresponding review results are integrated and generated into a complete review report according to a specific format and logic. The review report typically includes basic information about the image attachments, a list of review elements, the review results for each element, and corresponding modification suggestions, ensuring that users can fully and clearly understand the entire review process.
[0078] The review report is sent to the user, who can then view the review details of the image attachments. During the feedback process, to help users focus more intuitively on the parts that failed the review, the server will highlight the review results and modification suggestions corresponding to the specified review elements. This is usually done using eye-catching colors or special markers to highlight these important contents, guiding users to prioritize and address these issues.
[0079] The above are embodiments of the methods proposed in this application. Based on the same idea, some embodiments of this application also provide devices and non-volatile computer storage media corresponding to the above methods.
[0080] Figure 2 This is a schematic diagram of an image attachment review device based on OCR and NER models, provided as an embodiment of this application. Figure 2 As shown, it includes:
[0081] At least one processor; and,
[0082] At least one processor-communication-connected memory; wherein,
[0083] The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform an image attachment review method based on an OCR and NER model as described in any of the preceding claims.
[0084] This application provides a non-volatile computer storage medium storing computer-executable instructions, which are configured as follows:
[0085] An image attachment review method based on OCR and NER models, as described in any of the preceding items.
[0086] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.
[0087] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0088] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0089] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0090] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0091] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0092] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0093] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0094] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0095] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0096] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for reviewing image attachments based on OCR and NER models, characterized in that, The method includes: Determine the file type corresponding to the image attachments uploaded by the user; Based on the recognition patterns corresponding to each file type, image recognition is performed on the image attachments to extract the text structure features and layout features from the image attachments; The text structure features are preprocessed, and the preprocessed text structure features are fused with the layout features to obtain fused features; Load the preset feature extraction template, and extract the audit elements that match the feature extraction template from the fusion features based on the NER model; The pre-built rule engine performs rule verification on the review elements to generate a review report corresponding to the image attachment, and then sends the review report back to the user.
2. The image attachment review method based on OCR and NER models according to claim 1, characterized in that, The file types include structured documents and image files. Image recognition is performed on the image attachments to extract text structure features and layout features from the image attachments, specifically including: When the file type is the structured document, the structured document is image-recognized using a PDF parser to obtain the text structure features and layout features corresponding to the structured document; When the file type is an image file, the image file is preprocessed. The preprocessed image file is then subjected to OCR recognition using a preset feature recognition model to obtain the text structure features corresponding to the image file. Based on the text coordinates contained in the text structure features, the layout features corresponding to the image file are determined.
3. The image attachment review method based on OCR and NER models according to claim 2, characterized in that, Before performing image recognition on the structured document using a PDF parser, the method further includes: The python-docx library is used to parse the Word documents in the structured document to convert the Word documents into PDF documents.
4. The image attachment review method based on OCR and NER models according to claim 1, characterized in that, The layout features include at least table lines and heading levels. The preprocessed text structure features are integrated with the layout features, specifically including: The preprocessed text structure features are fused with the layout features to generate the positional encoding of the text structure features relative to the layout features; If the confidence level corresponding to the text structured feature does not meet the preset confidence level, the text structured feature is verified through the position encoding.
5. The image attachment review method based on OCR and NER models according to claim 1, characterized in that, The pre-built rule engine performs rule validation on the review elements to generate a review report corresponding to the image attachments, specifically including: The pre-built rule engine loads the review rules corresponding to the review elements. The review elements are validated using the review rules to determine the review results corresponding to the review elements, and a review report corresponding to the image attachments is generated based on the review results corresponding to each review element.
6. The image attachment review method based on OCR and NER models according to claim 5, characterized in that, Based on the review results corresponding to each review element, a review report corresponding to the image attachment is generated, specifically including: Based on the audit results, identify the specific audit elements that failed the audit, and generate corresponding modification suggestions for the specific audit elements; Based on the proposed modifications, each review element, and its corresponding review results, a review report corresponding to the image attachments is generated. The review results will be fed back to the user, specifically including: The review results and modification suggestions corresponding to the specified review elements are highlighted.
7. The image attachment review method based on OCR and NER models according to claim 1, characterized in that, The preprocessing of the text structure features specifically includes: Based on a preset regular expression, invalid characters in the text structure features are filtered out; wherein, the invalid characters include at least one or more of the following: newline character, page break character, and control character; Determine whether there are broken texts in the text structure features; if so, perform splicing on the broken texts.
8. The image attachment review method based on OCR and NER models according to claim 1, characterized in that, Based on the NER model, after extracting the audit elements that match the element extraction template from the fusion features, the method further includes: Based on the review elements and the actual review elements contained in the image attachments, determine the accuracy and recall corresponding to the NER model; Calculate the F1 score corresponding to the NER model based on the accuracy and the recall. If the F1 value does not meet the preset value, the NER model is optimized.
9. An image attachment review device based on OCR and NER models, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform an image attachment review method based on an OCR and NER model as described in any one of claims 1-8.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: A method for reviewing image attachments based on OCR and NER models as described in any one of claims 1-8.
Citation Information
Patent Citations
Consistency auditing method for different source files
CN109190092A
Document auditing method, device and system, equipment and storage medium
CN110852065A
Document layout analysis method, model training method and device and equipment
CN113361247A
Information extraction method and device
CN113961685A
Transaction background authenticity auditing method and system based on OCR and NLP technologies
CN114202755A