Intelligent commemorative inspection trial method based on OCR technology and NLP model

By applying OCR and NLP technologies in disciplinary inspection and review, the key information in the file text is automatically identified and extracted, and the problem of low efficiency of manual information filling in the existing technology is solved, and efficient and accurate document processing and data utilization are achieved.

CN119991011APending Publication Date: 2025-05-13数字广西集团有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411991053.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the disciplinary inspection and trial, the existing technology can only convert scanned images into semi-structured text, resulting in staff needing to manually extract information and fill in the marking record, which is very labor-intensive, inefficient and prone to errors.

Method used

The intelligent disciplinary inspection review method based on OCR technology and NLP model is adopted to identify the scan page of the file through the OCR model, generate semi-structured text, and use the NLP model to identify layout, character recognition and paragraph segmentation, extract key information and key fields, and automatically fill in the standard template for marking transcripts.

Benefits of technology

It improves the efficiency of document marking and documents processing, reduces the probability of human error, reduces the workload and difficulty of staff, and improves the utilization rate of file text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991011A_ABST
    Figure CN119991011A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a commemorative inspection intelligent trial method based on an OCR technology and an NLP model, and the method comprises the following steps: configuring a standard template of a paper marking record and a document; creating a case according to the information of the case, and scanning a paper file of the case to obtain a file scanning page of the case; s2, identifying the file scanning page in the step S2 through an OCR model to obtain a semi-structured file text; classifying file scanning pages corresponding to the file text through an NLP model to generate an index directory, and extracting key information and key fields associated with the key information to obtain processing data of the file text; s4, matching the key field in the step S4 with the preset field in the step S1, so as to fill the processed data into a position corresponding to the standard template; and processing the information of filling the data into the standard template, and performing sampling auditing and manual auditing. According to the invention, the processing efficiency of paper marking records and documents can be improved, and the probability of human errors is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of data processing technology, and in particular to an intelligent disciplinary inspection method based on OCR technology and NLP model. Background Art

[0003] In disciplinary inspection trials, the scanned pages of the entered case files are generally subjected to OCR recognition, thereby converting the scanned pages into semi-structured text. The disadvantage of converting into semi-structured text is that only the scanned images are converted into text and stored for archiving, and the staff is required to manually extract information from the case file materials and fill in the case reading minutes. The workload of the trial is large, it is difficult to effectively improve the efficiency of the staff, and it is easy to make mistakes. Summary of the invention

[0004] In order to solve the above problems, the present invention provides a disciplinary inspection intelligent trial method based on OCR technology and NLP model, which can improve the processing efficiency of examination records and documents and reduce the probability of human error.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A disciplinary inspection intelligent trial method based on OCR technology and NLP model includes the following steps:

[0007] S1. Configure standard templates for examination notes and documents;

[0008] S2. Create a case based on the case information, and scan the paper file of the case to obtain a scanned page of the case file;

[0009] S3. Recognize the scanned file page in step S2 by an OCR model to obtain a semi-structured file text;

[0010] S4. Performing layout recognition, character recognition and paragraph segmentation on the file text in step S3 through the NLP model, so as to classify the file scan pages corresponding to the file text to generate an index directory, and extracting key information and key fields associated with the key information from the file text to obtain processed data of the file text;

[0011] S5. Match the key fields of step S4 with the preset fields of step S1 to fill the processed data into the corresponding positions of the standard template to obtain the examination records and documents corresponding to the files;

[0012] S6. The information of the processed data filled into the standard template is sampled and reviewed, and the erroneous information is manually corrected.

[0013] Furthermore, in step S2, the case information includes the case name, file type, and case description.

[0014] Furthermore, in step S3, the OCR model construction method includes the following steps:

[0015] S3.1 Convert the scanned file page into a file image, annotate the file image with data using the labelimg tool, and use the label file in yolo format to determine the position of the layout area to obtain a data set, and divide the data set into a training set, a test set, and a validation set;

[0016] S3.2 performing enhancement processing on the image of the data set to obtain image data, normalizing the image data so that its pixel value range is between 0 and 1, and converting the image data and annotations into Tensor format;

[0017] S3.3 sets the model input size, batch size, learning rate parameters, and builds the YOLOv8 network structure to build the model framework;

[0018] S3.4 sets a loss function for the model framework, trains the model framework using the training set, adjusts the weights of different loss parts during the training process, and adjusts the hyperparameters of the learning rate and batch size according to the performance feedback of the validation set, and saves the model weights with the best performance to obtain a training model;

[0019] S3.5 evaluating the training model using the test set, and optimizing the training model according to the evaluation result to obtain an OCR model;

[0020] Furthermore, in step S3.5, the training model is evaluated using a test set to analyze and obtain performance indicators of the training model, the prediction results of the training model are analyzed to identify error types, and the training model is optimized based on the performance indicators and the errors of the error types.

[0021] Furthermore, in step S4, the key information includes text information of the subject information, case facts, evidence and case handling opinions.

[0022] Furthermore, the analysis of the NLP model includes the following steps:

[0023] S4.1 Identify the layout of the dossier text by using a layout detection algorithm based on YOLOv8 to determine the text area, header and footer, and title of the dossier text, and obtain a detection result; based on the detection result, select a page area with key information in the dossier text;

[0024] S4.2 presets the Schema list according to the required preset fields, and sets the recognition rules and recognition mode for each field through word segmentation and part-of-speech tagging;

[0025] S4.3 extracting the key fields associated with the key information from the page area according to the Schema list and by using the UIE model;

[0026] S4.4 Verify the key information and the key fields, and format the key information and the key fields.

[0027] Furthermore, the error information and correction information of step S6 are obtained, and the error information and the correction information are fed back to the NLP model.

[0028] Furthermore, in step S4.2, the preset fields include case name, case number, persons involved, amount involved, and evidence.

[0029] Furthermore, in step S5, the name of the evidence is matched with the index directory to associate the file page number where the evidence is located with the file scan page; the NLP model is used to extract keywords from the testimonies of each witness under the same evidence, and the semantic similarities of different testimonies are compared through semantic analysis to determine whether the facts described are consistent.

[0030] Furthermore, it also includes a trial system using the disciplinary inspection intelligent trial method, the trial system includes a business layer, an application layer and a support layer,

[0031] The business layer includes a user management module and a conversion module. The user management module is used for user registration and login; the conversion module is used for converting paper files into file scan pages;

[0032] The application layer includes a file creation module and a file processing module. The file creation module is used to create cases; the file processing module is used to call the support layer to generate the file reading records and documents corresponding to the files;

[0033] The support layer includes an identification module, an extraction module and a local file storage module. The identification module is used to identify the file scan pages and convert the file scan pages into semi-structured file texts; the extraction module is used to extract key information and key fields of the file text to obtain key information and its associated key fields; the local file storage module is used for data storage of the business layer, the application layer, the identification module and the extraction module.

[0034] The beneficial effects of the present invention are:

[0035] Scan the paper files of the case to obtain the file scan pages, and obtain the semi-structured file text through OCR model recognition. The OCR model is optimized for the layout recognition of the file, which can improve the recognition accuracy and recognition speed; the NLP model is used to perform word segmentation and grammatical analysis on the semi-structured file text data, and key information and key fields associated with the key information can be extracted from the file. According to the key information and key fields, the file processing data results are automatically filled into the standard template of the reading record, without manual filling, reducing the workload of the trial, improving the efficiency of the staff, and avoiding manual filling errors. In addition, by matching the name of the evidence with the index directory, the page number where the evidence is located can be associated with the corresponding file scan page. Through the NLP model, the factual contradictions in the description of the testimony of multiple witnesses of the same evidence fact can be detected to understand the deep meaning in the testimony. By comparing the semantic similarity between different testimonies, it can be judged whether they are consistent in describing the same fact, thereby realizing further mining and utilization of the file material data, improving the utilization rate of the file text data and reducing the workload of the case trial personnel, and reducing the difficulty of the staff's work. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flowchart of a disciplinary inspection intelligent trial method based on OCR technology and NLP model in a preferred embodiment of the present invention.

[0037] Figure 2 This is a structural diagram of a disciplinary inspection intelligent trial system based on OCR technology and NLP model in a preferred embodiment of the present invention.

[0038] In the figure, 1-business layer, 11-user management module, 12-conversion module, 2-application layer, 21-file creation module, 22-file processing module, 3-support layer, 31-identification module, 32-extraction module, 33-local file storage module. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0041] Please also see Figure 1 and Figure 2 A preferred embodiment of the present invention is a disciplinary inspection intelligent trial method based on OCR technology and NLP model, comprising the following steps:

[0042] S1. Configure standard templates for examination notes and documents.

[0043] S2. Create a case based on the case information, and scan the paper file of the case to obtain the case file scan page. In this embodiment, the paper file of the case is converted into the case file scan page by a high-speed scanner or a high-speed scanner.

[0044] In step S2, the case information includes the case name, file type, and case description.

[0045] S3. Use the OCR model to recognize the scanned file page in step S2 to obtain a semi-structured file text.

[0046] In step S3, the OCR model construction method includes the following steps:

[0047] S3.1 Convert the scanned file page into a file image, label the file image with the labelimg tool, and use the label file in yolo format to determine the position of the layout area to obtain a data set, and divide the data set into a training set, a test set, and a validation set. In this embodiment, the labeled data is cleaned to remove duplicate and unnecessary information.

[0048] S3.2 performs enhancement processing on the images of the data set to obtain image data, and normalizes the image data so that its pixel value range is between 0 and 1, and converts the image data and annotations into Tensor format.

[0049] In this embodiment, image enhancement operations such as rotation, scaling, cropping, color conversion, etc. are performed on the image to increase data diversity and improve the generalization ability of the model. The image data is normalized so that its pixel value range is between 0 and 1 to accelerate the convergence speed of model training. The Tensor format is the format required by the model.

[0050] S3.3 sets the model input size, batch size, and learning rate parameters, and builds the YOLOv8 network structure to build the model framework. The YOLOv8 network structure includes feature extraction layer, prediction layer, etc.

[0051] S3.4 sets the loss function for the model framework, uses the training set to train the model framework, adjusts the weights of different loss parts during the training process, and adjusts the hyperparameters of the learning rate and batch size based on the performance feedback of the validation set, and saves the model weights with the best performance to obtain the training model.

[0052] The loss function is a composite loss function, including classification loss, positioning loss and confidence loss. The weights of different loss parts are adjusted according to the actual situation to achieve the best training effect.

[0053] S3.5 evaluates the training model through the test set, and optimizes the training model according to the evaluation result to obtain the OCR model.

[0054] In step S3.5, the training model is evaluated using the test set to analyze and obtain the performance indicators of the training model, the prediction results of the training model are analyzed to identify the error type, and the training model is optimized based on the performance indicators and the errors of the error type.

[0055] Performance indicators include recognition accuracy; error types include false detection and missed detection; training model optimization includes adjusting the network structure, loss function or training strategy. In order to improve the model's reasoning speed and deployment efficiency, the model is pruned and quantized.

[0056] S4. Perform layout recognition, character recognition and paragraph segmentation on the file text in step S3 through the NLP model to classify the file scan pages corresponding to the file text to generate an index directory, and extract key information and key fields associated with it from the file text to obtain processed data of the file text.

[0057] In step S4, the information required to be filled in the examination record comes from the file material of the trial report. Through the layout analysis of the trial report, the scanned image is divided into paragraphs and character recognition is performed to extract the key information of the case: the name of the subject, education, position and criminal facts, case handling opinions and other text information; then, through the semantic analysis and grammatical analysis of the NLP model, the basic case information is matched with the preset corresponding fields to complete the association process of the key fields.

[0058] In step S4, the key information includes the subject information, case facts, evidence and text information of case handling opinions.

[0059] The analysis of the NLP model includes the following steps:

[0060] S4.1 identifies the layout of the dossier text through a layout detection algorithm based on YOLOv8 to determine the text area, header and footer, and title of the dossier text, and obtains the detection result; based on the detection result, selects the page area with key information in the dossier text;

[0061] S4.2 presets the Schema list according to the required preset fields, and sets the recognition rules and recognition patterns for each field through word segmentation and part-of-speech tagging. The fields are set through regular expressions and keyword lists to facilitate model recognition and information extraction.

[0062] The preset fields of this embodiment are based on the actual needs of the disciplinary inspection and supervision work, and the fields to be extracted are pre-defined. In step S4.2 of this embodiment, the preset fields include case name, case number, persons involved, amount involved, and evidence. In word segmentation and part-of-speech tagging, the continuous text is segmented into separate words or phrases and part-of-speech tagging based on the semi-structured text recognized by the OCR model to identify the part-of-speech meaning of each word.

[0063] S4.3 According to the Schema list, identify and extract key fields associated with key information in the UIE model extraction page area.

[0064] The UIE model can understand natural language and extract information according to the preset Schema. The UIE model automatically identifies and extracts corresponding key fields in the key information based on the preset Schema and page content.

[0065] S4.4 verifies key information and key fields, and formats the key information and key fields.

[0066] The data verification and correction in step S4.4 can clean the structured data and remove duplicate, erroneous or irrelevant data. Manual verification and correction are required to ensure the accuracy and completeness of the data.

[0067] In step S4.2 of this embodiment, the preset fields include case name, case number, persons involved in the case, amount involved in the case, and evidence.

[0068] After the key information and key fields are completed through the OCR model and NLP model, the data entry and page number association are manually reviewed to see if they are accurate and without omission. If the data is correct, proceed to step S5; if the data is incorrect, the user manually selects the file scan page and performs OCR recognition of the text again to fill in the standard template, and manually associates the page number where the evidence is located with the corresponding file scan page.

[0069] S5. Match the key fields of step S4 with the preset fields of step S1 to fill the processed data into the corresponding positions of the standard template to obtain the examination records and documents corresponding to the files;

[0070] In step S5, according to the standard templates of the examination record and documents, the correspondence between the pre-set processing data and the template, and the matching relationship between the blank spaces of the examination record template and the corresponding fields, the basic case information is automatically filled back into the examination record, thus realizing the corresponding filling of the file processing data results into the standard template of the examination record.

[0071] In step S5, the name of the evidence is matched with the index directory to associate the file page number where the evidence is located with the file scan page; the NLP model is used to extract keywords from the testimonies of each witness under the same evidence, and the semantic similarities of different testimonies are compared through semantic analysis to determine whether the facts described are consistent.

[0072] Detect factual contradictions in the testimonies of multiple witnesses on the same evidence. Extract keywords from each testimony, such as time, place, person, event, etc., and compare them. Use NLP models to perform semantic analysis to understand the deep meaning of the testimony. By comparing the semantic similarity between different testimonies, it can be determined whether they are consistent in describing the same fact.

[0073] S6. Process the data and fill it into the standard template for sampling and review, and manually correct the erroneous information. In this embodiment, the NLP model and the OCR model can be continuously iterated and optimized based on the feedback of the erroneous information to improve the accuracy of information extraction and the stability of layout detection.

[0074] It also includes a trial system using the disciplinary inspection intelligent trial method, the trial system includes a business layer 1, an application layer 2 and a support layer 3.

[0075] The business layer 1 includes a user management module 11 and a conversion module 12. The user management module 11 is used for user registration and login; the conversion module 12 is used to convert paper files into file scan pages. In this embodiment, you need to register an account and pass identity security authentication before use; or log in with an account that has been registered and authenticated.

[0076] The application layer 2 includes a file creation module 21 and a file processing module 22. The file creation module 21 is used for creating cases; the file processing module 22 is used to support the call of the layer 3 to generate the file reading records and documents corresponding to the file.

[0077] The support layer 3 includes an identification module 31, an extraction module 32 and a local file storage module 33. The identification module 31 is used to identify the file scan pages and convert the file scan pages into semi-structured file texts; the extraction module 32 is used to extract key information and key fields of the file text to obtain key information and its associated key fields; the local file storage module 33 is used for data storage of the business layer 1, application layer 2, the identification module 31 and the extraction module 32.

[0078] The application layer 2 of this embodiment is provided with a client, the file creation module 21 and the file processing module 22 are implemented in the client, and the business operation of the business layer 1 is sent to the application layer 2, the file processing request of the application layer 2 is sent to the support layer 3, the support layer 3 processes the data and sends it to the application layer 2, and the business execution result of the application layer 2 is sent to the business layer 3. The user can execute the business on the business layer through the client, so that the user can open the examination record and document management, and can preview, edit, export documents, print output and other operations on the examination record and documents.

Claims

1. A disciplinary inspection intelligent trial method based on OCR technology and NLP model, characterized in that: The steps include: S1. Configure standard templates for examination notes and documents; S2. Create a case based on the case information, and scan the paper file of the case to obtain a scanned page of the case file; S3. Recognize the scanned file page in step S2 by an OCR model to obtain a semi-structured file text; S4. Performing layout recognition, character recognition and paragraph segmentation on the file text in step S3 through the NLP model, so as to classify the file scan pages corresponding to the file text to generate an index directory, and extracting key information and key fields associated with the key information from the file text to obtain processed data of the file text; S5. Match the key fields of step S4 with the preset fields of step S1 to fill the processed data into the corresponding positions of the standard template to obtain the examination records and documents corresponding to the files; S6. The information of the processed data filled into the standard template is sampled and reviewed, and the erroneous information is manually corrected.

2. According to claim 1, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: In step S2, the case information includes the case name, file type, and case description.

3. According to claim 1, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: In step S3, the OCR model construction method includes the following steps: S3.1 Convert the scanned file page into a file image, annotate the file image with data using the labelimg tool, and use the label file in yolo format to determine the position of the layout area to obtain a data set, and divide the data set into a training set, a test set, and a validation set; S3.2 performing enhancement processing on the image of the data set to obtain image data, normalizing the image data so that its pixel value range is between 0 and 1, and converting the image data and annotations into Tensor format; S3.3 sets the model input size, batch size, learning rate parameters, and builds the YOLOv8 network structure to build the model framework; S3.4 sets a loss function for the model framework, trains the model framework using the training set, adjusts the weights of different loss parts during the training process, and adjusts the hyperparameters of the learning rate and batch size according to the performance feedback of the validation set, and saves the model weights with the best performance to obtain a training model; S3.5 The training model is evaluated using the test set, and the training model is optimized according to the evaluation result to obtain an OCR model.

4. According to claim 4, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: In step S3.5, the training model is evaluated using a test set to analyze and obtain performance indicators of the training model, the prediction results of the training model are analyzed to identify error types, and the training model is optimized based on the performance indicators and the errors of the error types.

5. According to claim 1, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: In step S4, the key information includes text information of the subject person, case facts, evidence and case handling opinions.

6. According to claim 5, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: The analysis of the NLP model includes the following steps: S4.1 Identify the layout of the dossier text by using a layout detection algorithm based on YOLOv8 to determine the text area, header and footer, and title of the dossier text, and obtain a detection result; based on the detection result, select a page area with key information in the dossier text; S4.2 presets the Schema list according to the required preset fields, and sets the recognition rules and recognition mode for each field through word segmentation and part-of-speech tagging; S4.3 extracting the key fields associated with the key information from the page area according to the Schema list and by using the UIE model; S4.4 Verify the key information and the key fields, and format the key information and the key fields.

7. According to claim 6, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: The error information and the correction information of step S6 are obtained, and the error information and the correction information are fed back to the NLP model.

8. According to claim 6, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: In step S4.2, the preset fields include case name, case number, persons involved, amount involved, and evidence.

9. According to claim 8, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: In step S5, the name of the evidence is matched with the index directory to associate the file page number where the evidence is located with the file scan page; The NLP model is used to extract keywords from the testimonies of each witness under the same evidence, and semantic similarities of different testimonies are compared through semantic analysis to determine whether the facts described are consistent.

10. According to claim 1, a disciplinary inspection intelligent trial method based on OCR technology and NLP model is characterized by: It also includes a trial system using the disciplinary inspection intelligent trial method, the trial system includes a business layer (1), an application layer (2) and a support layer (3), The business layer (1) comprises a user management module (11) and a conversion module (12), wherein the user management module (11) is used for user registration and login; and the conversion module (12) is used for converting paper files into file scan pages; The application layer (2) comprises a file creation module (21) and a file processing module (22), wherein the file creation module (21) is used for creating a case; and the file processing module (22) is used for calling the support layer (3) to generate a file reading record and documents corresponding to the file; The support layer (3) comprises an identification module (31), an extraction module (32) and a local file storage module (33), wherein the identification module (31) is used for identifying the file scan page and converting the file scan page into a semi-structured file text; the extraction module (32) is used for extracting key information and key fields of the file text to obtain key information and its associated key fields; the local file storage module (33) is used for data storage of the business layer (1), the application layer (2), the identification module (31) and the extraction module (32).

Citation Information

Cited By

  • Multi-modal large model-based law enforcement case handling document intelligent auditing system and method

    CN121212989A