License information extraction improvement method and system based on large model
Through the big model-based license information extraction method, combined with deep learning and license layout correction, the semantic understanding ability of the language big model is used to solve the problems of inaccurate and inefficient extraction of license information in traditional methods, and the rapid and accurate extraction of license information in the government system is achieved, and the efficiency of government services is improved.
Patent Information
- Application Number
- CN202510446930.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional methods of extracting certificate information are inefficient and are susceptible to factors such as light and angle, resulting in inaccurate and unstable information extraction.
The certificate information extraction method based on the big model is adopted, and the model is trained through deep learning and neural network structure, combined with the certificate layout correction and semantic understanding of the language big model, to achieve fast and accurate extraction of certificate information.
It significantly improves the accuracy and efficiency of the extraction of license information in the government affairs system and enhances the business efficiency of government affairs services.
Smart Images

Figure CN120340047A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information extraction method based on character recognition technology and a large model fine-tuned with license data, as well as a system implemented based on this method, which is used to extract key information when uploading license pictures and belongs to the field of image processing technology. Background Art
[0002] In the field of government services, with the increasing degree of informatization and digitalization, electronic pictures, as a carrier of information, are used more and more widely. For example, photos of various certificates are uploaded to the government affairs system as one of the supporting materials. Traditional license information extraction methods mainly rely on manual operations or simple character recognition technology. These methods are not only inefficient but also easily affected by factors such as lighting and angle, resulting in inaccurate and unstable information extraction.
[0003] With the breakthrough of deep learning and computer vision technologies, license information extraction methods based on large models have emerged. By using the labeled data and neural network structure of large models, a model with high recognition accuracy is trained to achieve fast and accurate extraction of license information, better serving the field of government services and improving the efficiency of business handling. Summary of the Invention
[0004] The object of the present invention is: in government services, to obtain the license pictures uploaded by users, extract the key data on the pictures, ensure the data is accurate and true as much as possible, and improve the efficiency of subsequent operations.
[0005] To achieve the above object, on the one hand, the present invention discloses an improved method for extracting license information based on a large model, which is characterized by including the following steps:
[0006] Step 1: The analysis module receives the picture sent by the external system;
[0007] Step 2: Send the decoded picture to the deep learning model classification module to obtain the license category of the current picture, and the analysis module queries the corresponding preset Prompt using the license category;
[0008] Step 3: The analysis module sends the decoded picture to the text extraction module, and the open-source character recognition model deployed by the text extraction module recognizes and extracts the characters in the picture;
[0009] Step 4: Send the license category calculated in Step 2 and the text content extracted in Step 3 to the text correction module, and the text correction module corrects the text content extracted from left to right and from top to bottom based on the license layout according to different license layouts;
[0010] Step 5: Send the text content corrected in Step 4 and the preset Prompt retrieved in Step 2 to the large model information extraction module, and use the semantic understanding ability of the large model to format the text content into a Key-Value JSON format.
[0011] Preferably, in Step 1, the picture is encoded in base64.
[0012] Another aspect of the present invention discloses an improved system for extracting certificate information based on a large model, which is characterized by being used to implement the above-mentioned improved method for extracting certificate information, including:
[0013] A deep learning training module, which is used to build a deep learning classification model, train the deep learning classification model, and deploy the trained deep learning classification model to the deep learning model classification module;
[0014] An analysis module, which is used for:
[0015] Receive the picture encoded in base64 sent by the calling party, decode the picture to obtain the decoded picture, and the analysis module sends the decoded picture to the deep learning model classification module and obtains the certificate classification information of the current picture;
[0016] Send the decoded picture to the text extraction module to extract the text on the picture;
[0017] Send the text fed back by the text extraction module and the certificate classification information fed back by the deep learning model classification module to the text correction module, and correct the scattered text recognition results according to the rules predefined based on different certificate formats;
[0018] Send the corrected result returned by the text correction module, the certificate type returned by the deep learning classification module, and the information extraction Prompt of the current type queried by the certificate type to the large model information extraction module, and return the finally extracted key information to the corresponding calling party;
[0019] A deep learning model classification module, which is used to use the trained deep learning classification model to discriminate the specific category of the certificate and return the result to the analysis module;
[0020] A text extraction module, which is used to deploy an open-source text recognition model to extract all the text positions and text contents in the certificate picture and return the extraction result to the analysis module;
[0021] A text correction module, which uses the rules formulated based on the prior knowledge of different certificate image formats to correct the text content and returns the corrected result to the analysis module;
[0022] The large model fine-tuning module uses the LoRA technology and combines it with an automatically generated dataset to fine-tune and train the large model;
[0023] The large model information extraction module uses the corrected text content extracted by the text correction module and the Prompt queried by the analysis module to complete the extraction of key text information and return it to the analysis module.
[0024] The present invention discloses a method and system for extracting license information based on a large model. By using the labeled data and neural network structure of the large model, a model with high recognition accuracy is trained to achieve fast and accurate extraction of license information, better serving the field of government affairs services and improving the efficiency of business handling. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart for extracting license photo information according to an embodiment of the present invention;
[0026] Figure 2 is a schematic structural diagram of a license photo information extraction system according to an embodiment of the present invention;
[0027] Figure 3 illustrates the structure of the residual module. DETAILED DESCRIPTION OF THE INVENTION
[0028] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0029] As Figure 1 shown, one aspect of the embodiment of the present invention is to disclose an improved method for extracting license information based on a large model, specifically including the following steps:
[0030] Step S101: The analysis module receives a base64-encoded picture sent by another system.
[0031] Step S102: After decoding, the obtained picture is sent to the deep learning model classification module to obtain the license category of the picture, and the analysis module queries the corresponding preset Prompt using this license type.
[0032] Step S103: The analysis module sends the decoded picture to the text extraction module, and the open-source text recognition model deployed by the text extraction module recognizes and extracts the text in the picture.
[0033] Step S104: Send the certificate type calculated in step S102 and the text content extracted in step S103 to the text correction module. The text correction module corrects the text content extracted from left to right and top to bottom based on different certificate layouts.
[0034] In an example of the present invention, the present invention realizes configurable correction of text recognition results based on certificate layouts. After correction, the readability of the text is significantly enhanced compared to before correction, making it easier for the large model to extract key text information. For example, taking the recognition results of some passports of the People's Republic of China as an example, the text content extracted by the text extraction model is: "Type of Passport of the People's Republic of China / Type Country Code / Country Code Passport Number / PassportNo P CHNE99999999 Name / Name PASSPORT Wu, Mou Gender / Sex / Nationality Date of Birth / Dateofbirth Male / M". The text content is arranged from left to right and top to bottom, so the data directly extracted by the text extraction module has poor readability and the effect of semantic understanding using the large model is also poor. Therefore, after analyzing the passport layout, it can be seen that the fields and field values in the passport data are arranged vertically. Therefore, in the text correction module, this type of certificate can be configured for vertical correction. The corrected content is: "Type of Passport of the People's Republic of China / Type P Name / Name Wu, Mou Gender / Sex Male / M Date of Birth / Dateofbirth 01 01 Month / Jan 2000 Country Code / Country Code CHN Passport Number / PassportNo E99999999". The readability of the corrected content is significantly enhanced. After testing, the effect of using the large model to extract key text information is also significantly enhanced.
[0035] Step S105: Send the text content corrected in step S104 and the Prompt queried in step S102 to the large model information extraction module, and use the semantic understanding ability of the large model to format the text content into a JSON format of Key-Value.
[0036] In an example of the present invention, a typical Prompt is: "This is a certificate of XX type, and the text recognized is as follows: 'XXXXXXX'. Please extract the following content 'Field A, Field B, Field C...', format the extracted content into JSON format, do not forge any content, and only return the completed JSON string in formatted form."
[0037] The improved method for extracting certificate information based on large models in the embodiments of the present invention aims to significantly enhance the ability to extract key information from certificate pictures in various government affairs system processes. By improving the existing process of simply using text content extraction, this method introduces text content correction based on the certificate format and Prompts separately set for different certificates. At the same time, leveraging the semantic understanding ability of large language models, it further improves the ability to extract information from certificate pictures. These improvement measures can extract information from certificate pictures more accurately and quickly, effectively improving the efficiency of information acquisition in government affairs system processes.
[0038] To implement the above embodiments, another aspect of the embodiments of the present invention is to propose a certificate information extraction system 210, including a deep learning model training module 211, an analysis module 212, a text extraction module 213, a text correction module 214, a deep learning model classification module 215, a large model information extraction module 216, and a large model fine-tuning module 217. Next, these modules will be introduced in detail.
[0039] I. Deep learning training module 211
[0040] The deep learning model training module 211 is mainly responsible for the construction and training of the deep learning classification model structure. This deep learning model mainly includes a residual module and a loss module.
[0041] 1) Residual module, the main structure of this module is as Figure 3 shown.
[0042] This module mainly solves the problem of the decline in model accuracy when the number of model layers is too deep. This module introduces a shortcut-connection, which can still backpropagate gradients even when the number of model layers is very large, solving the problem of gradient disappearance. The mathematical representation of this module is as follows:
[0043] y = σ(F(x, W) + x)
[0044] Among them, y represents the output of the residual block, σ() represents the activation function, F() represents the residual function, x represents the input (i.e., a three-channel picture), and W represents all the weights within the residual block.
[0045] 2) Loss module
[0046] The cross-entropy function is used to calculate the loss of the model, as shown in the following formula:
[0047]
[0048] Among them: y i represents the label of sample i, the positive class is 1, and the negative class is 0; p iIt represents the probability that sample i is predicted as the positive class; N represents the number of samples.
[0049] After the deep learning model 211 is constructed, a legal and private dataset is used to train the deep learning model 211, and then the deep learning model 211 can be deployed to the deep learning model classification module 214 for use.
[0050] II. Analysis Module 212
[0051] In an embodiment of the present invention, the analysis module 212 receives the base64-encoded picture sent by the caller, decodes the picture to obtain the decoded picture, and the analysis module 212: sends the decoded picture to the deep learning model classification module 215 and obtains the classification of the picture; sends the decoded picture to the text extraction module 213 to extract the text on the picture; sends the text fed back by the text extraction module 213 and the license type classification information fed back by the deep learning model classification module 215 to the text correction module 214 to correct the scattered text recognition results based on the layout according to the business rules of different licenses; sends the corrected text result returned by the text correction module 214, the license type returned by the deep learning discrimination module 215, and the information extraction Prompt of this type queried by the license type to the large model information extraction module 216 to obtain the finally extracted key information and return it to the corresponding caller.
[0052] III. Text Extraction Module 213
[0053] In an embodiment of the present invention, the text extraction module 213 is responsible for deploying an open-source text recognition model, which includes text position detection and text content detection.
[0054] IV. Text Correction Module 214
[0055] In an embodiment of the present invention, the text correction module 214 is responsible for correcting the text extracted by the text extraction module 213 based on the license layout. The algorithm of this module mainly includes the following steps:
[0056] Step 1: Set different thresholds for different license types, and this threshold controls the spacing when merging text content.
[0057] Step 2: Determine whether the license type is in horizontal layout or vertical layout.
[0058] Step 3.1: Horizontally, set the threshold. First, arrange the text rectangular boxes in the order from top to bottom, and align the text in the X-axis direction within the threshold range.
[0059] Step 3.2: Vertically, set thresholds. First, perform correction on the X-axis to distinguish left and right information, then sort. Next, perform threshold correction on the Y-axis to correct up and down misaligned information. Finally, sort again to obtain the final result.
[0060] IV. Deep learning model classification module 215
[0061] In an embodiment of the present invention, the deep learning model classification module 215 uses a trained deep learning model to detect the types of certificates, including the following steps:
[0062] Step 1: Receive the certificate image sent by the analysis module 212 and detect the corresponding type.
[0063] Step 2: Return the certificate type to the analysis module 212.
[0064] V. Large model fine-tuning module 217
[0065] In an embodiment of the present invention, the large model fine-tuning module 217 is responsible for the generation of the dataset and the fine-tuning training of the large model based on the LoRA (Low-Rank Adaptation of Large Language Models) technology. For the dataset generation part, taking the passport of the People's Republic of China as an example, the steps are as follows:
[0066] Step 1: Use photos of the passport at different angles to extract a number of original text contents by text extraction technology.
[0067] Step 2: Mask the field values in the extracted original texts. For example, mask the name Wu Moumou as name {name value}.
[0068] Step 3: Use random values to randomly replace the key contents. For example, replace name {name value} with name Liu Moumou, and replace date of birth {date of birth value} with date of birth 1990 / 01 / 01, etc. Since the data formats of different types of certificates are different, such as the date format being year / month / day, year-month-day, month / day / year, etc., random values need to be generated in the corresponding format according to different certificate types.
[0069] Step 4: Based on the randomly generated data items, assemble them into a training data item, and the data format is {"instruction": "Please extract the name, date of birth... from the given input and return it in JSON format", "input": {the data items randomly generated in Step 3}, "output": {the key information in JSON format}}.
[0070] After the dataset is generated, use the LoRA-based technology to fine-tune and train the large model (the llama2 model in the embodiments of the present invention). The LoRA technology will first freeze the large model parameters and not perform full-parameter fine-tuning training. Add the residual form to the large model parameters, and complete the optimization process of the large model parameters by training Δw:
[0071] w′ = w o + Δw
[0072] Among them, d is the rank of the parameter Δw.
[0073] VI. Large model information extraction module 216
[0074] In an embodiment of the present invention, the large model information extraction module 216 deploys the fine-tuned llama2 model. After receiving the license type and the assembled Prompt sent by the analysis module 212, the large model analyzes and returns the key information extraction result.
Claims
1. An improved method for extracting license information based on large models, characterized in that, It includes the following steps: Step 1: The analysis module receives the picture sent by the external system; Step 2: Send the decoded picture to the deep learning model classification module to obtain the license type of the current picture, and the analysis module queries the corresponding preset Prompt using the license type; Step 3: The analysis module sends the decoded picture to the text extraction module, and the open-source text recognition model deployed by the text extraction module recognizes and extracts the text in the picture; Step 4: Send the license type calculated in Step 2 and the text content extracted in Step 3 to the text correction module, and the text correction module corrects the text content extracted from left to right and top to bottom based on the license format; Step 5: Send the text content corrected in Step 4 and the preset Prompt queried in Step 2 to the large model information extraction module, and use the semantic understanding ability of the large model to format the text content into a Key-Value JSON format.
2. The improved method for extracting license information based on a large model according to claim 1, characterized in that, In Step 1, the picture is encoded in base64.
3. An improved system for extracting license information based on a large model, characterized in that, To implement the improved method for license information extraction as described in Claim 1, it includes: A deep learning training module, which is used to build a deep learning classification model, train the deep learning classification model, and deploy the trained deep learning classification model to the deep learning model classification module; An analysis module, which is used for: Receiving the picture encoded in base64 sent by the calling party, decoding the picture to obtain the decoded picture, sending the decoded picture to the deep learning model classification module and obtaining the license classification information of the current picture; Sending the decoded picture to the text extraction module to extract the text on the picture; Sending the text feedback by the text extraction module and the license classification information feedback by the deep learning model classification module to the text correction module, and correcting the scattered text recognition results according to the predefined rules based on different license formats; Sending the corrected result returned by the text correction module, the license type returned by the deep learning classification module, and the information extraction Prompt of the current type queried through the license type to the large model information extraction module, and returning the finally extracted key information to the corresponding calling party; A deep learning model classification module, which is used to use the trained deep learning classification model to determine the specific category of the license and return the result to the analysis module; A text extraction module, which is used to deploy an open-source text recognition model to extract all the text positions and text content in the license picture and return the extraction result to the analysis module; A text correction module, which uses the rules formulated based on the prior knowledge of different license image formats to correct the text content and returns the correction result to the analysis module; A large model fine-tuning module, which uses the LoRA technology and combines the automatically generated dataset to fine-tune and train the large model; A large model information extraction module, which uses the corrected text content extracted by the text correction module and the Prompt queried by the analysis module to complete the extraction of the text key information and return it to the analysis module.