Lightweight-based document key information extraction method and device, equipment and storage medium
By preprocessing and extracting features from document images using a lightweight text detection and classification model, the problem of low efficiency in extracting key information from documents is solved, achieving efficient and low-cost text content extraction, which is suitable for resource-constrained environments.
Patent Information
- Application Number
- CN202510270422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing technologies are inefficient and computationally expensive when extracting key information from documents, especially in resource-constrained environments where large and complex deep learning models face performance bottlenecks.
A lightweight text detection and classification model is adopted. By preprocessing the initial document image, including cropping, denoising and resolution optimization, a joint loss function of feature extraction and multi-task learning is trained using a lightweight neural network to achieve efficient extraction of text location and classification information.
It reduces computational resource consumption and improves inference speed while ensuring accuracy, making it suitable for applications with high real-time requirements and resource-constrained environments.
Smart Images

Figure CN120164225B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular to a lightweight-based document key information extraction method, device and equipment and storage medium. BACKGROUND
[0002] With the rapid development of information technology, image recognition technology has been widely used in document processing, information extraction and other fields. In the field of document management, especially in cadre archives management, the extraction of key information of fixed format "ten types" table has become an important link to improve work efficiency. The traditional manual input method is not only inefficient, but also prone to human error, which cannot meet the needs of modern archives management.
[0003] Therefore, how to efficiently extract the key information of the document is a technical problem to be solved at present. SUMMARY
[0004] The present application provides a lightweight-based document key information extraction method, device and equipment and storage medium, which solves the defect of low efficiency in extracting document key information in the prior art, realizes the determination of the document text information to be recognized in advance through the text detection classification model, and can efficiently, low-cost and accurately extract the target text content of the document.
[0005] In the first aspect, the present application provides a lightweight-based document key information extraction method, comprising the following steps:
[0006] Obtaining an initial document image of an initial document, and pre-processing the initial document image to determine a pre-processed document image;
[0007] Inputting the pre-processed document image into a trained text detection classification model for processing, and outputting text position information and corresponding text classification information;
[0008] Based on the text position information and the text classification information, performing text recognition processing on the initial document to obtain target text content corresponding to the initial document;
[0009] Wherein, the pre-processing at least includes image form processing and image classification processing; the text detection classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0010] Preferably, according to the lightweight-based document key information extraction method provided by the present application, the pre-processing of the initial document image to determine the pre-processed document image comprises:
[0011] Image form processing is performed on the initial document image to obtain a standard document image;
[0012] identifying a document layout type of the standard document image;
[0013] performing image classification processing on the standard document image based on the document layout type, to obtain the preprocessed document image after image classification.
[0014] Preferably, according to the lightweight-based document key information extraction method provided by the application, the image form processing on the initial document image to obtain a standard document image comprises:
[0015] performing image cropping processing on the initial document image to obtain a cropped document image;
[0016] performing image denoising processing on the cropped document image to obtain a denoised document image;
[0017] performing resolution optimization processing on the denoised document image to obtain the standard document image.
[0018] Preferably, according to the lightweight-based document key information extraction method provided by the application, after the step of performing text recognition processing on the initial document based on the text location information and the text classification information to obtain the target text content corresponding to the initial document, the method comprises:
[0019] determining a text mapping relationship between the text classification information and the target text content;
[0020] based on the text mapping relationship, structurally storing the text classification information, the text location information and the target text content.
[0021] Preferably, according to the lightweight-based document key information extraction method provided by the application, the determining step of the trained text detection classification model comprises:
[0022] performing feature extraction processing on the document image training sample to extract a multi-level feature map;
[0023] fusing feature maps of different levels to obtain a fused feature map, and updating the fused feature map based on a classification feature map for calculating a text box category to generate a detection classification feature map;
[0024] performing convolution processing and branch processing on the detection classification feature map in sequence to obtain an optimized probability map, a threshold value map and a classification map;
[0025] based on the probability map, the threshold value map, the classification map and a joint loss function, performing multiple learning and training processing on the lightweight neural network to obtain a trained text detection classification model.
[0026] Preferably, the lightweight-based document key information extraction method provided by the present application comprises a joint loss function composed of a detection loss and a joint loss, and the formula of the joint loss function is as follows:
[0027]
[0028] wherein L is the joint loss function, Lc represents a loss corresponding to the classification map, is a loss of the threshold map, is a loss corresponding to the approximate binary map, is a loss of the probability map, , is a constant value; the approximate binary map is obtained by performing differentiable binary calculation on the threshold map and the probability map.
[0029] In a second aspect, the present application further provides a lightweight-based document key information extraction device, comprising:
[0030] a preprocessing module configured to obtain an initial document image of an initial document, and perform preprocessing on the initial document image to determine a preprocessed document image;
[0031] a text detection and classification module configured to input the preprocessed document image into a trained text detection and classification model for processing, and output text position information and corresponding text classification information;
[0032] a text recognition module configured to perform text recognition processing on the initial document based on the text position information and the text classification information to obtain target text content corresponding to the initial document; wherein the preprocessing at least includes image form processing and image classification processing; and the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0033] In a third aspect, the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the lightweight-based document key information extraction method according to any one of the above aspects when executing the program.
[0034] In a fourth aspect, the present application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the lightweight-based document key information extraction method according to any one of the above aspects.
[0035] In a fifth aspect, the present application further provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the lightweight-based document key information extraction method according to any one of the above aspects.
[0036] The application provides a lightweight-based document key information extraction method, device and equipment and a storage medium. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0038] Figure 1 Fig. 1 is one of the flow diagrams of the lightweight-based document key information extraction method provided by the application.
[0039] Figure 2 Fig. 2 is a structural diagram of the lightweight-based document key information extraction device provided by the application.
[0040] Figure 3 Fig. 3 is a structural diagram of the electronic equipment provided by the application. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the application more clear, the technical solutions in the application will be clearly and completely described below in combination with the drawings in the application. Obviously, the described embodiments are some embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the application.
[0042] In the related art, at least the following technical problems exist:
[0043] Dry file, for personnel management and decision support is of great significance. Traditionally, such archives are saved in paper form, with the advancement of e-government, more and more archives are digitized. However, simple scanning cannot solve the problem of information retrieval and analysis.
[0044] In the related art, the current mainstream key information extraction adopts a document understanding and processing model, such as LayoutLMv3, which combines text and document image visual information, encodes text content and layout information through a unified Transformer architecture, and performs well in tasks such as table parsing, bill recognition, and document classification that need to consider text and layout information. Although this method can also achieve good results, it is not lightweight enough in actual training and deployment, and different models are needed for multiple structures of documents, and the space occupied by too many structures is also relatively large, requiring high performance of the deployment machine.
[0045] In practical applications, especially in resource-constrained environments, deploying large and complex deep learning models may encounter performance bottlenecks. Therefore, it is crucial to research and develop lightweight models.
[0046] The following will be combined Figures 1-3 A lightweight-based document key information extraction method, device, equipment and storage medium are described to solve the defects of low efficiency and high calculation cost in extracting document key information in the prior art, and to realize efficient and low-cost extraction of target text content of the document by determining the document text information to be recognized in advance through a text detection classification model. The lightweight text detection classification model not only reduces the consumption of computing resources under the premise of ensuring accuracy, but also speeds up the inference speed, and is suitable for application scenarios with high real-time requirements.
[0047] Figure 1 is one of the process flow diagrams of the lightweight-based document key information extraction method provided by the present application, as Figure 1 shown, the method can include but is not limited to steps S100-S300:
[0048] S100, obtaining an initial document image of an initial document, and preprocessing the initial document image to determine a preprocessed document image;
[0049] S200, inputting the preprocessed document image into a trained text detection classification model for processing, and outputting text position information and corresponding text classification information;
[0050] S300, based on the text position information and the text classification information, performing text recognition processing on the initial document to obtain target text content corresponding to the initial document;
[0051] The preprocessing at least includes image form processing and image classification processing; and the text detection classification model is obtained by training a light neural network using a large number of document image training samples.
[0052] In step S100 of some embodiments, an initial document image of an initial document is acquired, and the initial document image is preprocessed to determine a preprocessed document image.
[0053] It can be understood that the initial document can be imaged by an image acquisition terminal to acquire the initial document image containing the initial document.
[0054] The image acquisition terminal includes, but is not limited to, a smart phone, a camera, a document scanner, and other smart terminals with image acquisition functions.
[0055] Further, the electronic document can also be converted into an image document.
[0056] Since the initial document images obtained in different forms have many differences, for example, some initial document images obtained by shooting with a mobile phone are not clear enough, or the initial document in the obtained initial document image is not located in the middle position of the initial document image, and there are format differences between the initial document image obtained by a scanner and the initial document image obtained by a smart phone.
[0057] Therefore, the obtained initial document image needs to be preprocessed to determine a preprocessed document image with a unified format, so as to improve the text recognition efficiency and text recognition accuracy.
[0058] It should be noted that the preprocessing at least includes image form processing and image classification processing.
[0059] In some embodiments of the present application, the preprocessing of the initial document image to determine a preprocessed document image includes:
[0060] The initial document image is subjected to image form processing to obtain a standard document image.
[0061] The initial document image is subjected to image form processing to obtain a standard document image.
[0062] The initial document image is subjected to image cropping processing to obtain a cropped document image.
[0063] The cropped document image is subjected to image denoising processing to obtain a denoised document image.
[0064] The denoised document image is subjected to resolution optimization processing to obtain the standard document image.
[0065] Specifically, the initial document image is subjected to image cropping, denoising and resolution optimization processing to improve the quality and readability of the document image, thereby improving the recognition efficiency of the subsequent text recognition processing of the document image.
[0066] Further, the unnecessary background, frame or other irrelevant areas in the initial document image are removed, and the part containing the main content of the document is focused on, so that the main body of the document is more prominent, facilitating subsequent processing and analysis, while reducing the image data volume and improving the processing efficiency.
[0067] Further, image edge detection algorithms such as Canny operator and Sobel operator can be used to automatically identify the edges of the content and background in the document image, and the cropping area can be determined according to the edge information. This method has good effect on images with complex background and low contrast between document content and background, and can more accurately capture the outline of the document content to achieve automatic cropping.
[0068] Needless to say, if the document image has fixed format or proportion requirements, the image can be cropped according to the preset proportion or size parameters.
[0069] In some embodiments, the cropped document image is subjected to image denoising processing to obtain a denoised document image, which is used to eliminate noise interference generated in the image acquisition, transmission or storage process. These noises may appear as random spots, lines or blur, which can affect the clarity and accuracy of the document image and reduce the image quality. Denoising processing can smooth the image, making the text and graphics clearer and more distinguishable, and improving the visual effect and readability of the image.
[0070] The specific image denoising steps can be implemented by the following methods:
[0071] Mean filter denoising method: the average value of the pixel values in the neighborhood of each pixel point in the image is taken to replace the value of the pixel point, thereby smoothing the image and achieving the purpose of denoising.
[0072] Gaussian filter denoising method: a weighted mean filter method, which determines the weight of each pixel point in the neighborhood according to the Gaussian function, and the closer the pixel point to the center pixel point, the greater the weight. Then the weighted average value is calculated as the new value of the center pixel point. Gaussian filter can remove noise while preserving the edge and detail information of the image.
[0073] Median filter denoising method: the pixel values in the neighborhood of each pixel point in the image are sorted, and the middle value after sorting is taken as the value of the pixel point. Median filter is very effective in removing salt and pepper noise (i.e. black and white spot noise), which can effectively remove noise interference while protecting the edges of the image, making the image clearer.
[0074] In some embodiments, the denoised document image is subjected to resolution optimization processing to obtain the standard document image. The resolution of the document image is adjusted to reach a standard resolution suitable for specific needs. If the resolution is too low, the image may be blurred and it is difficult to identify the text and content therein; if the resolution is too high, the data volume and storage space occupation of the image will be increased, and the display capacity of the display device may also be exceeded. Resolution optimization can ensure the clear display and print output of the document image on different devices, while balancing the relationship between image quality and file size.
[0075] The specific resolution optimization processing steps can be determined according to actual needs and can include but are not limited to the following steps:
[0076] Upsampling to improve resolution: using interpolation algorithms to generate new pixel points based on the original image. Common interpolation algorithms include nearest neighbor interpolation, bilinear interpolation and cubic convolution interpolation, etc. Nearest neighbor interpolation algorithm is simple and fast, but may produce obvious jagged phenomenon; bilinear interpolation algorithm can improve the image quality to some extent, making the interpolated image smoother; cubic convolution interpolation algorithm can better preserve the detail information of the image and generate higher quality high-resolution image, but the calculation amount is relatively large.
[0077] Downsampling to reduce resolution: if the image resolution is too high, the pixel number can be reduced and the resolution can be reduced by downsampling. Downsampling can extract pixel points according to a certain ratio in the horizontal and vertical directions, for example, every other pixel point is extracted, so that the resolution of the image is reduced by half. In the downsampling process, in order to reduce information loss, low-pass filtering processing can be performed on the image first to remove high-frequency information, and then downsampling operation is performed.
[0078] Through the above steps of cropping, denoising and resolution optimization processing of the initial document image, a high-quality standard document image can be obtained.
[0079] The document format type of the standard document image is identified, and the standard document image is subjected to image classification processing according to the document format type of the standard document image, so as to improve the efficiency of text recognition and reduce the calculation cost of the recognition model.
[0080] It should be noted that the document format type includes but is not limited to: table format, book format, report format, letter format, poster format, etc.
[0081] The standard document image is subjected to image classification processing based on the document format type to obtain the preprocessed document image after image classification.
[0082] It should be noted that for each document format type, a corresponding template image or feature model is established. The pre-processed document image is compared and matched with these templates, and the similarity is calculated. For example, for the template of book format, the characteristics of its chapter structure, font size range, etc. can be defined, and when the similarity between the image to be classified and the book format template reaches a certain degree, it is classified as a book format.
[0083] Further, a large number of document images with annotated format types are used as training data sets to train machine learning models (such as neural networks, support vector machines, etc.). By letting the image classification model learn the feature patterns of different format images, when a new document image is input, the image classification model can classify it according to the learned knowledge. For example, a convolutional neural network (CNN) model is trained, and after the document image is input into the image classification model, the image classification model outputs the document format category to which the image most likely belongs.
[0084] In some embodiments of the present application, the classified document images are labeled with corresponding format types for subsequent retrieval and management. At the same time, according to different application scenarios, the classified images can be stored in specific folders or databases, classified and stored according to the format type.
[0085] In some embodiments of the present application, the classified document images are labeled with corresponding format types for subsequent retrieval and management. At the same time, according to different application scenarios, the classified images can be stored in specific folders or databases, classified and stored according to the format type.
[0086] The present application adopts a training strategy of transfer learning to improve the generalization ability and adaptability of the model, and the specific steps are as follows:
[0087] Image classification data set making: scanning images of the same format are used to form a data set, and a labeling tool is used for labeling. Only the content filling part needs to be labeled, including the position of the text box and its classification information (such as "name", "party time", etc.).
[0088] Training parameter setting: optimizer: Adam, parameter setting: β1=0.9, β2=0.999. Learning rate: 7e-3. Training rounds (epochs): 1000.
[0089] Transfer learning process: initial training: first, the initial training is performed on the "cadre resume table" data set to obtain a basic model. Fine-tuning training: the model obtained by the initial training is used as the basic model, and fine-tuning is performed on the specific cadre archive table data set to adapt to different table formats and improve the classification accuracy and model generalization ability.
[0090] In some embodiments, the image classification model: adopts a convolutional neural network (CNN) or a lightweight neural network, designed to accurately classify images into ten predefined layout categories. Transfer learning: through the transfer learning strategy, pre-training on a large-scale general image dataset, and fine-tuning on a specific cadre archive table dataset, to improve the classification accuracy and the generalization ability of the model.
[0091] In step S200 of some embodiments, the pre-processed document image is input into the trained text detection and classification model for processing, outputting text location information and corresponding text classification information.
[0092] It should be noted that the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0093] The determination step of the trained text detection and classification model includes:
[0094] The document image training samples are subjected to feature extraction processing to extract multi-level feature maps;
[0095] The feature maps of different levels are fused to obtain a fused feature map, and the fused feature map is updated based on the classification feature map for calculating the text box category to generate a detection and classification feature map;
[0096] The detection and classification feature map is subjected to convolution processing and branch processing in sequence to obtain an optimized probability map, a threshold map, and a classification map;
[0097] Based on the probability map, the threshold map, the classification map, and a joint loss function, the lightweight neural network is subjected to multiple learning and training processing to obtain the trained text detection and classification model.
[0098] Specifically, the document image training samples are subjected to feature extraction by a lightweight neural network to obtain multi-level feature maps. These feature maps can capture different levels of semantic information in the image, such as edges, textures, shapes, etc.
[0099] Specifically, a lightweight neural network MobileNetV3 is used as the backbone network to extract multi-level feature maps from the input image, and an FPN (Feature Pyramid Network) is connected to fuse features of different levels, and the feature fusion is completed through a concat operation to improve the detection capability of different scale texts.
[0100] The feature maps of different levels are fused to integrate the information of each level to obtain a fused feature map. This step helps to integrate the global and local features of the image and improves the understanding and recognition ability of the model for text.
[0101] The classification feature map of the text box category is calculated, and the fusion feature map is updated according to the classification feature map, so as to generate a detection classification feature map. The classification feature map contains important information about the text category, and by updating the fusion feature map, the model can pay more attention to features related to text classification.
[0102] After feature fusion, 3x3 convolution and deconvolution operations are performed twice to further optimize the feature map.
[0103] Specific convolution processing: the detection classification feature map is subjected to convolution operation to further extract features and enhance the expression ability of the features. Convolution operation can effectively capture local features and patterns in the image, which helps to improve the detection accuracy of the model.
[0104] Branch processing: after convolution processing, the obtained feature map is sent to different branches for processing to obtain optimized probability map, threshold map and classification map. The probability map is used to represent the probability of each pixel belonging to text or background; the threshold map is used to determine the boundary of the text region; and the classification map gives the category information of the text.
[0105] Based on the joint loss function training: using the probability map, threshold map, classification map and joint loss function, the lightweight neural network is subjected to multiple learning and training processes. During the training process, the model adjusts its parameters according to the value of the loss function to minimize the loss function, thereby improving the performance of the model. Obtain the trained model: after multiple iterations of training, when the loss function value of the model on the training data converges to a small range, it is considered that the model has been trained and the final text detection and classification model has been obtained.
[0106] In some embodiments of the present application, the formula of the joint loss function composed of the detection loss and the joint loss is:
[0107]
[0108] Wherein, L is the joint loss function, Lc represents the loss corresponding to the classification map, is the loss of the threshold map, is the loss corresponding to the approximate binary map, is the loss of the probability map, , is a constant value; the approximate binary map is obtained by differentiable binary calculation on the threshold map and the probability map.
[0109] In the text detection and classification model, in order to comprehensively consider the loss of multiple tasks or multiple aspects, a multi-task learning joint loss function is used for optimization. The joint loss function combines the losses of different tasks together, and adjusts the weight coefficients to balance the importance of each task.
[0110] Lc (loss corresponding to classification map): measures the prediction accuracy of the model on the text category. By comparing the difference between the predicted classification result and the true label, the model learns the correct text category information.
[0111] (Loss of threshold map): used to optimize the boundary positioning of the text region. It enables the model to more accurately determine the location and range of the text, improving the accuracy of detection.
[0112] (Loss corresponding to approximate binary map): the approximate binary map is a differentiable binary calculation of the threshold map and probability map. This loss helps further refine the distinction between text and non-text regions, enabling the model to accurately detect text in complex backgrounds.
[0113] (Loss of probability map): focuses on the prediction accuracy of the probability that each pixel belongs to text or background. By minimizing this loss, the model can better estimate the attribution probability of the pixel, thereby improving the overall detection effect.
[0114] 、 As weight coefficients, 、 can take values of 1.0 and 10 respectively, used to balance the contribution of each loss in the joint loss function. By reasonably adjusting these weight coefficients, the model can achieve better performance in different aspects according to specific application scenarios and requirements. For example, if the accuracy of text classification is required to be higher, the value of can be appropriately increased; if more attention is paid to the accurate positioning of the text boundary, the value of can be increased.
[0115] Further, Lc is a cross-entropy loss function, the formula is as follows:
[0116]
[0117] where, represents the value of the i-th class in the true label, which can take values of 0 or 1, represents the probability of the model predicting the i-th class, and Σ represents the cross-entropy loss obtained by summing the loss of each class.
[0118] Cross-entropy loss function is a commonly used loss function for measuring the difference between the predicted results of a classification model and the true labels. In the text detection classification task, it can effectively measure the difference between the predicted text category probability distribution of the model and the true category label.
[0119] In step S300 of some embodiments, based on the text location information and the text classification information, a text recognition process is performed on the initial document to obtain target text content corresponding to the initial document.
[0120] In some embodiments, the text location information is used to determine regions in the initial document that are likely to contain text. This location information can be obtained through pre-image analysis, layout detection, or manual annotation.
[0121] Further, for some documents with specific formats, such as tables, reports, etc., text location information can help us quickly lock in the text area within the table or the title, paragraph part in the report.
[0122] Further, according to the text classification information, the text area to be recognized is further filtered. Different categories of text may have different characteristics and importance. For example, if the document is classified as a "technical report", we may pay more attention to the technical parameters, data and conclusions in it; if it is a "news article", we focus on the title, lead and main content. For multi-language documents, text classification information can also help us determine the location of different language texts, so as to use appropriate language models for subsequent recognition.
[0123] According to the characteristics and needs of the document, select the appropriate optical character recognition (OCR) technology to perform text recognition processing on the initial document to obtain target text content corresponding to the initial document.
[0124] The selected OCR engine is used to recognize the filtered text area. During the recognition process, the OCR engine converts the text in the image into computer-understandable character encoding.
[0125] For some difficult-to-recognize parts, such as blurred text, special symbols or rare characters, manual intervention or additional language models may be needed for auxiliary recognition.
[0126] In some embodiments of the present application, an OCR model supporting multi-language and multi-font recognition is selected. Common OCR models include Tesseract, CRNN (Convolutional Neural Network Recurrent Neural Network), Attention-OCR, etc. These models have different characteristics and advantages, and can be selected according to specific needs, which are not limited here.
[0127] Tesseract is an open-source OCR engine that supports multiple operating systems and programming languages, with high recognition accuracy and good scalability. CRNN combines the advantages of convolutional neural networks and recurrent neural networks, effectively handling image scale changes and sequence information, and performing well on irregular text and handwritten text recognition. Attention-OCR introduces an attention mechanism to focus more on key areas of the text, improving recognition accuracy.
[0128] Further, if the pre-trained model does not meet the specific needs, a large number of labeled document images can be used to retrain the selected OCR model. The labeled data should cover multiple languages, fonts, and document format types to enable the model to learn the features of various text patterns.
[0129] For example, to build an OCR model that can recognize traditional Chinese characters and variant characters in ancient literature, a large number of images containing ancient literature characters need to be collected, and the characters in them need to be accurately labeled, and then these labeled data are used to train the model.
[0130] In some embodiments, the preprocessed document image is input into the trained OCR model, which will recognize the text in the preprocessed document image character by character or word by word, output the corresponding text encoding or character sequence, i.e., the target text content (key information).
[0131] For example, for an image of an English newspaper article, the OCR model will recognize the words, punctuation marks, etc. in it and convert them into a text format that computers can understand.
[0132] In some embodiments, the results of OCR recognition, i.e., the target text content, are preliminarily checked for obvious errors such as misspelled words, missing words, extra spaces, etc. This can be achieved through simple string matching, dictionary checking, or language models.
[0133] For problems found in the preliminary check, manual correction is performed, and in the correction process, the original document image or other reliable information sources can be referred to to ensure that the corrected text is consistent with the original intent. After checking and correcting, the correct target text content corresponding to the initial document is obtained.
[0134] In some embodiments of the present application, after the step of performing text recognition processing on the initial document based on the text position information and the text classification information to obtain the target text content corresponding to the initial document, the method comprises:
[0135] determining the text mapping relationship between the text classification information and the target text content;
[0136] Based on the text mapping relationship, the text classification information, the text position information and the target text content are stored in a structured manner.
[0137] Specifically, the text classification information is analyzed in depth to determine the characteristics, themes or other related attributes of each category of text. Key information is extracted from the target text content, which can be entities (such as names, place names, organization names, product names, etc.), keywords, theme sentences, etc.
[0138] The text classification information is associated with the extracted key information of the target text content to establish a mapping relationship. This can be achieved in various ways, such as using labels, indexes or knowledge graphs, etc.
[0139] For example, a unique identifier is created for each text category, and then the key information related to the category is marked in the target text content, and these key information is associated with the category identifier.
[0140] The text classification information and the text content classification mapping are combined to form key-value structured data and stored. First, choose JSON format, because JSON (JavaScript Object Notation) format has the characteristics of lightweight, easy to read and write, language independent, etc., which is very suitable for data exchange and storage. It can clearly represent hierarchical data, and is very suitable for the structured representation of text classification information and text content.
[0141] In some embodiments, storage is performed through a database management system, supporting efficient data retrieval and management, facilitating subsequent data processing and analysis.
[0142] The database of the embodiment example of the present application: such as MySQL, PostgreSQL, Oracle, etc. These databases have mature technical architecture and powerful functions, supporting transaction processing, concurrent access, data integrity constraints, etc.
[0143] For example, a table can be created to store the structured data of text classification information and text content. For example, the table name is text_data, which contains the following columns:
[0144] id: primary key, used to uniquely identify each record.
[0145] category: used to store text classification information.
[0146] position: used to store text position information.
[0147] content: used to store text content.
[0148] When storing data in JSON format, the entire JSON object can be stored as text in the content column, and extracted and processed by the corresponding parsing function when needed.
[0149] For example, the JSON file of text classification information and text content is stored in the file system, and is organized according to a certain directory structure, for example, a folder is established according to the text classification name, and the corresponding text file is stored under each folder. At the same time, an index library is established to record the path of the text file, the text classification information and the key text position information, so as to quickly locate and retrieve data.
[0150] The structured storage method provided by the embodiment of the application is simple and intuitive, easy to implement and manage. The file system combined with the index library has low cost and can meet the requirements.
[0151] The embodiment provided by the embodiment of the application can at least achieve the following technical effects, but is not limited to the following technical effects:
[0152] Simple labeling: when making a training data set, only the value field and its category in the table need to be labeled, which significantly reduces the labeling workload.
[0153] Lightweight model: compared with traditional deep learning models, the lightweight neural network design is adopted in the application, the model size is smaller, and it is suitable for deployment in resource-constrained environments.
[0154] Fast prediction speed: the optimized model has a faster prediction speed and can efficiently process a large number of documents, and is suitable for real-time or near real-time information extraction requirements.
[0155] High accuracy: through the joint loss function design of multi-task learning, combined with text detection, classification and recognition, the high accuracy of information extraction is ensured.
[0156] Wide application: the method is suitable for key information extraction of various fixed format tables, and can automatically adjust the parameter configuration of the text detection and classification model according to different formats, so as to ensure efficient and accurate information extraction under different table formats.
[0157] Resource optimization: the method can efficiently run on low-performance hardware platforms such as edge computing devices and mobile devices, significantly reducing the consumption of computing resources, while ensuring the accuracy of information extraction.
[0158] The application provides a lightweight-based document key information extraction method, device and equipment and a storage medium. The method comprises the following steps: obtaining an initial document image of an initial document, and preprocessing the initial document image to obtain a preprocessed document image; inputting the preprocessed document image into a trained text detection and classification model for processing, and outputting text position information and corresponding text classification information; performing text recognition processing on the initial document based on the text position information and the text classification information to obtain target text content corresponding to the initial document; wherein the preprocessing at least comprises image format processing and image classification processing; and the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples. The method can solve the problems of low efficiency and high cost in extracting document key information in the prior art, and can efficiently and at low cost extract target text content of a document by determining the document text information to be recognized in advance through a text detection and classification model.
[0159] The lightweight-based document key information extraction device provided by the application is described below, and the lightweight-based document key information extraction device described below can be referred to in conjunction with the lightweight-based document key information extraction method described above.
[0160] As shown in Figure 2 The structure of the lightweight-based document key information extraction device provided by the application is shown in the figure, and the lightweight-based document key information extraction device comprises the following modules:
[0161] The preprocessing module 210 is configured to obtain an initial document image of an initial document, and preprocess the initial document image to obtain a preprocessed document image.
[0162] The detection and classification module 220 is configured to input the preprocessed document image into a trained text detection and classification model for processing, and output text position information and corresponding text classification information.
[0163] The text recognition module 230 is configured to perform text recognition processing on the initial document based on the text position information and the text classification information to obtain target text content corresponding to the initial document; wherein the preprocessing at least comprises image format processing and image classification processing; and the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0164] Preferably, the lightweight-based document key information extraction device provided by the application is specifically configured to perform image format processing on the initial document image to obtain a standard document image.
[0165] The document format type of the standard document image is recognized.
[0166] perform image classification processing on the standard document image based on the document layout type, to obtain the preprocessed document image after image classification.
[0167] Preferably, the lightweight-based document key information extraction device provided by the application is also used for performing image cropping processing on the initial document image, to obtain a cropped document image.
[0168] perform image denoising processing on the cropped document image, to obtain a denoised document image.
[0169] perform resolution optimization processing on the denoised document image, to obtain the standard document image.
[0170] Preferably, the lightweight-based document key information extraction device provided by the application is also used for determining a text mapping relationship between the text classification information and the target text content.
[0171] based on the text mapping relationship, the text classification information, the text position information and the target text content are stored in a structured manner.
[0172] Preferably, the lightweight-based document key information extraction device provided by the application is also used for performing feature extraction processing on the document image training sample, to extract a multi-level feature map.
[0173] fuse feature maps of different levels to obtain a fused feature map, and update the fused feature map based on a classification feature map of a text box class, to generate a detection classification feature map;
[0174] perform convolution processing and branch processing on the detection classification feature map in sequence, to obtain an optimized probability map, a threshold value map and a classification map.
[0175] based on the probability map, the threshold value map, the classification map and a joint loss function, perform multiple learning and training processing on the lightweight neural network, to obtain a trained text detection classification model.
[0176] Preferably, the lightweight-based document key information extraction device provided by the application is also used for the formula of the joint loss function composed of a detection loss and a joint loss to be:
[0177]
[0178] wherein L is a joint loss function, Lc represents a loss corresponding to the classification map, is a loss of the threshold value map, is a loss corresponding to an approximate binaryzation map, is a loss of the probability map. , The value is a constant; the approximate binarized map is obtained by performing differentiable binarization calculations on the threshold map and the probability map.
[0179] This invention provides a lightweight method, apparatus, device, and storage medium for extracting key document information. The method involves acquiring an initial document image and preprocessing it to determine a preprocessed document image. The preprocessed document image is then input into a trained text detection and classification model for processing, outputting text location information and corresponding text classification information. Based on the text location information and the text classification information, text recognition processing is performed on the initial document to obtain the target text content corresponding to the initial document. The preprocessing includes at least image processing and image classification processing. The text detection and classification model is trained on a lightweight neural network using a large number of document image training samples. This invention addresses the shortcomings of existing technologies in extracting key document information due to low efficiency and high computational cost. By pre-determining the document text information to be recognized through a text detection and classification model, the target text content of a document can be extracted efficiently and at low cost.
[0180] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, communications interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a lightweight document key information extraction method. This method includes: acquiring an initial document image of an initial document and preprocessing the initial document image to determine a preprocessed document image; inputting the preprocessed document image into a trained text detection and classification model for processing, outputting text location information and corresponding text classification information; and performing text recognition processing on the initial document based on the text location information and the text classification information to obtain target text content corresponding to the initial document. The preprocessing includes at least image processing and image classification processing; the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0181] In addition, the logic instructions in the memory 330 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0182] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute a lightweight-based document key information extraction method provided by the above-mentioned methods, the method comprising: obtaining an initial document image of an initial document, and preprocessing the initial document image to determine a preprocessed document image; inputting the preprocessed document image into a trained text detection classification model for processing to output text position information and corresponding text classification information; based on the text position information and the text classification information, performing text recognition processing on the initial document to obtain target text content corresponding to the initial document; wherein the preprocessing at least includes image format processing and image classification processing; and the text detection classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0183] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute a lightweight-based document key information extraction method provided by the above-mentioned methods, the method comprising: obtaining an initial document image of an initial document, and preprocessing the initial document image to determine a preprocessed document image; inputting the preprocessed document image into a trained text detection classification model for processing to output text position information and corresponding text classification information; based on the text position information and the text classification information, performing text recognition processing on the initial document to obtain target text content corresponding to the initial document; wherein the preprocessing at least includes image format processing and image classification processing; and the text detection classification model is obtained by training a lightweight neural network using a large number of document image training samples.
[0184] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0186] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A lightweight document key information extraction method, characterized in that, include: Obtain the initial document image of the initial document, and preprocess the initial document image to determine the preprocessed document image; The preprocessed document image is input into the trained text detection and classification model for processing, and the text location information and corresponding text classification information are output. Based on the text location information and the text classification information, the initial document is subjected to text recognition processing to obtain the target text content corresponding to the initial document; The preprocessing includes at least image format processing and image classification processing; the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples. The steps for determining the trained text detection and classification model include: Feature extraction processing is performed on the document image training samples to obtain multi-level feature maps; By fusing feature maps from different levels to obtain a fused feature map, and updating the fused feature map based on the classification feature map of the text box category, a detection classification feature map is generated. The detection classification feature map is sequentially subjected to convolution and branching processes to obtain the optimized probability map, threshold map, and classification map; Based on the probability map, the threshold map, the classification map, and the joint loss function, the lightweight neural network is trained multiple times to obtain a trained text detection and classification model. The formula for the joint loss function, which consists of the detection loss and the joint loss, is as follows: L=L s +a*L b +β*L t +L c Where L is the joint loss function, Lc represents the loss corresponding to the classification map, and L t For the loss of the threshold map, L b L is the loss corresponding to the approximate binary image. s The loss of the probability map is denoted by α and β, which are constant values. The approximate binarized map is obtained by performing differentiable binarization calculations on the threshold map and the probability map.
2. The lightweight document key information extraction method according to claim 1, characterized in that, The step of preprocessing the initial document image to determine the preprocessed document image includes: The initial document image is processed in an image format to obtain a standard document image; Identify the document layout type of the standard document image; Based on the document layout type, the standard document image is subjected to image classification processing to obtain the preprocessed document image after image classification.
3. The lightweight document key information extraction method according to claim 2, characterized in that, The step of processing the initial document image to obtain a standard document image includes: The initial document image is cropped to obtain a cropped document image; The cropped document image is subjected to image denoising processing to obtain a denoised document image; The denoised document image is then subjected to resolution optimization processing to obtain the standard document image.
4. The lightweight document key information extraction method according to any one of claims 1 to 3, characterized in that, After the step of performing text recognition processing on the initial document based on the text location information and the text classification information to obtain the target text content corresponding to the initial document, the method includes: Determine the text mapping relationship between the text classification information and the target text content; Based on the text mapping relationship, the text classification information, the text location information, and the target text content are stored in a structured manner.
5. A lightweight document key information extraction device, characterized in that, include: The preprocessing module is used to acquire the initial document image of the initial document, preprocess the initial document image, and determine the preprocessed document image; The detection and classification module is used to input the preprocessed document image into the trained text detection and classification model for processing, and output text location information and corresponding text classification information. The text recognition module is used to perform text recognition processing on the initial document based on the text location information and the text classification information to obtain the target text content corresponding to the initial document; wherein, the preprocessing includes at least image processing and image classification processing; the text detection and classification model is obtained by training a lightweight neural network using a large number of document image training samples; The steps for determining the trained text detection and classification model include: Feature extraction processing is performed on the document image training samples to obtain multi-level feature maps; By fusing feature maps from different levels to obtain a fused feature map, and updating the fused feature map based on the classification feature map of the text box category, a detection classification feature map is generated. The detection classification feature map is sequentially subjected to convolution and branching processes to obtain the optimized probability map, threshold map, and classification map; Based on the probability map, the threshold map, the classification map, and the joint loss function, the lightweight neural network is trained multiple times to obtain a trained text detection and classification model. The formula for the joint loss function, which consists of the detection loss and the joint loss, is as follows: L=L s +a*L b +β*L t +L c Where L is the joint loss function, Lc represents the loss corresponding to the classification map, and L t For the loss of the threshold map, L b L is the loss corresponding to the approximate binary image. s The loss of the probability map is denoted by α and β, which are constant values. The approximate binarized map is obtained by performing differentiable binarization calculations on the threshold map and the probability map.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the lightweight document key information extraction method as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the lightweight document key information extraction method as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the lightweight document key information extraction method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Text classification and recognition method and device based on target detection
CN112036395A
Text detection model generation method and text detection method
CN112528976A
Text detection method and device, computer equipment and storage medium
CN116975266A
High-precision character segmentation method and system
CN118262362A