OCR intelligent image classification processing platform
By integrating convolutional neural networks and object detection algorithms in image classification and OCR processing platforms, a comprehensive image model is constructed, and the accuracy and efficiency problems in image classification and OCR text processing are solved, and the rapid and accurate extraction of text information in the image is achieved.
Patent Information
- Application Number
- CN202411672989.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art challenges in image data classification and OCR text processing, especially when dealing with images with noisy, blurry, incomplete or complex backgrounds.
A OCR intelligent image classification processing platform is proposed, which constructs an image comprehensive model through the integration of convolutional neural network and object detection algorithm to realize the intelligent image classification and OCR processing of images. The platform includes an image comprehensive model building module, an image recognition analysis module, an OCR model building module, an image analysis model building module, an image text filling module and an information data association storage module.
It realizes the rapid and accurate extraction of text information in the image, improves the accuracy of image classification and the efficiency of OCR text processing, and provides efficient and convenient data processing services.
Smart Images

Figure CN120014647A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of character recognition, and in particular relates to an OCR intelligent image classification processing platform. Background Art
[0002] Image data usually contain various types of data information such as text and numbers, and the types of image data are diverse and may belong to different fields (such as medicine, remote sensing, surveillance video, etc.). However, due to the different classification standards in different fields, the classification of image data may be incorrect due to different types and fields. Even with the development of artificial intelligence, machine learning and deep learning technologies, many fields have begun to use automated tools for image classification. However, automatic classification technology still faces some challenges, especially in terms of classification accuracy, algorithm adaptability and interpretability. For example, noise, blur, incompleteness or complex background in the image may affect the effect of the automatic classification system and further affect the application of OCR technology.
[0003] OCR (Optical Character Recognition) is used to extract text information and other key information from image materials, including document type, content subject, and even font and language. Therefore, the classification of image materials will affect the accuracy and efficiency of subsequent OCR text processing.
[0004] Therefore, how to improve the accuracy of image data classification and further improve the accuracy and efficiency of text processing of OCR technology has become a problem that needs to be solved urgently. Summary of the invention
[0005] The purpose of the present invention is to provide an OCR intelligent image classification and processing platform, which integrates advanced artificial intelligence and optical character recognition technology to intelligently classify and OCR various types of image data to achieve rapid and accurate extraction of text information in images, thereby providing users with efficient and convenient data processing services.
[0006] The purpose can be achieved through the following technical solutions:
[0007] In a first aspect, the present application embodiment provides an OCR intelligent image classification processing platform, including:
[0008] An image comprehensive model building module is used to build an image comprehensive model;
[0009] An image recognition and analysis module, used for acquiring an image to be processed and inputting it into the image comprehensive model for recognition and analysis to output an image to be recognized and an image to be analyzed respectively;
[0010] An OCR model building module, used for building an OCR model and extracting characters from the image to be recognized based on the OCR model to output image recognition text;
[0011] An image analysis model building module, used to build an image analysis model and perform content analysis on the image to be analyzed based on the image analysis model to output image analysis data;
[0012] An image text filling module fills the image recognition text with text content based on the image analysis data and outputs the filled comprehensive text information;
[0013] An information data association storage module, used to perform association analysis on the image recognition text, the image analysis data and the comprehensive text information and establish association rules, and input the association rules into the image content database for storage;
[0014] Wherein, the establishing of the comprehensive image model comprises:
[0015] Acquire image data from multiple fields and use them to build image datasets;
[0016] Based on the image dataset, a convolutional neural network is used to construct an image classification model;
[0017] Based on the image data set, an image recognition model is constructed using a target detection algorithm;
[0018] combining the image classification model and the image recognition model to form the image synthesis model;
[0019] Among them, image data in multiple fields include medical images, remote sensing images, industrial product images, transportation images and business document images.
[0020] Preferably, the use of a convolutional neural network to construct an image classification model includes:
[0021] Marking the type of each image in the image dataset to form a type label corresponding to each image;
[0022] Obtaining an original image and a type label corresponding to the original image;
[0023] Inputting the original image into a pre-built convolutional neural network model, wherein the convolutional neural network model includes a feature extraction network and a type prediction network, and the feature extraction network includes N convolutional layers connected in series;
[0024] Using the feature extraction network to extract features from the original image, obtaining a first feature map output by the Nth convolutional layer and a second feature map output by the N-1th convolutional layer;
[0025] Generate a first training set based on the first feature map, the type label of the original image, and a plurality of preset binary masks;
[0026] Generate a second training set based on the second feature map, the type label of the original image and the multiple binary masks;
[0027] The type prediction network is trained using the first training set and the second training set to obtain the trained image classification model.
[0028] Preferably, the training of the type prediction network using the first training set and the second training set includes:
[0029] Merging the first training set and the second training set into a training sample set;
[0030] Acquire multiple image samples with labeled type labels in the training sample set;
[0031] For each image sample, input the image sample into the type prediction network;
[0032] In any convolutional layer of the type prediction network, a plurality of feature images are extracted from the image sample and the type of the image sample and the type labels of the plurality of feature images are predicted;
[0033] Using the predicted type of the image sample and the type labels of the plurality of feature images, as well as the type label of the image sample that has been annotated, a loss value of the type prediction network is calculated;
[0034] Determine whether the loss value reaches the expected training target, if not, adjust the parameters of the type prediction network based on the loss value, and return to execute the step of inputting the image sample into the type prediction network for each image sample until the latest loss value reaches the expected training target;
[0035] If yes, the training is terminated to obtain the trained image classification model;
[0036] Among them, different feature images represent image samples lacking different channels, and different channels represent image samples with different image features.
[0037] Preferably, the use of a target detection algorithm to construct an image recognition model includes:
[0038] Data preparation: establishing a target detection dataset based on the image dataset; wherein the target detection dataset includes text images and non-text images;
[0039] Data labeling: labeling the text image as 0, labeling the non-text image as 1, and labeling the bounding box of the text area in the text image;
[0040] Data preprocessing: standardizing, enhancing and normalizing the images in the target detection dataset;
[0041] Model training: Based on the target detection data set, the YOLOv5 algorithm is used to perform model training and continuously optimize the loss function to generate the image recognition model; wherein the loss function includes classification loss, positioning loss and confidence loss;
[0042] Model output: For each image input into the image recognition model, after being processed by the model, the text image category and the non-text image category are output; if the text image category is output, the bounding box of the text area is also output and marked.
[0043] Preferably, establishing the target detection data set includes:
[0044] Acquire an initial target image from the image dataset;
[0045] Performing edge pre-recognition on the initial target image to obtain edge information of each local image contained in the initial target image;
[0046] Determine the edge range of each local image according to the edge information, classify the local image whose edge range exceeds a set threshold as a non-text image, otherwise classify it as a local text image, and combine the local text images whose adjacent edge intervals are less than a set threshold in the divided local text images into text images;
[0047] Identify the distance between the closest edges of the non-text image and the text image, and if the distance between the edges is lower than a set threshold, determine that the non-text image and the text image are associated and are marked as the same; otherwise, determine that the non-text image and the text image are not associated and are marked as different;
[0048] Recombining the non-text image and the text image with the same mark into a graphic image, and using the graphic image to replace the initial target image corresponding to the graphic image;
[0049] The object detection dataset is established based on the non-text image and the text image.
[0050] Preferably, constructing the OCR model includes:
[0051] Establishing a text correction unit, a text frame positioning unit and a text extraction unit respectively;
[0052] The text correction unit, the text box positioning unit and the text extraction unit are integrated to form the OCR model;
[0053] Wherein, the text correction unit comprises:
[0054] An image acquisition subunit is used to acquire the image to be recognized and input the image to be recognized into a text line detection network to obtain a text mask image corresponding to the image to be recognized; and use the text mask image to determine the text line contour and the text center line corresponding to each text line contour;
[0055] The sampling point construction subunit uses the text mask image to set a first control point corresponding to the image to be recognized; the first control point includes a text sampling point and a border sampling point, selects a plurality of text sampling points on the text midline, and constructs a text sampling point set for each text midline; and sets a plurality of border sampling points on the text mask image to form a border sampling point set;
[0056] A control point setting subunit, configured to set a second control point according to the first control point; the second control point includes a first source point set and a second source point set, the first source point set includes a plurality of text source point subsets, and the text source point subsets correspond to the text sampling point set one by one, the number of text source points in each text source point subset is the same as the number of text sampling points in the corresponding text sampling point set, and the number of source points in the second source point set is the same as the number of border sampling points in the border sampling point set;
[0057] The corrected image acquisition subunit is used to input the image to be recognized, the first control point and the second control point into an image correction network to obtain a corrected text image.
[0058] Preferably, the text box positioning unit includes:
[0059] An object detection subunit, configured to detect a first object in the corrected text image and obtain a first recognition image corresponding to the first object;
[0060] A feature extraction subunit processes the first recognition image using a text feature extraction network to obtain a first text feature; the first text feature is used to characterize a text box corresponding to the first recognition image;
[0061] A feature processing subunit, configured to rasterize the first text feature to generate first raster information, and process the second raster information detected from the first raster information according to a preset feature merging rule to generate third raster information;
[0062] A positioning calculation subunit, used to determine the position information and positioning result corresponding to the vertices of the text box according to the third grid information;
[0063] The first grid information is used to represent the position of the text element of the text box in the text box; the second grid information includes the first grid information corresponding to the text element in the text box.
[0064] Preferably, the text extraction unit comprises:
[0065] The positioning information collection subunit is used to obtain the position information and positioning results corresponding to the vertices of a series of text boxes, and to establish a positioning information set based on them;
[0066] A text box optimization subunit, used to optimize the text box based on the positioning information set and generate a text box array, each element of the array is a paragraph;
[0067] A text recognition subunit, used to recognize each element of the text box array in turn to obtain text content consisting of paragraphs;
[0068] The text output subunit is used to verify the text content and output the verified correct content as image recognition text.
[0069] Preferably, the image analysis model comprises:
[0070] An image encoding subunit, used for extracting features from the image to be analyzed and obtaining visual features and emotional features;
[0071] A multimodal mapping subunit, configured to convert the visual features into mapping features of a text feature embedding space;
[0072] A content understanding subunit, used for acquiring a task instruction text and inputting the task instruction text and the mapping features into a pre-trained large language model for content understanding, so as to acquire content understanding features and content description text;
[0073] A context analysis subunit performs emotion perception and scene understanding based on the emotion features and obtains context information;
[0074] An analysis data generation subunit generates the image analysis data based on the content description text and the context information and outputs the image analysis data;
[0075] The content description text is obtained by mapping the content understanding features to text based on the output layer of the large language model.
[0076] In a second aspect, an embodiment of the present application provides an OCR intelligent image classification processing method, which uses an OCR intelligent image classification processing platform as described above, and includes the following steps:
[0077] Establishing comprehensive image model;
[0078] Acquire the image to be processed and input it into the image comprehensive model for recognition and analysis to output the image to be recognized and the image to be analyzed respectively;
[0079] Constructing an OCR model and extracting characters from the image to be recognized based on the OCR model to output image recognition text;
[0080] Constructing an image analysis model and performing content analysis on the image to be analyzed based on the image analysis model to output image analysis data;
[0081] Filling the image recognition text with text content based on the image analysis data, and outputting the filled comprehensive text information;
[0082] The image recognition text, the image analysis data and the comprehensive text information are subjected to correlation analysis and correlation rules are established, and the correlation rules are input into an image content database for storage.
[0083] The beneficial effects of the present invention are:
[0084] (1) The intelligent image classification and OCR processing platform of the present invention integrates advanced artificial intelligence technology and OCR technology to provide users with efficient and accurate image classification and OCR processing services. Whether it is for industries that need to process a large amount of image data or for individual users who pursue efficient office work, the platform can provide strong support and help users improve data processing efficiency and work efficiency.
[0085] (2) The platform provided by the present invention focuses on intelligent classification and OCR processing of various types of image data to achieve rapid and accurate extraction of text information in images, thereby providing users with efficient and convenient data processing services. In terms of intelligent image classification, the platform uses deep learning algorithms to automatically classify input image data. By training and optimizing models, the platform can accurately identify key information in images, such as document type, content theme, etc., and classify them into corresponding categories. This helps users quickly find the required image data and improves the efficiency of data retrieval. In terms of OCR processing, the platform uses advanced OCR technology to identify and extract text in images. By optimizing algorithms and models, the platform can achieve high-accuracy text recognition, including multiple fonts such as printed and handwritten. At the same time, the platform also supports multi-language recognition to meet the needs of different users. In addition, the platform also supports batch processing of large amounts of image data and can process multiple files or images at the same time to improve processing efficiency. In addition, the platform also has high-performance computing capabilities to ensure fast response and stable operation when processing large amounts of data; at the same time, the platform also provides flexible customized development services, which can customize and expand functions according to the specific needs of users. At the same time, the platform also supports integration with other systems to achieve seamless data connection and sharing, making it convenient for users to perform cross-platform operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] For better understanding and implementation, the technical solution of the present application is described in detail below with reference to the accompanying drawings.
[0087] Figure 1 A schematic diagram of the structure of an OCR intelligent image classification processing platform provided in an embodiment of the present application;
[0088] Figure 2 A flowchart of the steps for constructing an image classification model provided in an embodiment of the present application;
[0089] Figure 3 A flowchart of the steps for constructing an image recognition model provided in an embodiment of the present application;
[0090] Figure 4 A flowchart of the steps of an OCR intelligent image classification processing method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0091] In order to further explain the technical means and effects taken by the present invention to achieve the predetermined invention purpose, exemplary embodiments will be described in detail here, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are only examples of methods and systems consistent with some aspects of the present application as detailed in the attached claims.
[0092] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in this article refers to any or all possible combinations of one or more associated listed items.
[0093] The specific implementation methods, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0094] Example 1
[0095] See also Figure 1 The present application embodiment provides an OCR intelligent image classification processing platform, including:
[0096] An image comprehensive model building module is used to build an image comprehensive model;
[0097] An image recognition and analysis module, used for acquiring an image to be processed and inputting it into the image comprehensive model for recognition and analysis to output an image to be recognized and an image to be analyzed respectively;
[0098] An OCR model building module, used for building an OCR model and extracting characters from the image to be recognized based on the OCR model to output image recognition text;
[0099] An image analysis model building module, used to build an image analysis model and perform content analysis on the image to be analyzed based on the image analysis model to output image analysis data;
[0100] An image text filling module fills the image recognition text with text content based on the image analysis data and outputs the filled comprehensive text information;
[0101] The information data association storage module is used to perform association analysis on the image recognition text, the image analysis data and the comprehensive text information and establish association rules, and input the association rules into the image content database for storage.
[0102] Specifically, since image data usually have many different types and involve different fields, such as medical images, remote sensing images, industrial product images, traffic images and business document images, and each field and category of images has different characteristics and features, for example, medical images are mostly of human body structures and various lesions, focusing on details and structures; remote sensing images are mostly from satellites, drones, etc., containing vegetation crops and geographic information, and need to consider geographical environmental factors; industrial product images mostly contain various products and their structures, focusing on defect identification; traffic images mostly contain traffic signs, lane lines, pedestrians, vehicles and other information; business document images refer to images taken in enterprises In the business process of an industry or organization, various types of documents, receipts and their image materials are involved, which are usually business data including tables, charts, documents, etc. that are saved by scanning, photographing or digitizing. Since there are obvious differences between the images of each of the above fields, if you want to obtain the specific text information contained in the image (some fields may not necessarily contain text information in the images, but even if they do not contain text information, they are still applicable to the image classification of this application), you need to classify the above multiple fields, first distinguish the field to which the image material to be identified belongs, and then use OCR technology to extract and recognize text to obtain the text information in the image. Since the existing multi-field image classification is not accurate and fast enough, it will further affect the subsequent OCR recognition efficiency, thereby reducing the data processing efficiency of text recognition and information extraction. Therefore, in order to solve the above problems, the present application proposes the above-mentioned OCR intelligent image classification processing platform, including an image comprehensive model construction module, which is used to obtain image data in multiple fields and establish an image comprehensive model based on it. The comprehensive model can not only identify the type of image to be processed, but also specifically identify and analyze the content of the image; an image recognition and analysis module, which is used to obtain the image to be processed and input it into the image comprehensive model for recognition and analysis to output the image to be recognized and the image to be analyzed respectively; an OCR model A construction module is used to construct an OCR model and extract characters from the image to be recognized based on the OCR model to output image recognition text; an image analysis model construction module is used to construct an image analysis model and perform content analysis on the image to be analyzed based on the image analysis model to output image analysis data; an image text filling module is used to fill the image recognition text with text content based on the image analysis data and output the filled comprehensive text information; an information data association storage module is used to perform association analysis on the image recognition text, the image analysis data and the comprehensive text information and establish association rules, and input the above association rules into an image content database for storage.This application establishes an image classification model to quickly classify the images to be processed, and then uses OCR to quickly identify the text content of the classified images, so as to improve data processing efficiency and work efficiency, thereby providing users with efficient and accurate image classification and OCR processing services.
[0103] It should be noted that the above-mentioned image to be identified means that the image has text content, while the image to be analyzed means that the image does not have text content, and most or all of the area is image content. The image to be identified is used for subsequent OCR processing, while the image to be analyzed is used for subsequent image content analysis, so as to provide contextual interpretation and supplement for the text content obtained after OCR processing, so that the final text content is more accurate and rich.
[0104] In an embodiment provided in the present application, the step of establishing an image comprehensive model includes:
[0105] Acquire image data from multiple fields and use them to build image datasets;
[0106] Based on the image dataset, a convolutional neural network is used to construct an image classification model;
[0107] Based on the image data set, an image recognition model is constructed using a target detection algorithm;
[0108] combining the image classification model and the image recognition model to form the image synthesis model;
[0109] Among them, image data in multiple fields include medical images, remote sensing images, industrial product images, transportation images and business document images.
[0110] Specifically, the image synthesis model of this embodiment is formed by combining an image classification model and an image recognition model, wherein the image classification model can distinguish the type of image to be processed, while the image recognition model preliminarily recognizes the different types of images distinguished by the image classification model, and detects whether the content of the image contains text information or non-text information. Non-text information means only image information without any text content. The above two models are combined into an image synthesis model to provide a data basis and support for subsequent image recognition and analysis.
[0111] like Figure 2 As shown, in one embodiment provided in the present application, the image classification model is constructed using a convolutional neural network, including:
[0112] Marking the type of each image in the image dataset to form a type label corresponding to each image;
[0113] Obtaining an original image and a type label corresponding to the original image;
[0114] Inputting the original image into a pre-built convolutional neural network model, wherein the convolutional neural network model includes a feature extraction network and a type prediction network, and the feature extraction network includes N convolutional layers connected in series;
[0115] Using the feature extraction network to extract features from the original image, obtaining a first feature map output by the Nth convolutional layer and a second feature map output by the N-1th convolutional layer;
[0116] Generate a first training set based on the first feature map, the type label of the original image, and a plurality of preset binary masks;
[0117] Generate a second training set based on the second feature map, the type label of the original image and the multiple binary masks;
[0118] The type prediction network is trained using the first training set and the second training set to obtain the trained image classification model.
[0119] Specifically, the present embodiment adopts a convolutional neural network to construct an image classification model. First, the type of each image in the image data set is annotated to form a type label corresponding to each image, and the type label represents the type and field corresponding to each image; then, the original image and the type label corresponding to the original image are obtained from the annotated image data set; then, the original image is input into a pre-constructed convolutional neural network model, and the convolutional neural network at this time has not been trained and does not have the function of classifying the image; the convolutional neural network model includes a feature extraction network and a type prediction network, wherein the feature extraction network includes N convolutional layers connected in series; then, the feature extraction network is used to extract features from the original image to obtain a first feature map output by the Nth convolutional layer and a second feature map output by the N-1th convolutional layer; then, based on the first feature map, the type label of the original image and a plurality of preset binary masks, a first training set is generated; based on the second feature map, the type label of the original image and a plurality of binary masks, a second training set is generated; finally, the type prediction network is trained using the first training set and the second training set to obtain a trained image classification model.
[0120] It is understandable that the N convolutional layers in this embodiment are connected in series and the specific values are limited, but in order to ensure that the convolutional layer feature extraction has a good effect, N>5 in this embodiment. In order to improve the accuracy of the model in the prior art, the original image is usually expanded several times before being input into the model for training, which will increase the convolution operation time and model overhead, thereby affecting the final effect of the model. Therefore, this embodiment generates a first training set and a second training set based on the first feature map and the second feature map output by the Nth convolutional layer and the N-1th convolutional layer, as well as the type label and multiple binary masks of the original image, and then uses the above two training sets to train the type prediction network to obtain an image classification model, thereby avoiding the convolution operation of the original image data expanded several times, reducing the additional overhead generated in the model training process, and also improving the accuracy and operation rate of the model.
[0121] It should be noted that the multiple binary masks preset in this embodiment refer to a technology for selecting or masking specific areas (pixels), which uses a binary image (i.e., a matrix in which each element is 0 or 1) to determine which areas should be retained and which should be ignored. In convolutional neural networks, binary masks optimize the training and reasoning process of the model by controlling which areas of pixels or features participate in the calculation, are ignored or weighted, which can effectively improve the model performance and computing efficiency.
[0122] In an embodiment provided in the present application, the training of the type prediction network using the first training set and the second training set includes:
[0123] Merging the first training set and the second training set into a training sample set;
[0124] Acquire multiple image samples with labeled type labels in the training sample set;
[0125] For each image sample, input the image sample into the type prediction network;
[0126] In any convolutional layer of the type prediction network, a plurality of feature images are extracted from the image sample and the type of the image sample and the type labels of the plurality of feature images are predicted;
[0127] Using the predicted type of the image sample and the type labels of the plurality of feature images, as well as the type label of the image sample that has been annotated, a loss value of the type prediction network is calculated;
[0128] Determine whether the loss value reaches the expected training target, if not, adjust the parameters of the type prediction network based on the loss value, and return to execute the step of inputting the image sample into the type prediction network for each image sample until the latest loss value reaches the expected training target;
[0129] If yes, the training is terminated to obtain the trained image classification model;
[0130] Among them, different feature images represent image samples lacking different channels, and different channels represent image samples with different image features.
[0131] Specifically, since a large amount of data sets will be used in the above-mentioned model training, which leads to an increase in time cost, in order to solve this problem, this embodiment uses the loss value to judge and adjust the network parameters when training the type prediction network to generate the final image classification model, and the feature image extracted in the convolution layer is only extracted in the convolution layer, and the data set is not expanded, so the number of images for model training is not increased, so the time for each model learning can be reduced, thereby improving the efficiency of model training. It can be understood that based on the type label that has been labeled with the image sample, the type label of each feature image extracted from the image sample is labeled; according to the channel missing from the feature image, the channel label of the image sample and the channel label of each feature image are labeled, so that the loss value can be determined later based on the type label and the channel label.
[0132] like Figure 3 As shown, in one embodiment provided in the present application, the image recognition model is constructed by using a target detection algorithm, including:
[0133] Data preparation: establishing a target detection dataset based on the image dataset; wherein the target detection dataset includes text images and non-text images;
[0134] Data labeling: labeling the text image as 0, labeling the non-text image as 1, and labeling the bounding box of the text area in the text image;
[0135] Data preprocessing: standardizing, enhancing and normalizing the images in the target detection dataset;
[0136] Model training: Based on the target detection data set, the YOLOv5 algorithm is used to perform model training and continuously optimize the loss function to generate the image recognition model; wherein the loss function includes classification loss, positioning loss and confidence loss;
[0137] Model output: For each image input into the image recognition model, after being processed by the model, the text image category and the non-text image category are output; if the text image category is output, the bounding box of the text area is also output and marked.
[0138] Specifically, this embodiment uses the YOLOv5 algorithm in the target detection algorithm for model training, and constructs a target detection model based on the YOLO algorithm to determine whether the input image contains text content. The YOLOv5 algorithm has a high reasoning speed, is suitable for real-time detection and recognition in this embodiment, and can improve the efficiency of target detection and model performance.
[0139] It is understandable that the trained image recognition model in this embodiment can recognize the image of the input model, distinguish whether there is text content in the image, and mark the location of the text content by a bounding box. The image recognized by the model may contain all text content, or all non-text content (all images), or may be a mixture of text content and image content. Therefore, this embodiment marks the text area with a bounding box for subsequent OCR processing, so that text information can be better extracted from it.
[0140] It should be noted that, unlike the above-mentioned image classification model used to determine the type and field of the image, the image recognition model in this embodiment mainly identifies the input image to determine whether it belongs to the text class or the image class. Although both functions include distinguishing categories, the objects they target are essentially different.
[0141] On the other hand, in the image recognition and analysis module, the image to be processed is input into the image synthesis model for recognition and analysis to output the image to be recognized and the image to be analyzed respectively. The image to be recognized here means the image containing text content, that is, the image of the text image category output by the above image recognition model; and the image to be analyzed means the image containing non-text content, that is, the image of the non-text image category output by the above image recognition model. This type of image is used for subsequent analysis of the image content using the image analysis model in the image analysis model construction module.
[0142] In one embodiment provided in the present application, establishing the target detection data set includes:
[0143] Acquire an initial target image from the image dataset;
[0144] Performing edge pre-recognition on the initial target image to obtain edge information of each local image contained in the initial target image;
[0145] Determine the edge range of each local image according to the edge information, classify the local image whose edge range exceeds a set threshold as a non-text image, otherwise classify it as a local text image, and combine the local text images whose adjacent edge intervals are less than a set threshold in the divided local text images into text images;
[0146] Identify the distance between the closest edges of the non-text image and the text image, and if the distance between the edges is lower than a set threshold, determine that the non-text image and the text image are associated and are marked as the same; otherwise, determine that the non-text image and the text image are not associated and are marked as different;
[0147] Recombining the non-text image and the text image with the same mark into a graphic image, and using the graphic image to replace the initial target image corresponding to the graphic image;
[0148] The object detection dataset is established based on the non-text image and the text image.
[0149] Specifically, this embodiment carries out a detailed development on the establishment of a target detection data set. By performing edge preprocessing and recognition on the acquired initial target image, text images and non-text images are divided, which greatly facilitates the classification processing of images and texts and provides a good data foundation for subsequent image recognition and text recognition.
[0150] In an embodiment provided in the present application, constructing the OCR model includes:
[0151] Establishing a text correction unit, a text frame positioning unit and a text extraction unit respectively;
[0152] The text correction unit, the text box positioning unit and the text extraction unit are integrated to form the OCR model;
[0153] Wherein, the text correction unit comprises:
[0154] An image acquisition subunit is used to acquire the image to be recognized and input the image to be recognized into a text line detection network to obtain a text mask image corresponding to the image to be recognized; and use the text mask image to determine the text line contour and the text center line corresponding to each text line contour;
[0155] The sampling point construction subunit uses the text mask image to set a first control point corresponding to the image to be recognized; the first control point includes a text sampling point and a border sampling point, selects a plurality of text sampling points on the text midline, and constructs a text sampling point set for each text midline; and sets a plurality of border sampling points on the text mask image to form a border sampling point set;
[0156] A control point setting subunit, configured to set a second control point according to the first control point; the second control point includes a first source point set and a second source point set, the first source point set includes a plurality of text source point subsets, and the text source point subsets correspond to the text sampling point set one by one, the number of text source points in each text source point subset is the same as the number of text sampling points in the corresponding text sampling point set, and the number of source points in the second source point set is the same as the number of border sampling points in the border sampling point set;
[0157] The corrected image acquisition subunit is used to input the image to be recognized, the first control point and the second control point into an image correction network to obtain a corrected text image.
[0158] Specifically, this embodiment uses a text line detection algorithm combined with an image correction network to interpolate the detected text lines, and then perform distortion correction based on the results. Furthermore, the text line detection algorithm using deep learning can well detect distorted text and has good robustness to various complex situations. Through the text correction unit, this embodiment can correct the distorted text appearing in the image to be recognized, thereby improving the efficiency and accuracy of subsequent text positioning and text extraction.
[0159] In an embodiment provided in the present application, the text box positioning unit includes:
[0160] An object detection subunit, configured to detect a first object in the corrected text image and obtain a first recognition image corresponding to the first object;
[0161] A feature extraction subunit processes the first recognition image using a text feature extraction network to obtain a first text feature; the first text feature is used to characterize a text box corresponding to the first recognition image;
[0162] A feature processing subunit, configured to rasterize the first text feature to generate first raster information, and process the second raster information detected from the first raster information according to a preset feature merging rule to generate third raster information;
[0163] A positioning calculation subunit, used to determine the position information and positioning result corresponding to the vertices of the text box according to the third grid information;
[0164] The first grid information is used to represent the position of the text element of the text box in the text box; the second grid information includes the first grid information corresponding to the text element in the text box.
[0165] Specifically, the present embodiment determines the position of the text box in the text image to be corrected in the above manner, so as to perform text extraction later. It should be noted that the first text feature includes the features of all text elements constituting the text box area, and all text elements constitute the shape of the text box, but only the boundary text elements are used to regress and predict the vertex coordinates of the text box. The above-mentioned text feature extraction network is obtained by training a variety of images to be recognized. It can be understood that the above-mentioned text box positioning unit is performed on the basis of the text correction unit, that is: the distorted image is first corrected by the text correction unit, and then the text box positioning unit is used to locate the text box of the corrected image, and finally the text extraction unit is used to extract text, which can greatly avoid the problem of text extraction errors caused by distorted images, thereby improving the accuracy and efficiency of text extraction.
[0166] In an embodiment provided in the present application, the text extraction unit includes:
[0167] The positioning information collection subunit is used to obtain the position information and positioning results corresponding to the vertices of a series of text boxes, and to establish a positioning information set based on them;
[0168] A text box optimization subunit, used to optimize the text box based on the positioning information set and generate a text box array, each element of the array is a paragraph;
[0169] A text recognition subunit, used to recognize each element of the text box array in turn to obtain text content consisting of paragraphs;
[0170] The text output subunit is used to verify the text content and output the verified correct content as image recognition text.
[0171] In one embodiment provided in the present application, the image analysis model includes:
[0172] An image encoding subunit, used for extracting features from the image to be analyzed and obtaining visual features and emotional features;
[0173] A multimodal mapping subunit, configured to convert the visual features into mapping features of a text feature embedding space;
[0174] A content understanding subunit, used for acquiring a task instruction text and inputting the task instruction text and the mapping features into a pre-trained large language model for content understanding, so as to acquire content understanding features and content description text;
[0175] A context analysis subunit performs emotion perception and scene understanding based on the emotion features and obtains context information;
[0176] An analysis data generation subunit generates the image analysis data based on the content description text and the context information and outputs the image analysis data;
[0177] The content description text is obtained by mapping the content understanding features to text based on the output layer of the large language model.
[0178] Specifically, this embodiment analyzes and recognizes the image content, combines the image content with emotion and context, and generates auxiliary text descriptions for the image. By using a multimodal learning model, images and text can be effectively combined to generate richer and more emotional descriptions, providing data support for the subsequent content supplement of the recognized text obtained after OCR recognition, making the subsequent comprehensive text content more complete and accurate.
[0179] In summary, the specific concept of the above-mentioned OCR intelligent image classification processing platform provided by this application is: first, by establishing an image comprehensive model to identify and classify the image to be processed, quickly distinguish the field of the image, the purpose of this is to improve the recognition efficiency by determining the field of the image, for example; if the input image to be processed is a medical image, then after the recognition of the image comprehensive model, its field can be quickly determined, and targeted recognition can be performed in the subsequent recognition; if the input image to be processed is a business document image, it means that more text recognition is required for documents and table data in the future. In other words, the result of the image comprehensive model recognition will provide different recognition methods for subsequent text extraction and image content recognition, so the establishment of this model can effectively improve the efficiency of subsequent text recognition. Secondly, after the above-mentioned model recognition, the image to be recognized and the image to be analyzed will be obtained. This application adopts different means for these two different results, performs OCR text extraction on the image to be recognized, obtains text recognition results (that is, image recognition text), and performs content understanding and emotional context analysis on the image to be analyzed to obtain content recognition results (that is, image analysis data). Finally, the content recognition results are used as supplementary explanations to improve and fill in the text recognition results, so that the final comprehensive text information is more complete and accurate. And finally, the information obtained from each module is used to establish association rules through correlation analysis and stored in the image content database. The advantage of this is that if the subsequent analysis is performed, relevant data can be called from the database for similarity analysis, which provides a data basis for subsequent image classification and text recognition. Through the above-mentioned conception content, this application greatly improves the accuracy of image data classification and the accuracy and efficiency of text processing of OCR technology.
[0180] In summary, the OCR intelligent image classification and processing platform proposed in this application is a comprehensive platform that integrates advanced artificial intelligence and optical character recognition (OCR) technology. The platform focuses on intelligent classification and OCR processing of various types of image materials to achieve fast and accurate extraction of text information in images, thereby providing users with efficient and convenient data processing services. In terms of intelligent image classification, the platform uses deep learning algorithms to automatically classify input image materials. By training and optimizing models, the platform can accurately identify key information in images, such as document types, content themes, etc., and classify them into corresponding categories. This helps users quickly find the required image materials and improve the efficiency of data retrieval. In terms of OCR processing, the platform uses advanced OCR technology to identify and extract text in images. By optimizing algorithms and models, the platform can achieve high-accuracy text recognition, including multiple fonts such as printed and handwritten. At the same time, the platform also supports multi-language recognition to meet the needs of different users. In addition, the platform also supports batch processing of large amounts of image materials, and can process multiple files or images at the same time to improve processing efficiency. In addition, the platform also has high-performance computing capabilities to ensure fast response and stable operation when processing large amounts of data; at the same time, the platform also provides flexible customized development services, which can customize and expand functions according to the specific needs of users. At the same time, the platform also supports integration with other systems to achieve seamless data docking and sharing, making it convenient for users to perform cross-platform operations.
[0181] Therefore, the above-mentioned intelligent image classification and OCR processing platform of this application provides users with efficient and accurate image classification and OCR processing services by integrating advanced artificial intelligence and OCR technology. Whether it is for industries that need to process a large amount of image data or for individual users who pursue efficient office work, the platform can provide strong support to help users improve data processing efficiency and work efficiency.
[0182] Example 2
[0183] See also Figure 4 The embodiment of the present application provides an OCR intelligent image classification processing method, which uses an OCR intelligent image classification processing platform as described above, and includes the following steps:
[0184] Establishing comprehensive image model;
[0185] Acquire the image to be processed and input it into the image comprehensive model for recognition and analysis to output the image to be recognized and the image to be analyzed respectively;
[0186] Constructing an OCR model and extracting characters from the image to be recognized based on the OCR model to output image recognition text;
[0187] Constructing an image analysis model and performing content analysis on the image to be analyzed based on the image analysis model to output image analysis data;
[0188] Filling the image recognition text with text content based on the image analysis data, and outputting the filled comprehensive text information;
[0189] The image recognition text, the image analysis data and the comprehensive text information are subjected to correlation analysis and correlation rules are established, and the correlation rules are input into an image content database for storage.
[0190] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0191] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0192] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0193] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An OCR intelligent image classification processing platform, characterized by: include: An image comprehensive model building module is used to build an image comprehensive model; An image recognition and analysis module, used for acquiring an image to be processed and inputting it into the image comprehensive model for recognition and analysis to output an image to be recognized and an image to be analyzed respectively; An OCR model building module, used for building an OCR model and extracting characters from the image to be recognized based on the OCR model to output image recognition text; An image analysis model building module, used to build an image analysis model and perform content analysis on the image to be analyzed based on the image analysis model to output image analysis data; An image text filling module fills the image recognition text with text content based on the image analysis data and outputs the filled comprehensive text information; An information data association storage module, used to perform association analysis on the image recognition text, the image analysis data and the comprehensive text information and establish association rules, and input the association rules into an image content database for storage; Wherein, the establishing of the comprehensive image model comprises: Acquire image data from multiple fields and use them to build image datasets; Based on the image dataset, a convolutional neural network is used to construct an image classification model; Based on the image data set, an image recognition model is constructed using a target detection algorithm; combining the image classification model and the image recognition model to form the image synthesis model; Among them, image data in multiple fields include medical images, remote sensing images, industrial product images, transportation images and business document images.
2. The OCR intelligent image classification processing platform according to claim 1, characterized in that: The image classification model is constructed by using a convolutional neural network, including: Marking the type of each image in the image dataset to form a type label corresponding to each image; Obtaining an original image and a type label corresponding to the original image; Inputting the original image into a pre-built convolutional neural network model, wherein the convolutional neural network model includes a feature extraction network and a type prediction network, and the feature extraction network includes N convolutional layers connected in series; Using the feature extraction network to extract features from the original image, obtaining a first feature map output by the Nth convolutional layer and a second feature map output by the N-1th convolutional layer; Generate a first training set based on the first feature map, the type label of the original image, and a plurality of preset binary masks; Generate a second training set based on the second feature map, the type label of the original image and the multiple binary masks; The type prediction network is trained using the first training set and the second training set to obtain the trained image classification model.
3. The OCR intelligent image classification processing platform according to claim 2, characterized in that: The using the first training set and the second training set to train the type prediction network includes: Merging the first training set and the second training set into a training sample set; Acquire multiple image samples with labeled type labels in the training sample set; For each image sample, input the image sample into the type prediction network; In any convolutional layer of the type prediction network, a plurality of feature images are extracted from the image sample and the type of the image sample and the type labels of the plurality of feature images are predicted; Using the predicted type of the image sample and the type labels of the plurality of feature images, as well as the type label of the image sample that has been annotated, a loss value of the type prediction network is calculated; Determine whether the loss value reaches the expected training target, if not, adjust the parameters of the type prediction network based on the loss value, and return to execute the step of inputting the image sample into the type prediction network for each image sample until the latest loss value reaches the expected training target; If yes, the training is terminated to obtain the trained image classification model; Among them, different feature images represent image samples lacking different channels, and different channels represent image samples with different image features.
4. The OCR intelligent image classification processing platform according to claim 1, characterized in that: The method of using a target detection algorithm to construct an image recognition model includes: Data preparation: establishing a target detection dataset based on the image dataset; wherein the target detection dataset includes text images and non-text images; Data labeling: labeling the text image as 0, labeling the non-text image as 1, and labeling the bounding box of the text area in the text image; Data preprocessing: standardizing, enhancing and normalizing the images in the target detection dataset; Model training: Based on the target detection data set, the YOLOv5 algorithm is used to perform model training and continuously optimize the loss function to generate the image recognition model; wherein the loss function includes classification loss, positioning loss and confidence loss; Model output: For each image input into the image recognition model, after being processed by the model, the text image category and the non-text image category are output; if the text image category is output, the bounding box of the text area is also output and marked.
5. The OCR intelligent image classification processing platform according to claim 4, characterized in that: Establishing the target detection data set includes: Acquire an initial target image from the image dataset; Performing edge pre-recognition on the initial target image to obtain edge information of each local image contained in the initial target image; Determine the edge range of each local image according to the edge information, classify the local image whose edge range exceeds a set threshold as a non-text image, otherwise classify it as a local text image, and combine the local text images whose adjacent edge intervals are less than a set threshold in the divided local text images into text images; Identify the distance between the closest edges of the non-text image and the text image, and if the distance between the edges is lower than a set threshold, determine that the non-text image and the text image are associated and are marked as the same; otherwise, determine that the non-text image and the text image are not associated and are marked as different; Recombining the non-text image and the text image with the same mark into a graphic image, and using the graphic image to replace the initial target image corresponding to the graphic image; The object detection dataset is established based on the non-text image and the text image.
6. The OCR intelligent image classification processing platform according to claim 1, characterized in that: Constructing the OCR model includes: Establishing a text correction unit, a text frame positioning unit and a text extraction unit respectively; The text correction unit, the text box positioning unit and the text extraction unit are integrated to form the OCR model; Wherein, the text correction unit comprises: An image acquisition subunit is used to acquire the image to be recognized and input the image to be recognized into a text line detection network to obtain a text mask image corresponding to the image to be recognized; and use the text mask image to determine the text line contour and the text center line corresponding to each text line contour; The sampling point construction subunit uses the text mask image to set a first control point corresponding to the image to be recognized; the first control point includes a text sampling point and a border sampling point, selects a plurality of text sampling points on the text midline, and constructs a text sampling point set for each text midline; and sets a plurality of border sampling points on the text mask image to form a border sampling point set; A control point setting subunit, configured to set a second control point according to the first control point; the second control point includes a first source point set and a second source point set, the first source point set includes a plurality of text source point subsets, and the text source point subsets correspond to the text sampling point set one by one, the number of text source points in each text source point subset is the same as the number of text sampling points in the corresponding text sampling point set, and the number of source points in the second source point set is the same as the number of border sampling points in the border sampling point set; The corrected image acquisition subunit is used to input the image to be recognized, the first control point and the second control point into an image correction network to obtain a corrected text image.
7. The OCR intelligent image classification processing platform according to claim 6, characterized in that: The text box positioning unit comprises: An object detection subunit, configured to detect a first object in the corrected text image and obtain a first recognition image corresponding to the first object; A feature extraction subunit processes the first recognition image using a text feature extraction network to obtain a first text feature; the first text feature is used to characterize a text box corresponding to the first recognition image; A feature processing subunit, configured to rasterize the first text feature to generate first raster information, and process the second raster information detected from the first raster information according to a preset feature merging rule to generate third raster information; A positioning calculation subunit, used to determine the position information and positioning result corresponding to the vertices of the text box according to the third grid information; The first grid information is used to represent the position of the text element of the text box in the text box; the second grid information includes the first grid information corresponding to the text element in the text box.
8. The OCR intelligent image classification processing platform according to claim 6, characterized in that: The text extraction unit comprises: The positioning information collection subunit is used to obtain the position information and positioning results corresponding to the vertices of a series of text boxes, and to establish a positioning information set based on them; A text box optimization subunit, used to optimize the text box based on the positioning information set and generate a text box array, each element of the array is a paragraph; A text recognition subunit, used to recognize each element of the text frame array in turn to obtain text content consisting of paragraphs; The text output subunit is used to verify the text content and output the verified correct content as image recognition text.
9. The OCR intelligent image classification processing platform according to claim 1, characterized in that: The image analysis model includes: An image encoding subunit, used for extracting features from the image to be analyzed and obtaining visual features and emotional features; A multimodal mapping subunit, configured to convert the visual features into mapping features of a text feature embedding space; A content understanding subunit, used for acquiring a task instruction text and inputting the task instruction text and the mapping features into a pre-trained large language model for content understanding, so as to acquire content understanding features and content description text; A context analysis subunit performs emotion perception and scene understanding based on the emotion features and obtains context information; An analysis data generation subunit generates the image analysis data based on the content description text and the context information and outputs the image analysis data; The content description text is obtained by mapping the content understanding features to text based on the output layer of the large language model.
10. An OCR intelligent image classification processing method, using an OCR intelligent image classification processing platform as claimed in any one of claims 1 to 9, characterized in that: The steps include: Establishing comprehensive image model; Acquire the image to be processed and input it into the image comprehensive model for recognition and analysis to output the image to be recognized and the image to be analyzed respectively; Constructing an OCR model and extracting characters from the image to be recognized based on the OCR model to output image recognition text; Constructing an image analysis model and performing content analysis on the image to be analyzed based on the image analysis model to output image analysis data; Filling the image recognition text with text content based on the image analysis data, and outputting the filled comprehensive text information; The image recognition text, the image analysis data and the comprehensive text information are subjected to correlation analysis and correlation rules are established, and the correlation rules are input into an image content database for storage.
Citation Information
Patent Citations
Multi-language text recognition method and device, computer equipment and storage medium
CN110569830A
Method and device for extracting text in picture
CN111612003A
Text correction method and device, electronic equipment and storage medium
CN111695554A
OCR image character recognition and paragraph output method based on deep learning
CN113435449A
Text box positioning method and device, electronic equipment and storage medium
CN114170592A
Cited By
Enhanced beam search decoding for transformer-based OCR with probability score optimization
US12711795B1