Artificial intelligence model construction method and device based on deep learning
By loading pretrained feature architectures in deep learning models and adding extended architectures, combined with the method of learning rate adjustment, the problem of decreasing recognition accuracy of traditional text recognition technology under complex backgrounds and diverse handwritten forms is solved, and efficient and accurate text recognition effect is achieved.
Patent Information
- Application Number
- CN202510094860.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional text recognition technology is susceptible to noise, background interference and font diversity when complex backgrounds or diverse handwritten characters, resulting in a decrease in recognition accuracy and requires a large amount of labeled data to train a high-precision model.
Using a deep learning-based artificial intelligence model construction method, the feature architecture of the pre-trained model is loaded and the extended architecture is added on it, and the training samples are used to train the pre-trained model based on the learning rate to obtain the target model. This method freezes the feature architecture and adjusts the weight of the extended architecture. After the learning rate reaches the set value, the feature architecture is thawed to avoid overfitting and improve the generalization ability of the model.
It realizes the generation of high accuracy and efficiency models under the scarcity of data, avoids the high computational cost and time in traditional methods, improves the accuracy and efficiency of the model, and is suitable for model training situations that lack a large amount of data.
Smart Images

Figure CN120014656A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning technology, and in particular, relates to a method and device for constructing an artificial intelligence model based on deep learning. Background Art
[0002] With the development of modern artificial intelligence and computer vision technology, text recognition technology has emerged, which is widely used in pre-scanned document digitization, license plate recognition, etc.
[0003] Currently, commonly used text recognition technologies mainly rely on traditional image processing algorithms and machine learning models. Most of these methods use traditional methods such as sliding windows, feature-based image processing, and support vector machines to identify text in images by manually extracting features for classification.
[0004] When faced with complex text backgrounds or handwritings of different styles, traditional text recognition methods rely on simple feature extraction algorithms and are easily affected by noise, background interference, and font diversity, resulting in reduced recognition accuracy and difficulty in accurately extracting text information. In addition, manually extracted features are difficult to cope with diverse text forms. When processing characters, traditional text recognition methods often rely on the classification of individual characters and ignore the relationship of context, and require a large amount of labeled data to train a high-precision model. Summary of the invention
[0005] Based on this, it is necessary to provide a method and device for building an artificial intelligence model based on deep learning that can generate high-accuracy and efficiency models under data scarcity conditions in order to address the above-mentioned technical problems.
[0006] In a first aspect, the present application provides a method for building an artificial intelligence model based on deep learning, comprising:
[0007] Obtain a data set, mark the coordinates of the data set, and obtain training samples; the data set includes image files;
[0008] Load the feature architecture of the pre-trained model and add an extended architecture based on the feature architecture; the feature architecture is a module in which the pre-trained model has obtained the target feature relationship; the target feature relationship indicates that the target model outputs the expected result; the extended architecture indicates that the target model produces the expected output after learning the effective features;
[0009] Freeze the feature architecture and adjust the weight of the extended architecture to the set value; set the feature architecture to be unfrozen when the learning rate increases to the set value; the learning rate is based on the maximum learning rate that does not damage the target feature relationship;
[0010] The pre-trained model is trained using the training samples according to the learning rate to obtain the target model.
[0011] In one embodiment, coordinate marking of a data set to obtain a training sample includes:
[0012] Cover several initial anchor boxes in the dataset; the initial anchor boxes are candidate regions of different scales and aspect ratios;
[0013] Determine whether the initial anchor frame contains the target object, and obtain anchor frame data of the initial anchor frame containing the target object; the anchor frame data includes the anchor frame position and the anchor frame size;
[0014] Extract and classify the anchor frame data to obtain the application anchor frame;
[0015] Output training samples using the applied anchor box.
[0016] In one embodiment, coordinate marking of a data set to obtain a training sample includes:
[0017] Crop the data set to obtain the cropped area;
[0018] Classify the cropped area and determine whether it is the target area, and obtain the target area image and the non-target area image;
[0019] The target area images and non-target area images constitute the training samples.
[0020] In one embodiment, using the training samples to train the pre-trained model according to the learning rate to obtain the target model includes:
[0021] Use convolutional neural networks to predict the key point coordinates of training samples;
[0022] Transform the key point coordinates to generate a standard input image;
[0023] Perform feature extraction on the standard input image to obtain feature data;
[0024] Convert feature data into feature sequences for sequence modeling;
[0025] Output the prediction results of the standard input image and determine the target model.
[0026] In one embodiment, the target model is an Attention HTR model;
[0027] The Attention HTR model consists of sequentially connected thin plate splines, a 32-layer residual neural network, a bidirectional long short-term memory network, and an attention-based decoder.
[0028] The thin plate spline localization network is used to generate standard input images;
[0029] 32-layer residual neural network is used for feature extraction;
[0030] Bidirectional long short-term memory network is used to generate feature sequences;
[0031] The attention-based decoder is used to output prediction results.
[0032] In one embodiment, the attention-based decoder outputs predictions by:
[0033] Use the unidirectional long short-term memory network to make predictions and output some preliminary prediction results;
[0034] The probability of a single preliminary prediction is calculated using the following formula:
[0035]
[0036] Z=[z1,z2,…,z n ]
[0037] Among them, Z represents the preliminary prediction result, Z i represents the i-th preliminary prediction result; Represents exponential operation, which performs exponential transformation on each preliminary prediction result; It represents the sum of the results after exponential changes, which is used for normalization;
[0038] The preliminary prediction result with the highest probability is selected as the prediction result.
[0039] In one embodiment, before predicting the key point coordinates of the training sample using a convolutional neural network, the method further includes:
[0040] Use filters to enhance the image edges of training samples;
[0041] Perform horizontal projection on the image to obtain a horizontal projection curve;
[0042] Determine the text position based on the horizontal projection curve;
[0043] According to the set threshold, the image is cropped into separate rows according to the text position to obtain a row image; the set threshold is obtained according to the middle value of the highest value and the lowest value of the horizontal projection curve;
[0044] Project the row image vertically to obtain a vertical projection curve and obtain space data;
[0045] Calculate the average value of the space data;
[0046] At the position where the space data is greater than the average value, the word image is segmented.
[0047] In one embodiment, before coordinate marking of a data set to obtain training samples, the following steps are included:
[0048] Perform image normalization on the size and color of image files;
[0049] The data set is enhanced; data enhancement includes horizontal flipping, noise enhancement, and brightness adjustment.
[0050] In one embodiment, it also includes:
[0051] Obtain the test set and input the test set into the target model to evaluate the model performance;
[0052] Optimization is performed based on model performance, including adjustment of learning rate and target model parameters.
[0053] In a second aspect, the present application also provides an artificial intelligence model construction device based on deep learning, comprising:
[0054] The data processing module is used to obtain the data set, mark the coordinates of the data set, and obtain training samples;
[0055] Model loading module, used to load the feature architecture of the pre-trained model and add an extended architecture based on the feature architecture;
[0056] The training setting module is used to freeze the feature architecture and adjust the weight of the extended architecture to the set value; it is set to unfreeze the feature architecture when the learning rate increases to the set value;
[0057] The model building module is used to train the pre-trained model using the training samples according to the learning rate to obtain the target model.
[0058] The above-mentioned method and device for constructing an artificial intelligence model based on deep learning, by loading the feature architecture in the pre-trained model and adding an extended architecture on its basis, utilizes the general features in the pre-trained model that conform to the target model and learns the task-related features through the extended architecture, so that the model can quickly transfer from known feature relationships to feature learning of specific tasks, avoiding the high computing cost and time of training the model from scratch, improving the model accuracy and efficiency of specific tasks, and is also suitable for model training situations where there is a lack of large amounts of data. The training is set to freeze the feature architecture and only update the extended architecture parameters. The feature architecture is staged after the learning rate reaches the set value, so as to prevent the effective features learned in the pre-trained model from being significantly adjusted during the initial training of the model. At the same time, it is ensured that after the feature architecture is gradually unfrozen in the later stage, the model can be further optimized without destroying the original features, avoiding the risk of overfitting and degradation of model performance during the model establishment process. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0060] Figure 1 This is an application environment diagram of the method for building an artificial intelligence model based on deep learning of the present invention;
[0061] Figure 2 A flowchart of a method for building an artificial intelligence model based on deep learning according to the present invention;
[0062] Figure 3 This is a flow chart of step S201 of the method for building an artificial intelligence model based on deep learning of the present invention;
[0063] Figure 4 This is a flow chart of step S201 of the method for building an artificial intelligence model based on deep learning of the present invention;
[0064] Figure 5 This is a flow chart of step S204 of the method for building an artificial intelligence model based on deep learning of the present invention;
[0065] Figure 6 This is a flow chart of the byte division method in the method for building an artificial intelligence model based on deep learning of the present invention;
[0066] Figure 7 This is a schematic diagram of the structure of the device for building an artificial intelligence model based on deep learning in the present invention. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0068] The deep learning-based artificial intelligence model construction method provided in the embodiments of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or it can be placed on the cloud or other network servers. The terminal 101 can be an operating system such as Windows, Linux, UNIX, Macos, etc., and the hardware is equipped with a CPU and GPU processor and memory and other memories that are sufficient to carry the machine deep learning model. The memory stores image data, label files, and data set configuration files. The software is equipped with Python (computer programming software) and above, and uses deep learning frameworks such as PyTorch and TensorFlow for model construction and training, and configures the CUDA (Compute Unified Device Architecture) environment. In addition, the terminal 101 can be, but is not limited to, various personal computers, laptops, tablets, etc. The server 102 can be implemented with an independent server or a server cluster consisting of multiple servers, and the model can be run on the server for online services.
[0069] In one embodiment, Figure 2 As shown, a method for building an artificial intelligence model based on deep learning is provided, which can be applied to the following application scenarios:
[0070] Document digitization: In the digitization process of historical documents and ancient books, marginal notes often contain important explanations, comments or contextual information. By identifying these marginal notes, we can have a more comprehensive understanding of the original text and enhance the research value of the document.
[0071] Academic research: In academic publications and research papers, marginalia and annotations are often used to provide additional insights or analysis of the main content. Identifying these marginalia can help researchers quickly obtain useful information and conduct literature analysis.
[0072] Educational applications: During the teaching process, students may make marginal notes or annotations on books or handouts. Through marginal note recognition technology, these handwritten annotations can be digitized, making it easier for students to review and search, and enhancing learning effects.
[0073] Legal documents and contracts: Legal documents, contracts, or agreements often have important marginal notes or amendments. By automatically identifying these marginal notes, legal workers can quickly review and organize documents, improving work efficiency.
[0074] Meeting minutes and notes: In meeting minutes or notes, marginal notes are often used to record additional comments or remarks. Automatically identifying these marginal notes can help quickly organize and archive meeting content and improve the efficiency of information retrieval.
[0075] In this embodiment, the method includes the following steps:
[0076] S201, obtain a data set, and mark the coordinates of the data set to obtain a training sample. The data set includes an image file.
[0077] The training samples are obtained by marking the coordinates of the data set. Schematically, the R-CNN (Region-based Convolutional Neural Network) or FR-CNN (Fast Region-based Convolutional Neural Network) method is used for coordinate marking, and the marking results are integrated into training samples to provide necessary basic data for subsequent model training.
[0078] S202, load the feature architecture of the pre-trained model, and add an extended architecture based on the feature architecture. The feature architecture is a module in which the pre-trained model has obtained the target feature relationship, the target feature relationship indicates that the target model outputs the expected result, and the extended architecture indicates that the target model generates the expected output after learning the effective features.
[0079] Since a large number of data sources and data annotations are required when training high-precision models, in specific application demand scenarios, data scarcity will lead to inadequate extraction of target feature relationships of deep learning artificial intelligence models. Therefore, a pre-trained model is loaded and fine-tuned using the pre-trained model through transfer learning methods to effectively improve the performance of the model on small data sets. Through transfer learning, the model is now pre-trained on a large data set, and then trained according to specific scenarios and specific data to improve the generalization ability of the model. Exemplarily, the Attention HTR target model is trained, the CNN (Convolutional Neural Network) feature extraction part of the scene text recognition model is pre-loaded, and the convolutional layer weights of the pre-trained model, such as ResNet (Residual Network), are retained for the feature extraction stage of the Attention HTR target model. In addition, a bidirectional long short-term memory network and an attention mechanism are added to the extended architecture of the character sequence modeling and character-by-character prediction module.
[0080] Different from the staged processing during traditional model building and training, the method of adding an extended architecture based on the feature architecture focuses on the integrity of the architecture, allowing the result output process to be jointly optimized in the same model architecture. Schematically, in the field of text recognition applications, the model with an overall architecture, from input images to output text results, all processing steps such as feature extraction, sequence modeling, decoding, etc. are carried out in the same model framework.
[0081] S203, freeze the feature architecture, and adjust the weight of the extended architecture to a set value. When the learning rate increases to a set value, unfreeze the feature architecture, and the learning rate is obtained based on the maximum learning rate that does not damage the target feature relationship.
[0082] Freeze the convolutional layers of the feature architecture in the pre-trained model to ensure that its weights do not change. Only adjust the relevant weights of the decoder of the upper-layer model such as the bidirectional long short-term memory network and the attention mechanism, so as to avoid overfitting of the target feature relationship and cause input and output deformation. By dynamically adjusting the learning rate, the model can adapt to different learning needs at different training stages, converge quickly and avoid overfitting. Regularization techniques such as weight decay are used schematically to prevent the model from overfitting on small data sets. In the process of gradually increasing from a smaller learning rate, the training degree of the model will gradually increase, and the integrity will gradually stabilize. At the same time, the convolutional layer is unfrozen synchronously to ensure that the model weights will not change drastically and damage the target feature relationship representation of the model.
[0083] S204: Use the training samples to train the pre-trained model according to the learning rate to obtain a target model.
[0084] By utilizing transfer learning and fine-tuning based on the pre-trained model, the computing resources and time required for training are shortened. This can also be used in scenarios where there is a large amount of insufficient data to make up for the problem of insufficient data. The accuracy and efficiency of the model can be further improved through targeted data training. The target feature relationships obtained by the pre-trained model in large data sets can help the target model converge quickly, while also improving the generalization ability of the target model.
[0085] Through transfer learning with pre-trained models, the target feature relationship parts of the pre-trained models that match the target model are retained and utilized. On this basis, extension modules are added to increase the functionality of the model to ensure that the input data produces the expected output. Then, a small data set that matches the target model is used for training. This greatly reduces the computational requirements and training time for model building. The big data features of the pre-trained model and the refinement of the small data set are used to enhance the generalization ability of the model, improve the accuracy and efficiency of the model output, and enable the model to be established flexibly to meet the model building requirements of various application scenarios. In text recognition applications, the model can gradually decode the characters in the image, generate accurate text sequences, and achieve the expected recognition effect. This architectural extension not only improves the flexibility and expressiveness of the model, but also ensures that the model is capable of handling more complex text recognition tasks.
[0086] In one embodiment, if Figure 3 As shown, the coordinates of the data set are marked to obtain training samples, which may include:
[0087] S301. Cover a number of initial anchor boxes in the data set, where the initial anchor boxes are candidate regions of different scales and aspect ratios.
[0088] Faster R-CNN combines the Region Proposal Network (RPN) and Fast R-CNN. RPN generates region proposals and shares convolution calculations with Fast R-CNN to reduce computing time. Schematically, a pre-trained ResNet-50 network (residual network) is used as a feature extractor. The ResNet-50 network is used to extract high-level features from the input image. The use of the ResNet network reduces the gradient vanishing problem when training deep networks. First, RPN generates a large number of anchor boxes, which are candidate regions of different scales and aspect ratios, covering individual data in the entire dataset, such as image text.
[0089] S302: Extract and classify the anchor frame data to obtain an application anchor frame.
[0090] For each anchor box, RPN predicts whether the area includes the target object, such as a marginal note. The candidate area box determined by RPN as possibly containing a marginal note will enter Fast R-CNN, where Fast R-CNN extracts features and regresses a more accurate bounding box, i.e., applies the anchor box, to further accurately locate the marginal note area.
[0091] S303: Output training samples using the application anchor frame.
[0092] By accurately applying the anchor boxes and rescanning the dataset, accurate training sample data is generated.
[0093] Running the model using the anchor box mechanism to generate multiple candidate areas at different scales and aspect ratios enables the model to detect complex target areas more flexibly. Through model learning, the anchor box can dynamically adjust its position and size to ultimately accurately locate the target area.
[0094] In one embodiment, if Figure 4 As shown, the coordinates of the data set are marked to obtain training samples, which may include:
[0095] S401: Crop the data set to obtain a cropping area.
[0096] Use R-CNN to use the region proposal method to generate a region of interest (ROI) on the image and crop the region of interest into a cropped area;
[0097] S402: Classify the cropped area and determine whether it is a target area, thereby obtaining a target area image and a non-target area image.
[0098] AlexNet (deep convolutional neural network) is used to classify the cropped area into target area images and non-target area images according to the target situation.
[0099] S403: The target area image and the non-target area image constitute training samples.
[0100] In one embodiment, if Figure 5 As shown, using the training samples to train the pre-trained model according to the learning rate to obtain the target model may include:
[0101] S501, using a convolutional neural network to predict the key point coordinates of the training samples.
[0102] The training samples are input into the convolutional neural network. CNN processes the input image through multiple layers of convolution, pooling and activation functions to automatically learn local features. CNN outputs the predicted key point coordinates, which usually refer to important positions in the image, such as the corners or center points of characters. These key points are used for subsequent image transformation.
[0103] S502: transform the key point coordinates to generate a standard input image.
[0104] Schematically, the predicted key point coordinates are used to transform the image to generate a standard input image. This step usually includes geometric transformations such as rotation, translation, and scaling to ensure that the target part in the image is within a unified standard framework and improve the effect of subsequent feature extraction. For example, handwritten tilted text is straightened according to the coordinate change. The generation of a standard input image can improve the robustness of the model to different writing styles, fonts, and tilted text.
[0105] S503: Extract features from the standard input image to obtain feature data.
[0106] The transformed standard input image is input into a CNN or other feature extraction network to extract high-level features in the image. For example, these features can represent the shape, texture, and structural information of the characters in the image, helping the model to better understand the image content.
[0107] S504: Convert feature data into feature sequences for sequence modeling.
[0108] The extracted feature data is converted into a feature sequence, usually by flattening the feature map or using a specific processing method such as a recurrent neural network to form sequence data.
[0109] S505: Output the prediction result of the standard input image and determine the target model.
[0110] Different from the traditional method in which feature extraction, sequence modeling and classifier are usually decoupled, that is, features are manually extracted first and then input into the classifier, the feature extraction, sequence modeling and prediction results of the present invention are tightly coupled, trained together and adjusted with each other, so that the features finally extracted are more suitable for the target task, and the accuracy of the model is improved. Moreover, feature extraction is obtained from data and can be dynamically adjusted according to task requirements without relying on manual design.
[0111] Convolutional neural networks can automatically learn features in images without manually designing feature extractors. Compared with traditional methods, this automation can improve the efficiency and effectiveness of feature extraction. By accurately predicting key points, the generated standard input image can effectively reduce noise and deformation in the image, improve the quality of subsequent feature extraction, and thus enhance the overall recognition performance of the model. Through geometric transformation and standardization, the model can better cope with changes in different writing styles, character spacing, and inclination, and enhance robustness to diverse inputs. The use of sequence modeling methods, such as BLSTM (Bidirectional Long Short-Term Memory), can make full use of the contextual relationship between characters, thereby improving the accuracy of character recognition, especially when processing long sequences or complex texts.
[0112] In one embodiment, the target model is an Attention HTR model:
[0113] The Attention HTR model consists of sequentially connected thin plate splines, a 32-layer residual neural network, a bidirectional long short-term memory network, and an attention-based decoder.
[0114] (1) The thin plate spline localization network is used to generate the standard input image;
[0115] The input handwritten marginalia image is normalized using Thin Plate Spline (TPS) transformation to eliminate irregular shapes such as tilt and curvature in the image. The key coordinates are predicted by Convolutional Neural Network (CNN) to generate a standardized input image.
[0116] (2) 32-layer residual neural network for feature extraction;
[0117] A 32-layer residual neural network (ResNet) is used to extract two-dimensional visual features from the transformed image and generate a column feature map, where each column represents a part of the image, consistent with the left-to-right order of the input image.
[0118] (3) Bidirectional long short-term memory network is used to generate feature sequences;
[0119] Bidirectional long short-term memory network (BLSTM) is used for sequence modeling. LSTM is a special recursive neural network (RNN) that can effectively capture long-term dependencies in time series. BLSTM can process sequence information from both the front and back directions, which helps to more accurately recognize the character sequence of handwritten text. Schematically, BLSTM is used for marginal handwritten text recognition. BLSTM converts the extracted feature map into a feature sequence and performs sequence modeling. By capturing the relationship of the context, character recognition is made more accurate and stable.
[0120] (4) The attention-based decoder is used to output prediction results.
[0121] A sequence-to-sequence model based on the attention mechanism, which pays more attention to the important parts of the input sequence through the attention mechanism, thereby improving the accuracy of text recognition. In handwritten text recognition, the attention mechanism helps the model focus on specific characters or words when recognizing long sequences. The decoder uses a unidirectional LSTM, combined with the context vector and hidden state of the BLSTM, to gradually output characters. The decoder can generate sequences of variable length until the "end sequence" symbol (EOS) is encountered. Specifically, the attention mechanism is used to find the most relevant features in the input image for each practice part during decoding, and each feature of the input sequence is weighted according to the needs of the current prediction, that is, a set of weights is calculated for each time step to represent the correlation between the time step and the input feature. These weights can be normalized using the Softmax function to form an attention distribution, guiding the decoder to focus on the most relevant part of the input feature, and the decoder forms a context vector by weighting and calculating the feature vector. This vector combines the most important information in the input feature for the current prediction, thereby generating recognized text.
[0122] Generate an Attention HTR model, which includes CNN, BLSTM and a decoder based on the attention mechanism. CNN can directly extract hierarchical features without manually designing features, making the model adaptable to complex application scenarios, especially in the field of text recognition, where it performs well under complex backgrounds and diverse fonts. BLSTM can capture text sequences in combination with context. Compared with the limitations of traditional methods that process characters independently, it performs well when there are dependencies between characters or words and is more suitable for long text recognition. The decoder based on the attention mechanism allows the decoding process to focus more on relevant areas, improving the accuracy of model predictions.
[0123] In one embodiment, the attention-based decoder outputs predictions by:
[0124] S61. Use a unidirectional long short-term memory network to make predictions and output some preliminary prediction results.
[0125] The probability of a single preliminary prediction is calculated using the following formula:
[0126]
[0127] Z=[z1,z2,…,z n ]
[0128] Among them, Z represents the preliminary prediction result, Z i represents the i-th preliminary prediction result; Represents exponential operation, which performs exponential transformation on each preliminary prediction result; It represents the sum of the results after exponential changes, which is used for normalization;
[0129] Schematically, the Softmax function converts the preliminary prediction result into a value between 0 and 1, and the sum of all output values is 1 to explain the probability of each preliminary prediction result. The Softmax function assigns probabilities according to the relative size of the input, with larger input values corresponding to higher probabilities and smaller input values corresponding to lower probabilities. For example, in character recognition, the model outputs a probability distribution of character categories, and the Softmax function can convert the original network output into probability values for each character category.
[0130] S62. Select the preliminary prediction result with the highest probability as the prediction result.
[0131] The character probability of each step is calculated through the Softmax function, and the character with the highest probability is selected as the prediction result.
[0132] In one embodiment, if Figure 6 As shown, before using the convolutional neural network to predict the key point coordinates of the training sample, it also includes:
[0133] S601: Use a filter to enhance the image edge of the training sample.
[0134] Optionally, a Sobel filter is used to highlight edge information.
[0135] S602: Perform horizontal projection on the image to obtain a horizontal projection curve.
[0136] The pixel density is analyzed through horizontal projection. By calculating the gray value of each row of pixels or the number of non-blank pixels in the image, a horizontal projection curve is generated. The horizontal projection curve shows the pixel density at different positions in the text image.
[0137] S603: Determine the text position according to the horizontal projection curve.
[0138] The position of the text line is determined by pixel density. When the pixel value of a certain part is larger, it means that the part may be the line where the text is located, and when the pixel value is smaller, it means that the area may be the blank area between lines.
[0139] S604 , according to the set threshold, the image is cut into separate rows according to the text position to obtain a row image; the set threshold is obtained according to the middle value of the highest value and the lowest value of the horizontal projection curve.
[0140] Therefore, the position of the line can be determined according to the horizontal projection, and the image is cropped according to the position of the text line to separate the independent lines, that is, the entire text image is divided into lines of word images.
[0141] S605 , vertically project the row image to obtain a vertical projection curve and obtain space data.
[0142] The row image is vertically projected to find the gaps between words, that is, each row of the image is vertically projected. The vertical projection generates a vertical projection curve by calculating the number of non-blank pixels in each column of pixels. When the pixel density of a column is, it usually represents the gap between characters or words, which is used to identify the blank areas between words or characters in the row.
[0143] S606: Calculate the average value of the space data.
[0144] The average white space between words or characters, i.e. the spacing, is calculated to derive the judgment value.
[0145] S607, segmenting the word image at the position where the space data is greater than the average value.
[0146] Segmentation is performed only when the gap is larger than the average value to avoid over-segmentation, and the word image is obtained through segmentation.
[0147] The byte division method of the vertical and horizontal projection methods belongs to the traditional image processing technology, but it can be used in combination with the deep learning method, that is, the divided word image is further input into the convolutional neural network for feature extraction, and then the character recognition is performed through the RNN or attention mechanism. The byte division method is more dependent on the neatness of the text and is generally used for printed text. It will cause segmentation errors in unstructured, complex background or curved handwritten text. The end-to-end model generally does not need to perform byte division in the early stage. It can directly learn features from the text image and perform text recognition, but the end-to-end requires a lot of data and computing resources for training. The combination of the two methods helps to reduce the computational burden of the model, especially when the text recognition model is applied to structured documents such as invoices, tables, and document texts. The byte division method effectively focuses on local areas and improves the efficiency of the model operation. When facing complex text, such as handwriting or text recognition in natural scenes, the model directly calls the end-to-end method without relying on byte division and recognizes text based on feature relationships. The combination of the two methods not only improves the model efficiency and performance but also reduces manual intervention in the model operation process.
[0148] In one embodiment, before coordinate marking of a data set to obtain training samples, the following steps are included:
[0149] S71, performing image normalization processing on the size and color of the image file.
[0150] The image files in the dataset are standardized to ensure that the size and channels of each image are consistent.
[0151] S72, perform data enhancement on the data set. Data enhancement includes horizontal flipping, noise enhancement, and brightness adjustment.
[0152] To improve the robustness of the model, data augmentation techniques are usually used to generate more diverse training data by flipping the training data set, adding noise, adjusting brightness, etc., to help the model improve its generalization ability.
[0153] In one embodiment, it also includes:
[0154] S81. Obtain a test set and input the test set into the target model to evaluate the model performance.
[0155] Exemplarily, the gap between the predicted character sequence and the actual character sequence is calculated, the difference between the predicted word and the actual word is calculated, and the recognition effect of the complete word is evaluated.
[0156] S82. Perform optimization processing according to the model performance, wherein the optimization processing includes adjusting the learning rate and target model parameters.
[0157] Debugging of hyperparameters such as learning rate, size, and convolutional layer stage strategy can be adjusted based on the performance of the model on the test set to ensure maximum model performance.
[0158] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0159] Based on the same inventive concept, the embodiment of the present application also provides a device for building an artificial intelligence model based on deep learning for implementing the above-mentioned method for building an artificial intelligence model based on deep learning. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in one or more embodiments of the artificial intelligence model building device based on deep learning provided below can be referred to the limitations of the artificial intelligence model building method based on deep learning above, and will not be repeated here.
[0160] In an exemplary embodiment, Figure 7 As shown, the present application provides an artificial intelligence model construction device based on deep learning, comprising:
[0161] The data processing module 701 is used to obtain a data set and mark the coordinates of the data set to obtain a training sample;
[0162] A model loading module 702 is used to load the feature architecture of the pre-trained model and add an extended architecture based on the feature architecture;
[0163] The training setting module 703 is used to freeze the feature architecture and adjust the weight of the extended architecture to a set value; when the learning rate increases to a set value, the feature architecture is unfrozen;
[0164] The model building module 704 is used to train the pre-trained model using the training samples according to the learning rate to obtain the target model.
[0165] In one of the embodiments, a deep learning data processing module is also included, which is used to cover a number of initial anchor frames in the data set, determine whether the initial anchor frame contains the target object, obtain anchor frame data of the initial anchor frame containing the target object, extract and classify features of the anchor frame data, and obtain an application anchor frame.
[0166] In one embodiment, a deep learning data processing module is further included, which is used to crop the data set to obtain a cropped area, classify the cropped area, and determine whether it is a target area to obtain a target area image and a non-target area image.
[0167] In one embodiment, the model building module 704 is also used to predict the key point coordinates of the training sample using a convolutional neural network; transform the key point coordinates to generate a standard input image; extract features from the standard input image to obtain feature data; convert the feature data into a feature sequence for sequence modeling; output the prediction result of the standard input image, and determine the target model.
[0168] In one embodiment, the model building module 704 is also used to use a unidirectional long short-term memory network to perform predictions and output a number of preliminary prediction results; and the preliminary prediction result with the highest probability is selected as the prediction result.
[0169] In one embodiment, the data processing module 701 is also used to enhance the image edge of the training sample using a filter; horizontally project the image to obtain a horizontal projection curve; determine the text position based on the horizontal projection curve; according to a set threshold, crop the image into separate rows according to the text position to obtain a row image; vertically project the row image to obtain a vertical projection curve to obtain space data; calculate the average value of the space data; and segment the word image at the position where the space data is greater than the average value.
[0170] In one embodiment, the data processing module 701 is further used to perform image normalization processing on the size and color of the image file and perform data enhancement on the data set. The data enhancement includes horizontal flipping, noise enhancement and brightness adjustment.
[0171] In one embodiment, the training setting module 703 is also used to obtain a test set, input the test set into the target model, and evaluate the model performance; perform optimization processing according to the model performance, and the optimization processing includes adjusting the learning rate and target model parameters.
[0172] For example, taking the construction of the Attention HTR model as an example, we can further help understand the deep learning-based artificial intelligence model construction method provided by this application. First, the application scenario and goal of the model are clarified as marginal annotated font recognition. Therefore, we collect and prepare a high-quality data set, that is, a certain amount of images with marginal annotations and corresponding text labels.
[0173] Before deep learning model training, data needs to be preprocessed to ensure input consistency and data quality, and the input images need to be standardized to ensure that the size and color channels of each input image are consistent. In order to increase the generalization ability of the model, the data can be enhanced, including rotation, translation, scaling, noise addition, etc.
[0174] Load the ResNet+BLSTM model pre-trained on the scene text recognition task as the initial weight. The pre-trained ResNet+BLSTM model has been trained with a large amount of data from standard datasets such as IAM, Imgur5K, and scene datasets such as MJSynth and SynthText, and retains the convolutional layer feature relationship available for text recognition. On this basis, add the architecture, namely the Attention HTR model, which uses the pre-trained convolutional neural network ResNet for feature extraction: input the standardized image into ResNet to generate a two-dimensional feature map, where each column of the feature map represents a part of the image and corresponds to the pixel points of the image; use BLSTM for sequence modeling: convert the feature map output by ResNet into a feature sequence. Bidirectional BLSTM is used to capture contextual information to ensure that the model can use the connection between previous and next characters for accurate character recognition. The generated feature sequence will be used for the next decoding prediction. The decoder with attention mechanism is used for character sequence prediction: a context vector is calculated for each time step, and the current character is predicted through the weighted feature sequence. A unidirectional LSTM is used to output the character sequence step by step until the end symbol (EOS) is encountered. At each time step, the output of LSTM is converted into the probability distribution of the character through the Softmax function, and the character with the highest probability is selected as the prediction result, and the marginal recognition text is output.
[0175] After training is complete, the model is deployed to the application environment for actual text recognition tasks. The model can be exported and deployed to servers, mobile devices, or edge devices using deep learning frameworks such as TensorFlow, PyTorch, or ONNX.
[0176] The Attention HTR model is applied to marginal notes recognition to identify and extract handwritten marginal notes or additional information written on books, documents or other texts.
[0177] The above-mentioned embodiments only express several implementation methods of the embodiments of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent of the embodiments of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the embodiments of the present application, and these all belong to the protection scope of the embodiments of the present application.
Claims
1. A method for constructing an artificial intelligence model based on deep learning, characterized in that: The method comprises: Acquire a data set, and mark the coordinates of the data set to obtain a training sample; the data set includes an image file; Loading a feature architecture of a pre-trained model, adding an extended architecture based on the feature architecture; the feature architecture is a module of the pre-trained model that has obtained a target feature relationship; the target feature relationship indicates that the target model outputs an expected result; the extended architecture indicates that the target model generates an expected output after learning effective features; Freeze the feature architecture and adjust the weight of the extended architecture to a set value; set the feature architecture to be unfrozen when the learning rate increases to a set value; the learning rate is derived based on a maximum learning rate that does not damage the target feature relationship; The pre-training model is trained using the training samples according to the learning rate to obtain a target model.
2. The method according to claim 1, characterized in that The step of marking the coordinates of the data set to obtain training samples includes: Covering a plurality of initial anchor frames in the data set; the initial anchor frames are candidate regions of different scales and aspect ratios; Determine whether the initial anchor frame contains the target object, and obtain anchor frame data of the initial anchor frame containing the target object; the anchor frame data includes the anchor frame position and the anchor frame size; Extracting and classifying the anchor frame data to obtain an application anchor frame; The training samples are outputted using the application anchor frame.
3. The method according to claim 1, characterized in that The step of marking the coordinates of the data set to obtain training samples includes: Cropping the data set to obtain a cropping area; Classifying the cropped area and determining whether it is a target area, thereby obtaining a target area image and a non-target area image; The target area image and the non-target area image constitute the training sample.
4. The method according to claim 1, characterized in that: The step of training the pre-trained model using the training sample according to the learning rate to obtain a target model includes: Predicting the key point coordinates of the training sample using a convolutional neural network; Transforming the key point coordinates to generate a standard input image; Perform feature extraction on the standard input image to obtain feature data; Converting the feature data into a feature sequence for sequence modeling; The prediction result of the standard input image is output, and the target model is determined.
5. The method according to claim 4, characterized in that The target model is the Attention HTR model; The Attention HTR model includes sequentially connected thin plate splines, a 32-layer residual neural network, a bidirectional long short-term memory network, and a decoder based on an attention mechanism; The positioning network of the thin plate spline is used to generate the standard input image; The 32-layer residual neural network is used for feature extraction; A bidirectional long short-term memory network is used to generate the feature sequence; The attention mechanism-based decoder is used to output the prediction result.
6. The method according to claim 5, characterized in that The attention-based decoder outputs predictions in the following way: Use the unidirectional long short-term memory network to make predictions and output some preliminary prediction results; The probability of a single preliminary prediction result is calculated using the following formula: Z=[z1,z2,…,z n ] Among them, Z represents the preliminary prediction result, Z i represents the i-th preliminary prediction result; Represents exponential operation, which performs exponential transformation on each preliminary prediction result; It represents the sum of the results after exponential changes, which is used for normalization; The preliminary prediction result with the highest probability is selected as the prediction result.
7. The method according to claim 4, characterized in that Before predicting the key point coordinates of the training sample using a convolutional neural network, the method further includes: Using a filter to enhance the image edge of the training sample; Performing horizontal projection on the image to obtain a horizontal projection curve; Determine the text position based on the horizontal projection curve; According to a set threshold, the image is cut into separate rows according to the text position to obtain a row image; the set threshold is obtained according to the middle value of the highest value and the lowest value of the horizontal projection curve; Projecting the row image vertically to obtain a vertical projection curve and obtain space data; Calculate the average value of the space data; At the position where the space data is greater than the average value, the word image is segmented.
8. The method according to claim 1, characterized in that: Before the coordinates of the data set are marked to obtain training samples, the following steps are included: Performing image normalization processing on the size and color of the image file; The data set is enhanced, wherein the data enhancement includes horizontal flipping, noise enhancement and brightness adjustment.
9. The method according to claim 1, characterized in that: Also includes: Obtain a test set, and input the test set into the target model to evaluate model performance; An optimization process is performed according to the model performance, wherein the optimization process includes adjusting the learning rate and the target model parameters.
10. A device for building an artificial intelligence model based on deep learning, characterized in that: The device comprises: A data processing module is used to obtain a data set and mark the coordinates of the data set to obtain a training sample; A model loading module is used to load the feature architecture of the pre-trained model and add an extended architecture based on the feature architecture; A training setting module, used to freeze the feature architecture and adjust the weight of the extended architecture to a set value; and to unfreeze the feature architecture when the learning rate increases to a set value; The model building module is used to train the pre-trained model using the training samples according to the learning rate to obtain a target model.
Citation Information
Patent Citations
Deep learning algorithm-based text recognition method, device, equipment and storage medium
CN112464945A
Font category visual detection method and system based on active learning and transfer learning
CN117649672A
Industrial character recognition method and device based on small sample target detection and storage medium
CN117809306A