A character recognition method and device based on PP-OCRv3 transfer learning
Patent Information
- Application Number
- CN202510168380.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-17
AI Technical Summary
In the coal mine environment, digital screen character recognition is affected by factors such as lighting conditions, acquisition technology, and equipment failure, resulting in low contrast, blurred, deformation of characters in the image, and interference lines and interference blocks appearing, making it difficult to distinguish the background area and character area, and is affected by the working environment and character diversity, resulting in small sample difficulties.
The character recognition method based on PP-OCRv3 transfer learning is adopted to train the PP-OCRv3 network through a public data set to build a pre-trained model, and combine a small amount of actual data and simulated data to fine-tune the pre-trained model multiple times to reduce the impact of the environment on recognition performance and improve the recognition accuracy and speed.
It effectively improves the accuracy and speed of character recognition of digital display screens of coal mines, reduces environmental impact, improves the generalization ability and robustness of the model, and solves the problem of insufficient data samples.
Smart Images

Figure CN119625738B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text image recognition, and in particular to a character recognition method and device based on PP-OCRv3 transfer learning. Background Art
[0002] Coal is the mainstay energy source for production and consumption in my country. As a "sensor" in coal production, the digital display screen can read and display various parameters of the equipment, reflect the operating status of the coal mine in real time, and maintain the safe and stable operation of coal production. In the process of coal production safety, the data of the equipment digital display screen is generally obtained through built-in acquisition devices and assisted by manual methods. Built-in devices are easily affected by the network environment and may cause transmission delays or failures and inability to obtain data. The manual method has the disadvantages of high financial resources, unstable data collection, low timeliness, safety hazards for staff, manual data misdetection and omission, and human tampering. In order to solve the above-mentioned problems of inaccurate data acquisition and data tampering, the research and application of automatic character recognition technology for digital display screens is particularly critical.
[0003] At present, in typical industrial scenarios, cameras are usually used to collect screen image information, and then target detection and recognition technology is used to extract character information. Optical Character Recognition (OCR), as a common text recognition technology, has achieved good results in many fields. However, it still faces some challenges in the coal mine environment: First, due to factors such as lighting conditions, acquisition technology, and equipment failure, the characters in the image have low contrast, blur, deformation, interference lines and interference blocks, etc.; the font color is similar to the background color, making it difficult to distinguish between the background area and the character area; second, due to the constraints of the working environment and the diversity of characters, small samples are difficult.
[0004] Therefore, how to invent a character recognition method based on PP-OCRv3 transfer learning to effectively reduce the impact of the environment on recognition performance and improve recognition accuracy and speed has become an urgent problem to be solved. Summary of the invention
[0005] To this end, the present invention provides a character recognition method and device based on PP-OCRv3 transfer learning. The method trains the PP-OCRv3 network through a public data set to build a pre-trained model; then based on the idea of transfer learning, the pre-trained model is fine-tuned multiple times in combination with a small amount of actual data and simulated data, effectively reducing the impact of the environment on the recognition performance and improving the recognition accuracy and speed.
[0006] In order to achieve the above object, the present invention provides the following technical solution: a character recognition method based on PP-OCRv3 transfer learning, comprising:
[0007] Build a pre-trained model based on the PP-OCRv3 network;
[0008] Performing a first training study on the pre-trained model using a public text recognition data set to obtain a first model parameter of the pre-trained model;
[0009] Substituting the pre-trained model into the first model parameters, performing a second training study on the pre-trained model using a laboratory-made digital display screen character data set, and obtaining second model parameters of the pre-trained model;
[0010] Substituting the pre-trained model into the second model parameters, iteratively training the pre-trained model through part of the real data and the simulated data until the loss function satisfies the set loss value, completing the training and learning, and obtaining the trained pre-trained model;
[0011] The target text image is input into the trained pre-trained model for recognition processing, and the recognition result is output.
[0012] As a preferred solution of a character recognition method based on PP-OCRv3 transfer learning, during the training and learning process of the pre-trained model, the pre-trained model is trained and learned by a model-based transfer learning strategy; the model-based transfer learning strategy expression is:
[0013] ;
[0014] Where w is the source task model parameter; w' is the model parameter after migration; L is the loss function of the model; D is the target task dataset; is the model; R(w) is the regularization; is the regularization coefficient.
[0015] As a preferred solution of a character recognition method based on PP-OCRv3 transfer learning, during the training and learning process of the pre-trained model, a detection module loss function and a recognition module loss function are respectively constructed based on character detection and character recognition, and the parameters of the pre-trained model are iteratively updated through a back-propagation algorithm;
[0016] The expression of the detection module loss function is:
[0017] ;
[0018] Where, L total is the total loss function of the detection module; L p is the probability map loss; L b is the approximate binary image loss; L t is the threshold map loss; and All are weight coefficients;
[0019] The expression of the recognition module loss function is:
[0020] ;
[0021] In the formula, is the optimal character label sequence obtained by solving the model; X is the feature expression sequence containing character images, which are obtained by the detection module and corrected by the correction module; Y is the character label sequence obtained by solving; conditional probability It represents the probability that the character label is Y given that the input is X.
[0022] As a preferred solution of the character recognition method based on PP-OCRv3 transfer learning, the performance of the trained pre-trained model is evaluated by two indicators: accuracy and number of frames per second. The calculation formula of the accuracy is:
[0023] ;
[0024] In the formula, The number of positive samples identified as positive samples; The positive sample recognition result is the number of positive samples; P+N is the total number of samples in the data set.
[0025] As a preferred solution of a character recognition method based on PP-OCRv3 transfer learning, in the process of inputting the target text image into the trained pre-trained model for recognition processing, the target text image is detected by a text detection module, and a text detection result is output; the text detection result is calibrated by a text calibration module, and a calibration effect is output; the calibration effect is recognized by a text recognition module, and a recognition result is output.
[0026] The present invention also provides a character recognition device based on PP-OCRv3 transfer learning, based on the above character recognition method based on PP-OCRv3 transfer learning, comprising:
[0027] A pre-training model building unit, used to build a pre-training model based on the PP-OCRv3 network;
[0028] A pre-training model first training learning unit is used to perform a first training learning on the pre-training model through a text recognition public data set to obtain a first model parameter of the pre-training model;
[0029] A second training learning unit for the pre-trained model is used to substitute the pre-trained model into the first model parameters, perform a second training learning on the pre-trained model through a laboratory-made digital display screen character data set, and obtain second model parameters of the pre-trained model;
[0030] A pre-trained model iterative training learning unit, used to substitute the pre-trained model into the second model parameters, perform iterative training learning on the pre-trained model through part of the real data and the simulated data, until the loss function satisfies the set loss value, the training learning is completed, and the trained pre-trained model is obtained;
[0031] The text image recognition processing unit is used to input the target text image into the trained pre-trained model for recognition processing and output the recognition result.
[0032] As a preferred solution of a character recognition device based on PP-OCRv3 transfer learning, in the pre-training model first training learning unit, the pre-training model second training learning unit and the pre-training model iterative training learning unit, during the training and learning process of the pre-training model, the pre-training model is trained and learned by a model-based transfer learning strategy; the model-based transfer learning strategy expression is:
[0033] ;
[0034] Where w is the source task model parameter; w' is the model parameter after migration; L is the loss function of the model; D is the target task dataset; is the model; R(w) is the regularization; is the regularization coefficient.
[0035] As a preferred solution of a character recognition device based on PP-OCRv3 transfer learning, in the first training learning unit of the pre-training model, the second training learning unit of the pre-training model and the iterative training learning unit of the pre-training model, during the training and learning process of the pre-training model, a detection module loss function and a recognition module loss function are respectively constructed based on character detection and character recognition, and the parameters of the pre-training model are iteratively updated through a back-propagation algorithm;
[0036] The expression of the detection module loss function is:
[0037] ;
[0038] Where, L total is the total loss function of the detection module; L p is the probability map loss; L b is the approximate binary image loss; L tis the threshold map loss; and All are weight coefficients;
[0039] The expression of the recognition module loss function is:
[0040] ;
[0041] In the formula, is the optimal character label sequence obtained by solving the model; X is the feature expression sequence containing character images, which are obtained by the detection module and corrected by the correction module; Y is the character label sequence obtained by solving; conditional probability It represents the probability that the character label is Y given that the input is X.
[0042] As a preferred solution of a character recognition device based on PP-OCRv3 transfer learning, in the iterative training learning unit of the pre-trained model, the performance of the trained pre-trained model is evaluated by two indicators, accuracy and number of frames per second; the calculation formula of the accuracy is:
[0043] ;
[0044] In the formula, The number of positive samples identified as positive samples; The positive sample recognition result is the number of positive samples; P+N is the total number of samples in the data set.
[0045] As a preferred solution of a character recognition device based on PP-OCRv3 transfer learning, in the text image recognition processing unit, when the target text image is input into the trained pre-trained model for recognition processing, the target text image is detected by a text detection module, and a text detection result is output; the text detection result is calibrated by a text calibration module, and a calibration effect is output; the calibration effect is recognized by a text recognition module, and a recognition result is output.
[0046] The present invention has the following advantages: the present invention constructs a pre-training model based on a PP-OCRv3 network; the pre-training model is trained for the first time through a public data set for text recognition to obtain a first model parameter of the pre-training model; the pre-training model is substituted into the first model parameter, and the pre-training model is trained for the second time through a laboratory-made digital display screen character data set to obtain a second model parameter of the pre-training model; the pre-training model is substituted into the second model parameter, and the pre-training model is iteratively trained and learned through part of the real data and the simulated data until the loss function meets the set loss value, the training and learning are completed, and the trained pre-training model is obtained; the target text image is input into the trained pre-training model for recognition processing, and the recognition result is output. The present invention is based on deep learning and transfer learning, and proposes a coal mine digital display screen character recognition method based on PP-OCRv3 transfer learning. The method uses a public data set to train the PP-OCRv3 network to construct a pre-training model, and then based on the idea of transfer learning, the pre-training model is fine-tuned multiple times in combination with a small amount of actual data and simulated data, effectively reducing the impact of the environment on the recognition performance and improving the recognition accuracy and speed. Aiming at the problems of poor recognition effect in complex environment and insufficient sample data in digital display screen character recognition in coal mine environment, the present invention improves the recognition effect by introducing the PP-OCRv3 model with leading recognition effect in the OCR field in order to improve the problem of unsatisfactory recognition effect caused by factors such as small recognition area, complex illumination changes, and poor image quality in complex coal mine environment; in order to solve the problem of insufficient data samples, the PP-OCRv3 model is trained by multiple transfer learning methods to effectively reduce the model learning cost and avoid model overfitting. The present invention has significantly improved the recognition accuracy and generalization ability of characters on digital display screens in coal mines, laying a solid theoretical and technical foundation for the research on coal mine safety production and intelligent construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0048] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0049] Figure 1 A flowchart of a character recognition method based on PP-OCRv3 transfer learning provided in Example 1 of the present invention;
[0050] Figure 2 A schematic diagram of a specific implementation framework flow of a character recognition method based on PP-OCRv3 transfer learning provided in Example 1 of the present invention;
[0051] Figure 3 A schematic diagram of multiple migration training of a pre-trained model in a character recognition method based on PP-OCRv3 migration learning provided in Example 1 of the present invention;
[0052] Figure 4 A schematic diagram of a loss function drop curve during a pre-training model training process in a character recognition method based on PP-OCRv3 transfer learning provided in Example 1 of the present invention;
[0053] Figure 5 A schematic diagram of a character recognition process in a character recognition method based on PP-OCRv3 transfer learning provided in Example 1 of the present invention;
[0054] Figure 6 A schematic diagram of characters on a digital display screen in a coal mine environment in a character recognition method based on PP-OCRv3 transfer learning provided in Example 1 of the present invention;
[0055] Figure 7 A schematic diagram of a character image of a digital display screen simulating a coal mine in a possible embodiment provided in Embodiment 1 of the present invention;
[0056] Figure 8 Schematic diagram of character recognition results of a coal mine digital display screen in different scenarios in a possible embodiment provided in Embodiment 1 of the present invention; wherein (a) is the recognition result of the PP-OCRv3 model; (b) is the recognition result of the PP-OCRv3 transfer learning model; (c) is the recognition result under interference state;
[0057] Fig. 9 Schematic diagram of the architecture of a character recognition device based on PP-OCRv3 transfer learning provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0058] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] Example 1
[0060] See also Figure 1 and Figure 2 Embodiment 1 of the present invention provides a character recognition method based on PP-OCRv3 transfer learning, comprising the following steps:
[0061] S1. Build a pre-trained model based on the PP-OCRv3 network;
[0062] S2. Performing a first training study on the pre-trained model using a public text recognition data set to obtain a first model parameter of the pre-trained model;
[0063] S3, substituting the pre-trained model into the first model parameters, performing a second training study on the pre-trained model using a laboratory-made digital display screen character data set, and obtaining second model parameters of the pre-trained model;
[0064] S4, substituting the pre-trained model into the second model parameters, iteratively training the pre-trained model through part of the real data and the simulated data, until the loss function satisfies the set loss value, completing the training and learning, and obtaining the trained pre-trained model;
[0065] S5. Input the target text image into the trained pre-trained model for recognition processing, and output the recognition result.
[0066] In this embodiment, in step S1, a pre-trained model is constructed based on the PP-OCRv3 network;
[0067] Specifically, in view of the complex production environment of coal mines, the processed digital display screen images present the characteristics of non-uniform detection area size and complex background of recognized characters, and the early OCR technology cannot achieve the ideal recognition effect. The present invention introduces the PP-OCRv3 network with advanced text recognition accuracy as a pre-training model.
[0068] PP-OCRv3 is a practical ultra-lightweight OCR network that can meet the needs of industrial implementation. The network mainly includes three key modules: text detection, text box calibration and text recognition. Figure 5The text detection module uses the DB (Differentiable Binarization) text detection model to detect the input text image and output the text detection result, such as marking the location of the text in the image.
[0069] Among them, in the text calibration module, the text box calibration process uses an ultra-light backbone network. The ultra-light backbone network uses MobileNetV3, which is a lightweight neural network architecture that helps reduce the amount of calculation and improve the processing speed. The BDA and TIA data enhancement strategy algorithms are used to increase the diversity of training data and improve the generalization ability of the model. By increasing the resolution of the feature map, the model can obtain more detailed information. The text calibration module finally outputs the calibration effect, that is, the detected text box is corrected.
[0070] Among them, in the SVTR recognition model, the Progressive Overlapping Patch Embedding (POPE) method is used to process text features. After passing through multiple Local Mixing Blocks and Global Mixing Blocks, the text features are mixed and interacted locally and globally. Merging and output: After operations such as merging and combining, the final recognition result is output through the fully connected layer (FC), such as text content such as "refreshing and not fake white" and "moist and not greasy". In this way, starting from the text image input, it goes through the three links of detection, calibration and recognition in sequence, and finally obtains the recognition content of the text.
[0071] The accuracy test results in various scenarios show that the PP-OCRv3 model still achieves significant detection and recognition results under noise interference. In addition, PP-OCRv3 supports offline deployment and service deployment to meet the OCR application needs in complex environments in industrial scenarios.
[0072] In this embodiment, in step S2, the pre-trained model is trained for the first time using a public data set for text recognition to obtain a first model parameter of the pre-trained model;
[0073] Specifically, Figure 3 As shown, considering that some features of the models on different data sets are common, the first transfer learning uses a text recognition public data set to train the pre-trained model. After the model converges, the first model parameter of the pre-trained model is obtained, so that the pre-trained model learns the universal representation of the text.
[0074] Specifically, the pre-trained model PP-OCRv3 initializes the weight W 0 =(W 0 BD , W 0 SVTR ), the public text recognition dataset is the source task training sample, and the model converges to the parameter W 1= (W 1 BD , W 1 SVTR ).
[0075] In this embodiment, in step S3, the pre-trained model is substituted into the first model parameters, and the pre-trained model is trained for a second time using a laboratory-made digital display screen character data set to obtain second model parameters of the pre-trained model;
[0076] Specifically, in the second transfer learning, the pre-trained model parameters are first initialized to the first model parameters of the last transfer learning, and then the laboratory-made digital display screen character data set is selected for training to obtain the second model parameters of the pre-trained model; at the same time, the pre-trained model learns the characteristics of the digital display screen character data to improve the generalization ability of the model.
[0077] Specifically, use W in step S2 1 Initialize the pre-trained model weights and freeze the parameters W of BD in the PP-OCRv3 network 1 BD , set the self-made dataset as the target domain for transfer learning, and obtain the model parameters W after transfer 2 =(W 1 BD , W 2 SVTR ).
[0078] In this embodiment, in step S4, the pre-trained model is substituted into the second model parameters, and the pre-trained model is iteratively trained and learned through part of the real data and the simulated data until the loss function satisfies the set loss value, the training and learning are completed, and the trained pre-trained model is obtained;
[0079] Specifically, in iterative training and learning, the pre-trained model parameters are first initialized to the second model parameters saved by the second transfer learning, and then a small amount of real data and simulated data are used for training to obtain the trained pre-trained model, so that the pre-trained model learns the character features of the digital display screen under special environments.
[0080] Specifically, freeze the parameter W of the BD part in the PP-OCRv3 network 1 BD , and use the parameter W in step S32 SVTR Initialize the network parameters of the recognition module, use a small amount of real data and simulated data to perform transfer learning on the target domain, and obtain the model parameters W after transfer 3 =(W 1 BD , W 3 SVTR ).
[0081] In this embodiment, W i BD and W i SVTR (i=0, 1, 2, 3) represent the parameters of DBNet and SVTR_LCNet of the PP-OCRv3 pre-trained model. After multiple migrations, the pre-trained model is easier to obtain the character features of the coal mine digital display screen, and uses less training time and storage cost to solve the problem of poor generalization ability caused by insufficient model training samples. During the training process, the loss function decreases as shown in 4. Figure 4 It can be seen that the model of transfer learning decreases quickly, and after about 10 rounds, it converges stably and the loss value approaches 6. It can be seen that multiple transfer learning helps the model converge quickly and improves the accuracy of target recognition.
[0082] In this embodiment, the transfer learning strategy adopted can play a positive role in improving the performance of the training model and the prediction performance through the domain transfer of knowledge. On the one hand, the pre-trained model of a large data set can learn good general features, thereby enhancing its generalization ability. On the other hand, in the process of multiple migrations, the model gradually focuses on richer domain features, thereby improving its robustness.
[0083] Transfer learning includes three methods: model-based, feature-based, and relationship-based transfer learning. The present invention adopts model-based transfer learning, which can share the model and parameters of the source task and is applicable to migration tasks in various scenarios. Assume that the source task is S and the target task is T. The loss functions of the source task and the target task are LS(w) and LT(w'), respectively, where w and w' represent the model parameters of the source task and the target task, respectively. The goal of transfer learning is to find the parameter w so that LT(w') is small while keeping the performance of the source task from being significantly reduced. The model-based transfer learning method uses part or all of the source task model parameters w as the initial value of the target task model parameters w', and adjusts the parameters of the source task model by minimizing LT(w'), as shown in the following formula:
[0084] ;
[0085] Where w is the source task model parameter; w' is the model parameter after migration; L is the loss function of the model; D is the target task dataset; is the model; R(w) is the regularization; is the regularization coefficient.
[0086] In this embodiment, the loss function is constructed from the two aspects of character detection and recognition during the training process of PP-OCRv3 model transfer learning, and the model parameters are iteratively updated through the back propagation algorithm. The detection module uses DBNet, and the output of this network includes probability map, threshold map and approximate binary map. Figure 3 Obviously, the loss function of DBNet consists of three parts, each of which corresponds to the loss of a prediction graph, as shown in the following formula:
[0087] ;
[0088] Where, L total is the total loss function of the detection module; L p is the probability map loss; L b is the approximate binary image loss; L t is the threshold map loss; and All are weight coefficients;
[0089] The recognition module uses a three-stage highly progressive decreasing network structure SVTR_LCNet, which consists of Transformer and CTC decoding, and realizes character recognition through a simple parallel linear classifier. For a pair of input and output (X, Y), CTC decoding can remove duplicate labels and blank labels, select the character sequence with the highest score for output, and the loss calculation is shown in the following formula:
[0090] ;
[0091] In the formula, is the optimal character label sequence obtained by solving the model; X is the feature expression sequence containing character images, which are obtained by the detection module and corrected by the correction module; Y is the character label sequence obtained by solving; conditional probability It represents the probability that the character label is Y given that the input is X.
[0092] In the experiment, the transfer learning based on the PP-OCRv3 model is divided into two stages. The first stage is the frozen training stage, in which the parameters of the detection module DBNet in PP-OCRv3 are frozen, and only the parameters of the recognition module SVTR_LCNet network are trained, including character components, three-level cascade highly decreasing feature mixing, merging and splicing networks, and linear prediction layers. In the frozen training, the model weight initialization data comes from the results of the previous transfer learning, and the learning rate is set to 0.0002, so that the model converges quickly, so that it has a certain ability to recognize the number of characters on the digital display screen. The second stage is the thawing stage. When the model converges, it is the thawing stage. The previously frozen detection module DBNet parameters and the recognition SVTR_LCNet module parameters involved in the training are the parameters of the model after migration.
[0093] In this embodiment, the multiple migration PP-OCRv3 model is evaluated from two aspects: recognition accuracy and speed. The evaluation index of recognition accuracy is accuracy (ACC), and the calculation formula of the accuracy is:
[0094] ;
[0095] In the formula, The number of positive samples identified as positive samples; The positive sample recognition result is the number of positive samples; P+N is the total number of samples in the data set. It reflects the overall accuracy of the model character recognition. The higher the value, the better the system performance. The recognition speed evaluation index is the number of frames per second (FpS).
[0096] In this embodiment, in step S5, the target text image is input into the trained pre-trained model for recognition processing, and the recognition result is output.
[0097] Specifically, Figure 6 As shown, the characters on the digital display screen in the coal mine environment are input into the trained pre-trained model for recognition processing, such as Figure 5 As shown, the characters on the digital display screen are detected by a text detection module, and a text detection result is output; the text detection result is calibrated by a text calibration module, and a calibration effect is output; the calibration effect is recognized by a text recognition module, and a recognition result is output.
[0098] In a possible embodiment, a specific identification example is provided as follows:
[0099] The problem of character recognition on digital display screens in coal mines based on the PP-OCRv3 transfer learning model requires three data sets, namely, a public data set for text recognition, a laboratory-made digital display screen character data set, and a coal mine digital display screen character data set that includes a small amount of real data and simulated data. Among them, the public data set for text recognition includes 127K text detection, 18.5K text recognition, and 18.7K+500 text verification data, and the data set data has been processed and can be downloaded and used directly. The laboratory-made digital display screen character data set comes from the digital display screens of various equipment in actual production and the digital display screen images collected online, totaling 25,000 images. The coal mine digital display screen character data set has a total of 5,000 images, including real images obtained from actual projects and simulated coal mine complex environments such as low resolution, horizontal or vertical line interference, blur, and block interference images. The simulation methods include changing the contrast of the image, perspective transformation, adding Gaussian noise and salt and pepper noise, adding Gaussian blur, etc., to simulate the environment of the detected character image, as follows Figure 7 shown.
[0100] The character recognition model for coal mine digital display screen based on PP-OCRv3 transfer learning is a method based on deep learning. This method requires a computing power support platform to meet the needs of rapid processing and calculation of large amounts of data. To this end, all experiments were completed on the GPU hardware acceleration platform. The specific experimental platform configuration and model parameters used in the experiment are shown in Tables 1 and 2 respectively.
[0101] Table 1 Experimental platform configuration
[0102]
[0103] Table 2 Model parameter configuration
[0104]
[0105] In the experiment, the model training parameter settings fully considered the platform GPU performance issues. On the one hand, the batch_size can be adjusted to improve the impact of insufficient GPU memory on the recognition effect. The minimum batch_size is set to 2, and the learning rate value increases as its value increases. On the other hand, if the GPU memory allows, the batch_size is set to 8 during frozen training, the learning rate is 0.001 during frozen training, and the number of Epochs for frozen training is 100; the batch_size is 4 during unfrozen training, the learning rate is 0.0002 during unfrozen training, and the number of Epochs for unfrozen training is 100.
[0106] In order to verify the effectiveness and necessity of the PP-OCRv3 transfer learning model in character recognition on coal mine digital display screens, a comparative experiment between the original PP-OCRv3 and PP-OCRv3 transfer learning models was designed. The comparative experiment was run in the same experimental environment and used the same data set. The evaluation indicators of the experimental results are shown in Table 3, and the recognition effect is shown in Figure 8.
[0107] Table 3 Comparison of recognition results between PP-OCRv3 transfer learning and PP-OCRv3 model
[0108]
[0109] As can be seen from Table 3, the recognition accuracy of the PP-OCRv3 transfer learning model is higher than that of the comparison model in four different scenarios, achieving higher recognition accuracy. Specifically, in the scenes of light contrast, blur and interference blocks, the average recognition accuracy of the two models increased by 13.28%. In the scene with interference lines, the accuracy was as high as 77.53%, which was 29.32% higher than the model before improvement. On average, the accuracy increased by 17.29%. In terms of recognition speed, the FPS of the improved model in the four scenarios was better than that before improvement, and the average state increased by 27.295f / s. These results show that the PP-OCRv3 transfer learning model has been significantly improved in both recognition accuracy and recognition speed, and has achieved good results in character recognition accuracy and performance of coal mine digital display screens.
[0110] Randomly select coal mine digital instrument images from the data set for character recognition. The results are as follows: Figure 8 As shown. Among them, Figure 8 Parts (a) and (b) in the figure are the comparison of the recognition results of the PP-OCRv3 model and the PP-OCRv3 transfer learning model. It can be seen that the transfer learning model can achieve very ideal results regardless of blur or horizontal and vertical line interference scenes. Figure 8 The first figure in part (c) is the recognition result of the PP-OCRv3 model in an undisturbed image, followed by the recognition results of the PP-OCRv3 transfer learning model under the same image lighting conditions, blur, horizontal and vertical line interference, and noise interference. It can be seen that the PP-OCRv3 transfer learning model has better feature extraction capabilities and is more suitable for digital instrument character recognition in coal mine environments.
[0111] In summary, the present invention constructs a pre-trained model based on the PP-OCRv3 network; the pre-trained model is trained for the first time through a public data set for text recognition to obtain the first model parameters of the pre-trained model; the pre-trained model is substituted into the first model parameters, and the pre-trained model is trained for the second time through a laboratory-made digital display screen character data set to obtain the second model parameters of the pre-trained model; the pre-trained model is substituted into the second model parameters, and the pre-trained model is iteratively trained and learned through part of the real data and simulated data until the loss function meets the set loss value, and the training and learning are completed to obtain the trained pre-trained model; the target text image is input into the trained pre-trained model for recognition processing, and the recognition result is output. Based on deep learning and transfer learning, the present invention proposes a coal mine digital display screen character recognition method based on PP-OCRv3 transfer learning. The method uses a public data set to train the PP-OCRv3 network to construct a pre-trained model, and then based on the idea of transfer learning, the pre-trained model is fine-tuned multiple times in combination with a small amount of actual data and simulated data, effectively reducing the impact of the environment on the recognition performance and improving the recognition accuracy and speed. Aiming at the problems of poor recognition effect in complex environment and insufficient sample data in digital display screen character recognition in coal mine environment, the present invention improves the recognition effect by introducing the PP-OCRv3 model with leading recognition effect in the OCR field in order to improve the problem of unsatisfactory recognition effect caused by factors such as small recognition area, complex illumination changes, and poor image quality in complex coal mine environment; in order to solve the problem of insufficient data samples, the PP-OCRv3 model is trained by multiple transfer learning methods to effectively reduce the model learning cost and avoid model overfitting. The present invention has significantly improved the recognition accuracy and generalization ability of characters on digital display screens in coal mines, laying a solid theoretical and technical foundation for the research on coal mine safety production and intelligent construction.
[0112] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
[0113] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0114] Example 2
[0115] See also Fig. 9 Embodiment 2 of the present invention further provides a character recognition device based on PP-OCRv3 transfer learning, comprising:
[0116] A pre-training model building unit 001 is used to build a pre-training model based on the PP-OCRv3 network;
[0117] A pre-training model first training learning unit 002 is used to perform a first training learning on the pre-training model using a text recognition public data set to obtain a first model parameter of the pre-training model;
[0118] The pre-trained model second training learning unit 003 is used to substitute the pre-trained model into the first model parameters, perform a second training learning on the pre-trained model through a laboratory-made digital display screen character data set, and obtain the second model parameters of the pre-trained model;
[0119] The pre-trained model iterative training learning unit 004 is used to substitute the pre-trained model into the second model parameters, and iteratively train and learn the pre-trained model through part of the real data and the simulated data until the loss function meets the set loss value, and the training learning is completed to obtain the trained pre-trained model;
[0120] The text image recognition processing unit 005 is used to input the target text image into the trained pre-trained model for recognition processing and output the recognition result.
[0121] In this embodiment, in the pre-trained model first training learning unit 002, the pre-trained model second training learning unit 003 and the pre-trained model iterative training learning unit 004, during the training and learning process of the pre-trained model, the pre-trained model is trained and learned by a model-based transfer learning strategy; the model-based transfer learning strategy expression is:
[0122] ;
[0123] Where w is the source task model parameter; w' is the model parameter after migration; L is the loss function of the model; D is the target task dataset; is the model; R(w) is the regularization; is the regularization coefficient.
[0124] In this embodiment, in the pre-training model first training learning unit 002, the pre-training model second training learning unit 003 and the pre-training model iterative training learning unit 004, during the training and learning process of the pre-training model, a detection module loss function and a recognition module loss function are respectively constructed based on character detection and character recognition, and the parameters of the pre-training model are iteratively updated through a back propagation algorithm;
[0125] The expression of the detection module loss function is:
[0126] ;
[0127] Where, L total is the total loss function of the detection module; L p is the probability map loss; L b is the approximate binary image loss; L t is the threshold map loss; and All are weight coefficients;
[0128] The expression of the recognition module loss function is:
[0129] ;
[0130] In the formula, is the optimal character label sequence obtained by solving the model; X is the feature expression sequence containing character images, which are obtained by the detection module and corrected by the correction module; Y is the character label sequence obtained by solving; conditional probability It represents the probability that the character label is Y given that the input is X.
[0131] In this embodiment, in the pre-training model iterative training learning unit 004, the performance of the trained pre-training model is evaluated by two indicators: accuracy and number of frames transmitted per second; the calculation formula of the accuracy is:
[0132] ;
[0133] In the formula, The number of positive samples identified as positive samples; The positive sample recognition result is the number of positive samples; P+N is the total number of samples in the data set.
[0134] In this embodiment, in the text image recognition processing unit 005, when the target text image is input into the trained pre-trained model for recognition processing, the target text image is detected by a text detection module, and a text detection result is output; the text detection result is calibrated by a text calibration module, and a calibration effect is output; the calibration effect is recognized by a text recognition module, and a recognition result is output.
[0135] It should be noted that the information interaction, execution process and other contents between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.
[0136] Example 3
[0137] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code of a character recognition method based on PP-OCRv3 transfer learning is stored, and the program code includes instructions for executing embodiment 1 or any possible implementation thereof.
[0138] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0139] Example 4
[0140] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0141] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute a character recognition method based on PP-OCRv3 transfer learning of Example 1 or any possible implementation thereof.
[0142] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor implemented by reading software codes stored in a memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0143] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.
[0144] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing system, they can be concentrated on a single computing system, or distributed on a network composed of multiple computing systems, and optionally, they can be implemented by a program code executable by a computing system, so that they can be stored in a storage system and executed by the computing system, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0145] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. A method for character recognition on coal mine digital display screen based on PP-OCRv3 transfer learning, characterized in that: include: Based on the PP-OCRv3 network, a pre-trained model is built; Performing a first training study on the pre-trained model using a public text recognition data set to obtain a first model parameter of the pre-trained model; Substituting the pre-trained model into the first model parameters, performing a second training study on the pre-trained model using a laboratory-made digital display screen character data set, and obtaining second model parameters of the pre-trained model; Substituting the pre-trained model into the second model parameters, iteratively training the pre-trained model through part of the real data and the simulated data until the loss function satisfies the set loss value, completing the training and learning, and obtaining the trained pre-trained model; Input the target text image into the trained pre-trained model for recognition processing, and output the recognition result; Pre-trained model PP-OCRv3 initialization weight W0=(W0 BD , W0 SVTR ), the public text recognition dataset is the source task training sample, and the model converges to the parameter W 1= (W1 BD , W1 SVTR ); Initialize the pre-trained model weights with W1 and freeze the parameters W1 of BD in the PP-OCRv3 network BD , set the self-made dataset as the target domain for transfer learning, and obtain the model parameters after migration W2 = (W1 BD , W2 SVTR ); Freeze the parameter W1 of the BD part in the PP-OCRv3 network BD , and use parameter W2 SVTR Initialize the recognition module network parameters, use the data set composed of real data and simulated data to perform transfer learning again for the target domain, and obtain the model parameters after transfer W3 = (W1 BD , W3 SVTR ); During the training and learning process of the pre-trained model, a detection module loss function and a recognition module loss function are respectively constructed based on character detection and character recognition, and the parameters of the pre-trained model are iteratively updated through a back propagation algorithm; The expression of the detection module loss function is: L total =L p +αL b +βL t Where, L total is the total loss function of the detection module; L p is the probability map loss; L b is the approximate binary image loss; L t is the threshold map loss; α and β are weight coefficients; The expression of the recognition module loss function is: Where Y * is the optimal character label sequence obtained by solving the model; X is the feature expression sequence containing character images, which are obtained by the detection module and corrected by the correction module; Y is the character label sequence obtained by solving; the conditional probability p(Y|X) represents the probability that the character label is Y under the condition that the input is X; The performance of the trained pre-trained model is evaluated by two indicators: accuracy and number of frames per second. The calculation formula of the accuracy is: Where, T P is the number of positive samples identified as positive samples; T N The positive sample recognition result is the number of positive samples; P+N is the total number of samples in the data set.
2. According to claim 1, a method for character recognition on a coal mine digital display screen based on PP-OCRv3 transfer learning is characterized in that: During the training and learning process of the pre-trained model, the pre-trained model is trained and learned through a model-based transfer learning strategy; the model-based transfer learning strategy expression is: Where w is the source task model parameter; w' is the model parameter after migration; L is the loss function of the model; D is the target task dataset; f(x,w) is the model; R(w) is the regularization; λ is the regularization coefficient.
3. The method for character recognition on a coal mine digital display screen based on PP-OCRv3 transfer learning according to claim 1, characterized in that: In the process of inputting the target text image into the trained pre-trained model for recognition processing, the target text image is detected by the text detection module and the text detection result is output; the text detection result is calibrated by the text calibration module and the calibration effect is output; the calibration effect is recognized by the text recognition module and the recognition result is output.
4. A coal mine digital display screen character recognition device based on PP-OCRv3 transfer learning, using a coal mine digital display screen character recognition method based on PP-OCRv3 transfer learning as described in any one of claims 1 to 3, characterized in that: include: A pre-training model building unit, used to build a pre-training model based on the PP-OCRv3 network; A pre-training model first training learning unit is used to perform a first training learning on the pre-training model through a text recognition public data set to obtain a first model parameter of the pre-training model; A second training learning unit for the pre-trained model is used to substitute the pre-trained model into the first model parameters, perform a second training learning on the pre-trained model through a laboratory-made digital display screen character data set, and obtain second model parameters of the pre-trained model; A pre-trained model iterative training learning unit, used to substitute the pre-trained model into the second model parameters, and iteratively train and learn the pre-trained model through part of the real data and the simulated data until the loss function satisfies the set loss value, and the training learning is completed to obtain the trained pre-trained model; The text image recognition processing unit is used to input the target text image into the trained pre-trained model for recognition processing and output the recognition result.
5. A coal mine digital display screen character recognition device based on PP-OCRv3 transfer learning according to claim 4, characterized in that: In the first training learning unit of the pre-training model, the second training learning unit of the pre-training model and the iterative training learning unit of the pre-training model, during the training and learning process of the pre-training model, the pre-training model is trained and learned by a model-based transfer learning strategy; the model-based transfer learning strategy expression is: Where w is the source task model parameter; w' is the model parameter after migration; L is the loss function of the model; D is the target task dataset; f(x,w) is the model; R(w) is the regularization; λ is the regularization coefficient.
6. A coal mine digital display screen character recognition device based on PP-OCRv3 transfer learning according to claim 5, characterized in that: In the first training learning unit of the pre-training model, the second training learning unit of the pre-training model and the iterative training learning unit of the pre-training model, during the training and learning process of the pre-training model, a detection module loss function and a recognition module loss function are respectively constructed based on character detection and character recognition, and the parameters of the pre-training model are iteratively updated through a back-propagation algorithm; The expression of the detection module loss function is: L total =L p +αL b +βL t Where, L total is the total loss function of the detection module; L p is the probability map loss; L b is the approximate binary image loss; L t is the threshold map loss; α and β are weight coefficients; The expression of the recognition module loss function is: Where Y * is the optimal character label sequence obtained by solving the model; X is the feature expression sequence containing character images, which are obtained by the detection module and corrected by the correction module; Y is the solved character label sequence; the conditional probability p(Y|X) represents the probability that the character label is Y under the condition that the input is X.
7. A coal mine digital display screen character recognition device based on PP-OCRv3 transfer learning according to claim 6, characterized in that: In the pre-training model iterative training learning unit, the performance of the trained pre-training model is evaluated by two indicators: accuracy and number of frames transmitted per second; the calculation formula of the accuracy is: Where, T P is the number of positive samples identified as positive samples; T N The positive sample recognition result is the number of positive samples; P+N is the total number of samples in the data set.
8. The coal mine digital display screen character recognition device based on PP-OCRv3 transfer learning according to claim 7, characterized in that: In the text image recognition processing unit, in the process of inputting the target text image into the trained pre-trained model for recognition processing, the target text image is detected by a text detection module, and a text detection result is output; the text detection result is calibrated by a text calibration module, and a calibration effect is output; the calibration effect is recognized by a text recognition module, and a recognition result is output.
Citation Information
Patent Citations
End-to-end method for text detection and recognition based on deep learning
CN116758552A