Text Detector Training Method and Text Detection Method for Scene Text Detection
By introducing unsupervised intermediate training stage UNITS in scene text detection, unsupervised training is used for unsupervised training, the domain difference problem between synthetic data and real data is solved, and the model's representation ability and detection performance are improved.
Patent Information
- Application Number
- CN202210492865.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-05-07
AI Technical Summary
The prior art has a domain difference problem between synthetic data and real data in scene text detection, resulting in limited performance of the model when fine-tuning the real data and failing to effectively utilize a large amount of unlabeled real data.
A text detector training method for scene text detection is proposed, including UNITS, an unsupervised intermediate training stage. This stage connects the pre-training stage and the fine-tuning stage, and uses unsupervised training to improve the representation ability of the pre-trained model, thereby improving the performance of the final fine-tuning.
Through UNITS in the unsupervised intermediate training stage, the model can better perceive text information in real data, provide an initialization closer to real data, and improve the performance and generalization capabilities of scene text detection.
Smart Images

Figure CN114913531B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer software, and particularly relates to a method for training a text detector for scene text detection and a text detection method. Background Art
[0002] Obtaining text in a scene image mainly includes two steps: text detection and text recognition. Among them, text detection is similar to general object detection, that is, to locate the position where the text appears in the image, while text recognition is to convert the text in the image into text content stored in a computer. As a pre-step of text recognition, text detection plays a crucial role in the whole text extraction process and directly affects the accuracy of text extraction. In recent years, methods based on deep learning have become the mainstream methods for scene text detection, and these methods require a large amount of data for training. Most methods use synthetic data SynthText that is relatively easy to obtain and whose annotations can be automatically generated to pre-train the model and then fine-tune it on the target dataset. The existing technical methods have the following defects:
[0003] 1. The shapes, colors, fonts, sizes, and directions of text instances vary greatly and blend with the background relatively naturally, which results in a large domain difference between synthetic data and real data. When fine-tuning on real data, the results obtained by directly initializing the model with the parameters pre-trained on synthetic data may not be optimal.
[0004] 2. Generally speaking, more rounds of pre-training will bring more performance improvement to fine-tuning. However, due to the data inconsistency problem between the pre-training stage and the fine-tuning stage, this improvement is limited.
[0005] 3. Most of the existing methods are dedicated to designing a better model and rarely consider the data inconsistency problem between the pre-training stage and the fine-tuning stage.
[0006] 4. There are a large number of unlabeled images and video data containing text in the real world, and these data have not been well utilized in previous methods. Summary of the Invention
[0007] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a method for training a text detector for scene text detection and a text detection method. The method includes an unsupervised intermediate training stage UNITS, which can connect the pre-training stage and the fine-tuning stage and does not introduce additional calculations and model parameters during the inference process. In addition, the method can utilize more unlabeled real data to improve the representation ability of the pre-trained model so as to improve the performance of the final fine-tuning.
[0008] The unsupervised intermediate training stage UNITS proposed by the present invention can connect the pre-training stage and the fine-tuning stage, introduce the information in the real data into the pre-trained model, and provide a more real-data-close initialization for fine-tuning.
[0009] Three unsupervised training strategies are explored in UNITS to improve the representation ability of the pre-trained model by using synthetic data and a large amount of unlabeled real data, thereby improving the performance of the final fine-tuning.
[0010] The technical solution of the present invention is as follows:
[0011] A method for training a text detector for scene text detection, the steps of which include:
[0012] 1) Pre-train the selected text detector using the training data set; wherein, the images in the training data set are scene images embedded with text information;
[0013] 2) Initialize each branch of the set model UNITS with the text detector parameters obtained by pre-training; wherein, the model structure of the branch corresponds to the text detector structure;
[0014] 3) According to the unsupervised training strategy set in UNITS, use the unlabeled real data to perform unsupervised training on UNITS initialized in step 2), and at the same time use the training data set to perform supervised training on UNITS to update the model parameters of UNITS;
[0015] 4) Initialize the text detector with the model parameters of UNITS finally obtained in step 3), and then use the labeled target data set to perform supervised training on the initialized text detector to obtain the finally trained text detector.
[0016] Further, in step 3), the unsupervised training strategy is a double-branch single-supervision strategy, and the model UNITS includes two branches with the same structure but different parameters; wherein, the unlabeled scene image X u is input into the first branch to obtain the corresponding segmentation result P1; the image X' u obtained by data augmentation of X u is input into the second branch to obtain the corresponding segmentation result P2; then the segmentation result P1 is binarized according to the set threshold to obtain the pseudo-label, and the pseudo-label Y1 corresponding to the segmentation result P2 is obtained according to the inverse transformation corresponding to the data augmentation; then use the pseudo-label Y1 to supervise the segmentation result P2; then in step 4), the text detector is initialized with the parameters of the second branch in the finally obtained UNITS.
[0017] Further, in step 3), the unsupervised training strategy is a dual-branch dual-supervision strategy, and the model UNITS includes two branches with the same structure but different parameters;
[0018] 31) Input the unlabeled scene image X u into the first branch to obtain the corresponding segmentation result P1; input the image X' u obtained by performing data augmentation on X u into the second branch to obtain the corresponding segmentation result P2; then binarize the segmentation result P1 according to a set threshold to obtain pseudo-labels, and obtain the pseudo-labels Y1 corresponding to the segmentation result P2 according to the inverse transformation corresponding to the data augmentation; then use the pseudo-labels Y1 to supervise the segmentation result P2;
[0019] 32) Input X' u into the first branch trained in step 31) to obtain the corresponding segmentation result P1'; input X u into the second branch trained in step 31) to obtain the corresponding segmentation result P2'; then binarize the segmentation result P2' according to a set threshold to obtain pseudo-labels, and obtain the pseudo-labels Y1' corresponding to the segmentation result P1' according to the inverse transformation corresponding to the data augmentation; then use the pseudo-labels Y1' to supervise the segmentation result P1';
[0020] 33) Initialize the text detector using the parameters of the first branch or the second branch in the finally obtained UNITS.
[0021] Further, in step 3), the unsupervised training strategy is a single-branch single-supervision strategy, and the model UNITS includes one branch; perform data augmentation on the unlabeled scene image X u to obtain the image X' u , input X u , X' u into the branch respectively to obtain the corresponding segmentation results P1, P2, and then use the segmentation result P1 corresponding to X u to construct pseudo-labels for supervising the prediction result of X' u ; then in step 4), initialize the text detector using the parameters of the branch.
[0022] Further, the loss function for training the UNITS is where is the standard detection loss for the scene images in the training dataset, and is the supervision loss for the unlabeled real scene images.
[0023] Further, the training dataset is the synthetic dataset SynthText.
[0024] A method for text detection in a scene image, the steps of which include: inputting a scene image I to be detected into a text detector trained by the method described in claim 1 to obtain the positions of the text in the scene image I.
[0025] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above method.
[0026] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are realized when the computer program is executed by a processor.
[0027] The advantages of the present invention are as follows:
[0028] The present invention proposes a new training paradigm for scene text detection, which includes an unsupervised intermediate training stage UNITS. This training stage can connect the pre-training stage and the fine-tuning stage, and does not introduce additional calculations and model parameters during the inference process. In the unsupervised intermediate training stage UNITS, the present invention explores a series of unsupervised training strategies that only use synthetic data and unannotated real data, enabling the pre-trained model to perceive text information from real data and providing a more suitable model initialization for the subsequent fine-tuning stage. In addition, the training paradigm proposed by the present invention is not limited to a specific scene text detection model. The present invention has conducted extensive experiments on several classic scene text detection methods such as EAST, PSENet, PAN, and DB, and the experimental results on different datasets verify the effectiveness and generalization of the proposed method. Description of the Drawings
[0029] Figure 1 It is a schematic diagram of the model for the unsupervised intermediate training method for scene text detection.
[0030] Figure 2 It is a comparison diagram between the traditional training paradigm and the training paradigm proposed by the present invention in scene text detection;
[0031] (a) Traditional training paradigm, (b) Training paradigm proposed by the present invention.
[0032] Figure 3 It is a schematic diagram of different training strategies in the unsupervised intermediate training stage;
[0033] (a) Dual-branch single supervision, (b) Dual-branch double supervision, (c) Single-branch single supervision. Detailed Implementation Manner
[0034] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0035] After pre-training the model using a synthetic dataset, the model has acquired certain detection capabilities. Then, it is natural to think of semi-supervised methods. These methods classify the dataset into two parts: labeled data and unlabeled data. The goal is to obtain a stronger model for tasks such as classification, detection, and segmentation by utilizing some labeled data and a large amount of unlabeled data. Can we use some unlabeled real datasets during pre-training to enhance the capabilities of the pre-trained model? Based on this idea, the present invention proposes a method for improving the detection performance on a real target dataset using data from an unlabeled dataset. This method can achieve better initialization of the model in this unsupervised mode and achieve better fine-tuning performance. The overall idea of the method is as Figure 1 shown. First, pre-train the model using a large-scale synthetic data, and then perform unsupervised feature representation learning using unlabeled real data, that is, enable the model to perceive the text information in the real data while maintaining the text detection ability. Finally, fine-tune the target real dataset on this basis to improve the final detection performance.
[0036] Compared with the traditional "pre-training - fine-tuning" training paradigm, the present invention additionally adds an unsupervised intermediate training stage to obtain a new training paradigm UNITS. The comparison between the two training paradigms is as Figure 2 shown. Inspired by semi-supervised learning, the present invention explores three training strategies (two-branch single-supervised, two-branch double-supervised, single-branch single-supervised) in UNITS to enable the model to obtain an initialization closer to the real data when fine-tuning on the target dataset, as Figure 3 shown. Below, taking DB as the detection model, the three training strategies will be introduced.
[0037] The two-branch single-supervised consists of two branches with the same structure but different parameters (the branch is the DB detection model). One branch inputs the unlabeled image X u to obtain the corresponding segmentation result P1, and the other branch inputs the data augmentation version X′ u of X u (such as rotation data augmentation, color change augmentation, etc.) to obtain the corresponding segmentation result P2. First, set a threshold to binarize the segmentation result P1 to obtain pseudo-labels, and then obtain the pseudo-labels Y1 corresponding to P2 according to the inverse transformation corresponding to different data augmentations, and use it to supervise P2. It can be seen that for the unlabeled data, only the branch f(θ 1 ) has its parameters updated. In the fine-tuning stage, use the branch f(θ 2) Initialize the model and train for specific rounds for different scene text detectors.
[0038] Double-branch double-supervision is a modification of double-branch single-supervision. In double-branch single-supervision, it is expected that in the branches, f(θ 2 ) can learn better. Then, can the two branches promote each other to obtain better f(θ 1 ) and f(θ 2 ). On the basis of double-branch single-supervision, swap X u and X′ u , and also update the parameters of the branch f(θ 1 ). Since the two branches are completely symmetric during training, in the fine-tuning stage, arbitrarily select the branch f(θ 1 ) or f(θ 2 ) to initialize the model.
[0039] In the implementation of single-branch single-supervision, for the input unannotated image X u , perform data augmentation to obtain the image X′ u , input both of them into the model f(θ) to obtain the segmentation result, and then use the segmentation result of X u to construct pseudo-labels for supervising the prediction result of X′ u . If the parameters of f(θ 1 ) and f(θ 2 ) in double-branch single-supervision are shared, the implementation of single-branch single-supervision can be obtained. In the fine-tuning stage, use the branch f(θ) to initialize the model.
[0040] The parameters pre-trained on the synthetic dataset are used to initialize each of the above branches. To maintain the detection ability of each branch, use the annotated synthetic images to supervise each branch. In this way, the model can not only maintain a certain detection ability for the synthetic data but also perceive information from the real data. Finally, use the obtained weights to initialize the model and fine-tune it on the target dataset. Thus, the loss function of UNITS is:
[0041]
[0042] where is the standard detection loss for the annotated synthetic images (the images in the synthetic dataset SynthText) in each text detection method, is the supervision loss of the unannotated real images in UNITS.
[0043] The entire process of the present invention is divided into the following steps:
[0044] 1. Use the synthetic dataset SynthText to pre-train the detection model.
[0045] 2. Since each branch in UNITS is exactly the same as the detection model, the parameter of the detection model obtained by pre-training can be directly used to initialize each branch in UNITS.
[0046] 3. Then, according to different unsupervised training strategies in UNITS, the unlabeled real data X u and X' u are used for unsupervised training. At the same time, the labeled data in the synthetic dataset is used for supervised training to enable the model to maintain the text detection ability, and the model parameters are further updated.
[0047] 4. Finally, the model parameters obtained by training with the unsupervised training strategy are used for model initialization, and fine-tuning is performed on the final labeled target dataset, that is, the target dataset is used to retrain the detection model to obtain the final text detection model.
[0048] Effect analysis:
[0049] The present invention has been experimented on multiple classical scene text detection methods, including EAST, DB, PSENet, and PAN. At the same time, we have also conducted experiments on multiple datasets. Among them, SynthText is a large synthetic dataset, in which text images are embedded into 80,000 scene images in a natural simulation manner, with a total of 800,000 images, and there are about 8 million synthetic word instances in total. This dataset is used for pre-training and unsupervised training of the model. The TotalText dataset is a word-level labeled dataset, which contains horizontal text, multi-directional text, and curved text, with a total of 1,555 images, of which 1,255 images are used for training and the rest are used for testing. The ICDAR2015 dataset has a total of 1,500 images, of which two-thirds are used for training and one-third are used for testing. The MSRA-TD500 dataset is a multi-directional text dataset, which contains Chinese and English text instances with a large aspect ratio in natural scenes. Among them, 300 images are used for training and 200 images are used for testing. At the same time, an additional 400 images from HUST-TR400 are added as training images.
[0050] To verify the effectiveness of the three different training strategies in UNITS and their performance comparison, we conducted experiments on DB for verification and experiments on three datasets: ICDAR2015, TotalText, and MSRA-TD500. Specifically, first, train for 600 epochs in UNITS, then use the trained parameters for initialization, and finally fine-tune for 1,200 epochs on the real dataset. The experimental results are shown in Table 1. It is worth mentioning that the baseline method uses the branch f(θ in the double-branch single-supervised 1) Initialize the model to eliminate the influence of using synthetic data in UNITS. It can be seen that the three training strategies can all bring performance improvements on the three datasets, which verifies the effectiveness of the new training paradigm proposed in the present invention. The dual-branch single-supervision achieves the best performance among the three training strategies, so it is adopted as the default training strategy for subsequent experiments. In addition, after fine-tuning for 1800 epochs on the baseline method DB, it can be seen that the performance does not increase or even decreases due to overfitting, which confirms that the performance improvement does not come from the additional training epochs in UNITS.
[0051] Meanwhile, the present invention verifies the influence of data augmentation on UNITS through experiments. First, an experiment is designed to verify whether weak augmentation or strong augmentation should be used in UNITS. On the ICDAR2015 dataset, experiments are carried out using the methods DB and EAST, and the training strategy is dual-branch single-supervision. The experimental results are shown in Table 2. It can be seen that when using weak data augmentation (color jitter), the performance remains basically unchanged. Since the input images of the two branches are basically the same and the outputs are also almost the same, the provided supervision is roughly equivalent to the supervision of pre-training using synthetic data alone. When using strong data augmentation (random rotation), it can bring a 1.2% and 1.4% performance improvement on EAST and DB respectively. Then the present invention explores the roles of various data augmentations. On the basis of random rotation, two additional data augmentations are added: random cropping and random scaling. The experimental results on DB are shown in Table 3. On MSRA-TD500, both can bring significant performance improvements. On ICDAR2015 and TotalText, the performance of single data augmentation is slightly better than that of multiple data augmentations. Therefore, random rotation is used as the default data augmentation setting subsequently.
[0052] To verify whether the method proposed in the present invention can alleviate the domain difference problem, we directly use the pre-trained model and the UNITS model to conduct tests on the test set. The performance of the UNITS model is always better than that of the pre-trained model. For example, under the dual-branch single-supervised training strategy, on MSRA-TD500 (39.9% → 44.6%), ICDAR2015 (49.5% → 52.6%), and TotalText (49.8% → 51.5%). The domain adaptation method aims to achieve the highest possible performance in the target domain. In the dual-branch single-supervised implementation, labeled real data is used to approximate the best performance of the domain adaptation method, reaching an F-measure of 83.9% on MSRA-TD500. After fine-tuning, its F-measure increases to 86.5%, 0.5% higher than the baseline, and this slight performance improvement may come from strong data augmentation. In contrast, the F-measure of the dual-branch single-supervised after fine-tuning increases from 44.6% to 88.1%. This shows that better performance before fine-tuning does not necessarily mean better performance after fine-tuning. The domain adaptation method migrates the model to the target domain as much as possible, resulting in the loss of knowledge learned from pre-training on large-scale synthetic data. UNITS is a trade-off between the two and can serve as a good transition between the pre-training and fine-tuning stages.
[0053] As Figure 1 shown, the present invention explores the possibility of using other unlabeled data. We constructed a large dataset for UNITS training using existing publicly available training datasets (ART, MLT2017, MSRA-TD500, HUST-TR400, and ICDAR2015), named it UnlabeldDataset. By training UNITS on UnlabeldDataset, an initialized model can be obtained and directly used to fine-tune all target datasets, which can reduce the time for re-training UNITS for each dataset. Under the dual-branch single-supervised training strategy, the experimental results on DB are shown in Table 4. It can be seen that more data can bring similar performance improvements for ICDAR2015 and TotalText, but the improvement on MSRA-TD500 is small. This may be because MSRA-TD500 mainly consists of text instances with a large aspect ratio and the proportion of such data in UnlabeldDataset is small.
[0054] To further verify the effectiveness of the method proposed in the present invention, experiments were conducted on different scene text detection methods. Specifically, a dual-branch single-supervised training strategy was used in UNITS, and the experimental dataset was ICDAR2015. After adding UNITS, the F-measure on four scene text detection methods, namely DB, EAST, PSENet, and PAN, was improved by 1.3%, 1.2%, 0.9%, and 0.5% respectively, indicating that the method proposed in the present invention has good generalization ability on different detectors.
[0055] In addition, the method proposed in the present invention was also compared with the state-of-the-art methods, and the results are shown in Table 1. The method proposed in the present invention can achieve better performance than the previous methods on MSRA-TD500 and TotalText. In particular, the F value on MSRA-TD500 reached 88.1%, significantly higher than other methods. For ICDAR2015, the method proposed in the present invention is lower than CRAFT that uses additional character information.
[0056] Table 1 shows the detection results of different training strategies in UNITS and the comparison with previous methods
[0057]
[0058] Among them, "*" indicates the reproduced results.
[0059] Table 2 shows the ablation experiment of data augmentation intensity in UNITS on the ICDAR2015 dataset
[0060]
[0061]
[0062] Table 3 shows the ablation experiment of multiple data augmentations
[0063]
[0064] Among them, "single data augmentation" means random rotation, and "multiple data augmentations" means random rotation, random cropping, and random scaling.
[0065] Table 4 shows the ablation experiment of using a single dataset and multiple datasets in UNITS
[0066]
[0067] Among them, the single dataset represents using only the target dataset, and the multiple datasets refer to using UnlabeldDataset.
[0068] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined by the claims.
Claims
1. A method for training a text detector for scene text detection, the steps of which include: 1) Pre-train the selected text detector using a training dataset; wherein, the images in the training dataset are scene images embedded with text information; 2) Initialize each branch of the set model UNITS with the text detector parameters obtained from pre-training; wherein, the model structure of the branch corresponds to the structure of the text detector; UNITS is an unsupervised intermediate training stage for connecting the pre-training stage and the fine-tuning stage; 3) According to the unsupervised training strategy set in UNITS, use unlabeled real data to perform unsupervised training on UNITS initialized in step 2), and at the same time use the training dataset to perform supervised training on UNITS to update the model parameters of UNITS; 4) Initialize the text detector with the model parameters of UNITS finally obtained in step 3), and then use the labeled target dataset to perform supervised training on the initialized text detector to obtain the finally trained text detector.
2. The method according to claim 1, wherein, In step 3), the unsupervised training strategy is a two-branch single-supervised strategy, and the model UNITS includes two branches with the same structure but different parameters; among them, the unlabeled scene image X u is input into the first branch to obtain the corresponding segmentation result P1; the image X u obtained by performing data augmentation on X ′ u is input into the second branch to obtain the corresponding segmentation result P2; then the segmentation result P1 is binarized according to a set threshold to obtain pseudo-labels, and the pseudo-labels Y1 corresponding to the segmentation result P2 are obtained according to the inverse transformation corresponding to the data augmentation; then the segmentation result P2 is supervised using the pseudo-labels Y1; then in step 4), the text detector is initialized using the parameters of the second branch in the finally obtained UNITS.
3. The method according to claim 1, wherein, In step 3), the unsupervised training strategy is a dual-branch dual-supervision strategy, and the model UNITS includes two branches with the same structure but different parameters; (31) Input the unannotated scene image X u into the first branch to obtain the corresponding segmentation result P1; input the image X u obtained by performing data augmentation on X ′ u into the second branch to obtain the corresponding segmentation result P2; then perform binarization processing on the segmentation result P1 according to the set threshold to obtain pseudo-labels, and obtain the pseudo-label Y1 corresponding to the segmentation result P2 according to the inverse transformation corresponding to the data augmentation; then use the pseudo-label Y1 to supervise the segmentation result P2; 32) Input X ′ u Input the trained first branch in step 31) to obtain the corresponding segmentation result P1'; Input X u Input the trained second branch in step 31) to obtain the corresponding segmentation result P2'; Then, perform binarization processing on the segmentation result P2' according to the set threshold to obtain pseudo-labels, and obtain the pseudo-label Y1' corresponding to the segmentation result P1' according to the inverse transformation corresponding to the data augmentation; Then, use the pseudo-label Y1' to supervise the segmentation result P1'. 33) Initialize the text detector with the parameters of the first branch or the second branch in the finally obtained UNITS.
4. The method according to claim 1, wherein, In step 3), the unsupervised training strategy is a single-branch single-supervised strategy, and the model UNITS includes one branch; the unannotated scene image X u is subjected to data augmentation to obtain the image X ′ u , and X u , X ′ u are respectively input into the branch to obtain the corresponding segmentation results P1 and P2, and then the pseudo-label is constructed using the segmentation result P1 corresponding to X u to supervise the prediction result of X ′ u ; then in step 4), the text detector is initialized using the parameters of the branch.
5. The method according to any one of claims 1 to 4, wherein, The loss function for training the UNITS is where is the standard detection loss for the scene images in the training dataset, is the supervision loss for the unannotated real scene images.
6. The method according to claim 1, wherein, The training dataset is the synthetic dataset SynthText.
7. A method for text detection in a scene image, the steps of which include: Input a scene image I to be detected into the text detector trained by the method according to claim 1 to obtain the positions of the text in the scene image I.
8. A server, wherein, includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 7.
9. A computer-readable storage medium, on which a computer program is stored, wherein, the steps of the method according to any one of claims 1 to 7 are implemented when the computer program is executed by a processor.