Text detection method and device, electronic equipment and storage medium
By inputting the target image into the pre-trained student model and performing feature map fusion processing, combining the head network to generate probability maps and threshold maps, the problems of complex existing text detection models and poor detection effects on curved text are solved, and efficient and accurate text detection is achieved.
Patent Information
- Application Number
- CN202510276229.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
AI Technical Summary
The existing text detection model is complex and has poor detection effect on curved text.
A text detection method is adopted, by inputting the target image into the pre-trained student model, obtaining the feature map group and fusion processing are performed, combining the head network to generate a probability map and a threshold map, and binarized calculations are performed to determine the text area. The student model is jointly trained with the teacher model, and an adversarial training method is adopted to improve detection accuracy.
Effective detection of texts of various shapes is realized, reducing the amount of model parameters and calculation time, and improving the performance and efficiency of text detection.
Smart Images

Figure CN120182984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a text detection method, apparatus, electronic device, and storage medium. Background Art
[0002] OCR (Optical Character Recognition) is one of the important directions in computer vision and has rich application scenarios. OCR itself is a specific vision task, which involves various technologies, mainly including two major tasks: text detection and text recognition. Among them, the task of text detection is to locate the text regions in the input image for subsequent OCR tasks.
[0003] Currently, text detection is usually performed using regression-based text detection algorithms. The regression-based text detection method draws on the idea of general object detection algorithms. It regards the text in the image as the object to be detected and the rest as the background, and realizes detection by predicting the position of the text box. Although this method has achieved good results in text detection, it is often difficult to obtain a smooth text bounding curve for curved text, and the model is relatively complex without performance advantages.
[0004] Therefore, there is an urgent need for a text detection method to solve the problems of the existing text detection model being complex and having poor detection effects on curved text. Summary of the Invention
[0005] The present invention provides a text detection method, apparatus, electronic device, and storage medium to solve the defects of the existing technology that the text detection model is complex and has poor detection effects on curved text.
[0006] The present invention provides a text detection method, including the following steps: Input a target image into a pre-trained student model to obtain a set of feature maps output by the student model. The set of feature maps includes feature maps of multiple different scales. The student model is jointly trained with a teacher model, and the number of parameters of the student model is less than that of the teacher model; Perform fusion processing on the feature maps of multiple different scales in the set of feature maps to obtain a fused feature map; Input the fused feature map into a head network to obtain a probability map and a threshold map output by the head network. The head network includes a first branch and a second branch. The first branch is used to generate the probability map, and the probability map reflects the probability that each pixel point in the target image belongs to the text region. The second branch is used to generate the threshold map, and the threshold map contains the threshold used by each pixel point in the target image during binarization; Based on the probability map and the threshold map, perform binarization calculation to obtain a binary map; Based on the binary map, determine the text region.
[0007] According to a text detection method provided by the present invention, the fusion processing of multiple feature maps with different scales in the feature map group to obtain a fused feature map includes: Input multiple feature maps with different scales in the feature map group into a feature pyramid network to obtain the fused feature map output by the feature pyramid network.
[0008] According to a text detection method provided by the present invention, the joint training of the student model and the teacher model is performed based on the following steps: Take the student model and the feature pyramid network as the student feature extractor, and take the teacher model and the feature pyramid network as the teacher feature extractor, and perform adversarial training on the student feature extractor, the teacher feature extractor, and the feature discriminator.
[0009] According to a text detection method provided by the present invention, the adversarial training of the student feature extractor, the teacher feature extractor, and the feature discriminator includes: After loading the pre-trained model parameters into the student feature extractor and the teacher feature extractor respectively, freeze the model parameters of the student feature extractor and the teacher feature extractor, and perform multiple rounds of training on the feature discriminator; Freeze the model parameters of the trained feature discriminator, thaw the model parameters of the student feature extractor and the teacher feature extractor, and simultaneously perform multiple rounds of training on the student feature extractor and the teacher feature extractor; Perform cyclic iterative training on the feature discriminator, the student feature extractor, and the teacher feature extractor until a preset iteration stop condition is reached.
[0010] According to a text detection method provided by the present invention, the text detection method further includes: When training the feature discriminator, label the fused feature map output by the teacher feature extractor as 1 and label the fused feature map output by the student feature extractor as 0; When training the student feature extractor and the teacher feature extractor, label the fused feature map output by the teacher feature extractor and the fused feature map output by the student feature extractor as 1.
[0011] According to a text detection method provided by the present invention, when performing adversarial training on the student feature extractor, the teacher feature extractor, and the feature discriminator, train based on a preset loss function; The preset loss function is determined based on the sum of the true label loss function and the generative adversarial loss function; the true label loss function is determined based on the probability map loss, the binary map loss, and the threshold map loss; the generative adversarial loss function is the cross-entropy loss function.
[0012] The present invention also provides a text detection device, including the following modules: A feature extraction module, configured to: input a target image into a pre-trained student model to obtain a group of feature maps output by the student model, the group of feature maps including feature maps of multiple different scales, the student model being jointly trained with a teacher model, and the number of parameters of the student model being less than that of the teacher model; A feature fusion module, configured to: perform fusion processing on the feature maps of multiple different scales in the group of feature maps to obtain a fused feature map; A probability calculation module, configured to: input the fused feature map into a head network to obtain a probability map and a threshold map output by the head network, the probability map reflecting the probability that each pixel point in the target image belongs to a text region, and the threshold map including the threshold used by each pixel point in the target image during binarization; A binary calculation module, configured to: perform binarization calculation based on the probability map and the threshold map to obtain a binary map; A text region prediction module, configured to: determine a text region based on the binary map.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the text detection method described in any one of the above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the text detection method described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the text detection method described in any one of the above is implemented.
[0016] The text detection method, device, electronic device, and storage medium provided by the present invention input a target image into a pre-trained student model to obtain a set of feature maps output by the student model. The set of feature maps includes feature maps of multiple different scales. The student model is jointly trained with a teacher model, and the number of parameters of the student model is less than that of the teacher model. The feature maps of multiple different scales in the set of feature maps are fused to obtain a fused feature map. The fused feature map is input into a head network to obtain a probability map and a threshold map output by the head network. The head network includes a first branch and a second branch. The first branch is used to generate the probability map, and the probability map reflects the probability that each pixel point in the target image belongs to the text region. The second branch is used to generate the threshold map, and the threshold map contains the threshold used by each pixel point in the target image during the binarization process. Based on the probability map and the threshold map, binarization calculation is performed to obtain a binary map. Based on the binary map, the text region is determined. In the present invention, the probability map of the text region is obtained through feature extraction and prediction, and then the text segmentation region is obtained, and texts of various shapes can be effectively detected. The teacher model and the student model are jointly trained, so that the detection accuracy of the student model approaches that of the teacher model, while the number of parameters of the student model is less than that of the teacher model, reducing the time consumption of text detection and improving the performance of text detection. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is a flowchart of the text detection method provided by the present invention; Figure 2 is a schematic diagram of the network architecture of DBNet in the prior art; Figure 3 is a flowchart of the adversarial training provided by the present invention; Figure 4 is a schematic diagram of the structure of the text detection device provided by the present invention.
[0019] Figure 5 is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Embodiments
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.
[0021] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. Unless otherwise clearly specified and limited, the terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0022] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0023] Figure 1 is a schematic flowchart of the text detection method provided by the present invention. As Figure 1 shown, the method includes the following: S110. Input the target image into a pre-trained student model to obtain a set of feature maps output by the student model. The set of feature maps includes feature maps of multiple different scales. The student model is jointly trained with a teacher model, and the number of parameters of the student model is less than that of the teacher model. S120. Perform fusion processing on the feature maps of multiple different scales in the set of feature maps to obtain a fused feature map. S130. Input the fused feature map into a head network to obtain a probability map and a threshold map output by the head network. The head network includes a first branch and a second branch. The first branch is used to generate the probability map, and the probability map reflects the probability that each pixel point in the target image belongs to the text region. The second branch is used to generate the threshold map, and the threshold map contains the threshold used by each pixel point in the target image during the binarization process. S140. Perform binarization calculation based on the probability map and the threshold map to obtain a binary map. S150. Determine the text region based on the binary map.
[0024] It should be noted that the execution subject of the text detection method provided in the embodiments of the present invention may be a server or a computer device, such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.
[0025] In the embodiments of the present invention, the network structure of the teacher model is relatively complex. The backbone network of the teacher model can use ResNet18_v or other larger CNN model structures. Since the backbone network can be initialized with a large-scale pre-trained model, it usually has better generalization ability and performance. However, due to its large number of model parameters, in practical applications, the resources and time consumed during model inference are more, and the cost of model training is also higher.
[0026] In the embodiments of the present invention, the student model is the model used for final text detection, and its network structure is simpler than that of the teacher model. Its backbone network can use lightweight CNN architectures such as MobileNetV3. The student model has a small number of parameters, requires less system resources, and has a faster inference verification speed, which is suitable for actual industrial applications.
[0027] In the embodiments of the present invention, the student model and the teacher model are jointly trained, so that the detection accuracy of the student model approaches that of the teacher model, thereby realizing high-precision text detection with a smaller model size.
[0028] In the embodiments of the present invention, the student model performs multi-scale feature extraction on the target image. For example, feature maps with sizes of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original size of the target image are extracted as a group of feature maps. The feature maps of multiple scales are fused, and finally a fused feature map of a preset scale is obtained. For example, a feature map with a size of 1 / 4 of the original image scale is finally obtained.
[0029] In S130, the fused feature map is decoded by a Head network to generate a probability map and a threshold map. Among them, the probability map is a two-dimensional matrix that is the same as or similar to the size of the input image. The value of each pixel represents the probability that the position belongs to the text region. These probability values usually range from 0 to 1. A value close to 1 means a higher confidence that the pixel belongs to the text. The threshold map is also a two-dimensional matrix that is the same as or similar to the size of the input image, but it records the threshold to be used for each pixel during the binarization process. This threshold determines the standard for marking the pixel as text or non-text.
[0030] In S140, for each pixel, a new value is calculated using the probability value and the threshold at its corresponding position as a part of the approximate binary map. Among them, the value of each pixel represents whether the pixel is a text region.
[0031] In S150, through post-processing steps such as contour detection, the text region in the binary map is converted into a rectangular box or a polygon to obtain the detection result. In the specific implementation process, since the detection result may be a polygon, it needs to be converted into a horizontal rectangle through affine transformation.
[0032] The text detection method provided by the embodiment of the present invention inputs a target image into a pre-trained student model to obtain a feature map group output by the student model. The feature map group includes feature maps of multiple different scales. The student model is jointly trained with a teacher model, and the number of parameters of the student model is less than that of the teacher model; the feature maps of multiple different scales in the feature map group are fused to obtain a fused feature map; the fused feature map is input into a head network to obtain a probability map and a threshold map output by the head network. The probability map reflects the probability that each pixel point in the target image belongs to a text region, and the threshold map contains the threshold used by each pixel point in the target image during binarization; based on the probability map and the threshold map, binarization calculation is performed to obtain a binary map; based on the binary map, the text region is determined. In the present invention, a probability map of the text region is obtained through feature extraction and prediction, and then the text segmentation region is obtained, and texts of various shapes can be effectively detected; the teacher model and the student model are jointly trained, so that the detection accuracy of the student model approaches that of the teacher model, and the number of parameters of the student model is less than that of the teacher model, reducing the time consumption of text detection and improving the performance of text detection.
[0033] In an optional embodiment, the fusing the feature maps of multiple different scales in the feature map group to obtain a fused feature map includes: Inputting the feature maps of multiple different scales in the feature map group into a feature pyramid network to obtain a fused feature map output by the feature pyramid network.
[0034] Figure 2 is a schematic diagram of the network architecture of DBNet in the prior art. As Figure 2 shown, the text detection network provided by the embodiment of the present invention is based on DBNet and is divided into the following three parts: (1) The backbone network is responsible for extracting multi-scale features of the image, that is, the above-mentioned student model / teacher model; (2) The Feature Pyramid Networks (FPN) constructs a feature pyramid to fuse feature maps of different scales. Through the bottom-up and top-down paths and combined with lateral connections, the fusion of multi-scale features is realized, thereby improving the detection ability of the model for different scales; (3) The head network is used to calculate the probability map of the text region.
[0035] The text detection method provided by the embodiment of the present invention fuses feature maps of different scales through a feature pyramid network to make full use of information of different scales and improve the subsequent detection accuracy.
[0036] In an alternative embodiment, the student model and the teacher model are jointly trained based on the following steps: Taking the student model and the feature pyramid network as the student feature extractor, and taking the teacher model and the feature pyramid network as the teacher feature extractor, adversarial training is performed on the student feature extractor, the teacher feature extractor, and the feature discriminator.
[0037] In the embodiment of the present invention, integrating the idea of adversarial training, the student feature extractor and the teacher feature extractor are used as generators to generate samples that are close to real but adversarial. The feature discriminator is responsible for distinguishing whether the input sample is a real sample or an adversarial sample generated by the generator; by alternately iteratively training the generator and the feature discriminator, the generator and the discriminator are jointly optimized, thereby improving the robustness of the entire model and enabling the model to maintain stable performance when facing adversarial samples.
[0038] It can be understood that when training the generator, the student feature extractor and the teacher feature extractor are trained simultaneously to align the feature space of the student model with the feature space of the teacher model to obtain better performance.
[0039] The text detection method provided by the embodiment of the present invention integrates the idea of adversarial training, and optimizes the feature space of the student model through continuous alternating training, so that the model maintains stable performance when facing adversarial samples and improves the robustness of the entire model.
[0040] In an alternative embodiment, the adversarial training of the student feature extractor, the teacher feature extractor, and the feature discriminator includes: After loading the pre-trained model parameters into the student feature extractor and the teacher feature extractor respectively, freezing the model parameters of the student feature extractor and the teacher feature extractor, and performing multiple rounds of training on the feature discriminator; Freezing the model parameters of the trained feature discriminator, thawing the model parameters of the student feature extractor and the teacher feature extractor, and simultaneously performing multiple rounds of training on the student feature extractor and the teacher feature extractor; Performing cyclic iterative training on the feature discriminator, the student feature extractor, and the teacher feature extractor until a preset iteration stop condition is reached.
[0041] Different from using noise as the input of the generator in a general Generative Adversarial Network (GAN), in the embodiments of the present invention, the input is an image with text. The input is subjected to feature extraction through a backbone network (student model / teacher model) + FPN network, and the backbone network + FPN network is used as the generator. The finally extracted fused feature map is the output of the generator.
[0042] Figure 3 is a schematic diagram of the adversarial training provided by the present invention. As Figure 3 shown, in the embodiments of the present invention, the input of the feature discriminator is two kinds of fused feature maps output by the student feature extractor (i.e., the Student feature extractor in Figure 3 ) and the teacher feature extractor (i.e., the Teacher feature extractor in Figure 3 ), which are the Student feature map and the Teacher feature map respectively. The output of the feature discriminator is 0 or 1, representing "false" or "true" respectively. Essentially, the feature discriminator is a binary classifier used to classify the two different fused feature maps output by the student feature extractor and the teacher feature extractor.
[0043] Optionally, the feature discriminator includes several convolutional layers for compressing the fused feature map, and an output layer for outputting the classification probability.
[0044] In the embodiments of the present invention, first, the network parameters of the generator are frozen, and the feature discriminator is trained to obtain a discriminator with a relatively high accuracy. Then, the parameters of the feature discriminator are frozen, and the generator is trained. After one round of alternating training, the accuracy of the feature discriminator cannot meet the requirement of distinguishing the fused feature maps output by the teacher feature extractor and the student feature extractor. Therefore, the network parameters of the generator are frozen again, and the feature discriminator is repeatedly trained to obtain a higher accuracy. When the accuracy of the feature discriminator is improved to be high enough, the training process of the generator is repeated. By continuously repeating such training in a cycle, the discriminator and the generator co-evolve, and the performance of the overall network can be continuously improved. As the above iterative training progresses, the overall loss of the network gradually converges, and the performance of the discriminator network also continuously improves. When the discriminator can no longer effectively distinguish the fused feature maps output by the teacher feature extractor and the student feature extractor after hundreds of epochs (the process in which the entire training dataset is completely traversed by the model once), it means that the feature extractor already has good enough performance, and the two feature spaces output by it are aligned. At this time, the training can be stopped. The finally trained student model not only has a high accuracy on the binary probability map, but also is closer to the teacher model in the feature space and has better performance.
[0045] The text detection method provided by the embodiments of the present invention alternately trains a student feature extractor, a teacher feature extractor, and a feature discriminator, enabling the generator and the feature discriminator to co-evolve and continuously improving the overall network performance. At the same time, the teacher feature extractor and the student feature extractor are trained to align the feature space of the student feature extractor with the feature space of the teacher feature extractor, narrowing the feature outputs of the two models, thereby enhancing the performance of the student model.
[0046] Further, the text detection method further includes: When training the feature discriminator, mark the fused feature map output by the teacher feature extractor as 1 and the fused feature map output by the student feature extractor as 0; When training the student feature extractor and the teacher feature extractor, mark the fused feature map output by the teacher feature extractor and the fused feature map output by the student feature extractor as 1.
[0047] In the embodiments of the present invention, when training the feature discriminator, first load the pre-trained model parameters. Before starting the training, freeze the network parameters of the generator, and then input an image into the network to obtain two different fused feature maps through the student feature extractor and the teacher feature extractor respectively. Since the performance and pre-trained parameters of the teacher feature extractor are better, in the first few rounds of training, it will output a feature map with higher accuracy. Therefore, when training the feature discriminator, mark the fused feature map output by the teacher feature extractor as 1 and the fused feature map output by the student feature extractor as 0. Based on this, train the feature discriminator. After training for N epochs (the process in which the entire training dataset is completely traversed by the model once), a feature discriminator with higher accuracy can be obtained. Using this feature discriminator, it is easy to distinguish whether the fused feature map comes from the teacher feature extractor or the student feature extractor.
[0048] It can be understood that since the feature differences between the two fused feature maps are relatively large in the initial stage of training, a feature discriminator with relatively high accuracy can be obtained after training for a relatively small number of rounds.
[0049] In the embodiments of the present invention, after the feature discriminator is trained for N epochs, the parameters of the feature discriminator are frozen, and at the same time, the model parameters of the teacher feature extractor and the student feature extractor are unfrozen. In order to achieve the purpose of adversarial training, at this time, the labels of the feature maps are reversed, that is, the labels of the fused feature maps output by the teacher feature extractor and the student feature extractor are both recorded as 1. At this time, the entire network structures of the teacher feature extractor and the student feature extractor can be trained simultaneously. After training for N epochs, a pair of feature generators with better performance can be obtained. Since the labels of the feature maps of the two networks are both recorded as 1 here, that is, the same classification, even the feature discriminator with higher accuracy in the previous step will be difficult to distinguish the source of the feature maps after training for several rounds. In this way, the feature spaces of the student feature extractor and the teacher feature extractor are made as close as possible. When the two feature maps are similar enough to "deceive" the discriminator, a pair of generators with better performance is obtained.
[0050] In an alternative embodiment, when performing adversarial training on the student feature extractor, the teacher feature extractor, and the feature discriminator, training is performed based on a preset loss function; The preset loss function is determined based on the sum of the true label loss function and the generative adversarial loss function; the true label loss function is determined based on the probability map loss, the binary map loss, and the threshold map loss; the generative adversarial loss function is the cross-entropy loss function.
[0051] In the embodiments of the present invention, different loss functions are used to train the probability map, the threshold map, and the binary map, as shown in Table 1 below.
[0052] Table 1 Loss Function Table Here, the calculation formula of the final true label loss function GT Loss is as follows: ; Wherein, gt represents the probability map and threshold map labels, S out represents the output of the final head network corresponding to the student model, L p is the probability map cross-entropy loss, L b is the binary map Dice loss, L t is the threshold map L1 loss, and α and β are hyperparameters.
[0053] In the embodiments of the present invention, the generative adversarial loss function is as follows: ; Wherein,x Represents the input image sample, X is the image training dataset, T represents the teacher feature extractor, S represents the student feature extractor, D represents the feature discriminator.
[0054] Here, part of the purpose of optimizing the loss is to improve the discriminator D 's resolution ability, that is, when receiving S the output features will give a low score, while T 's output gives a high score, that is to say, during training D we want log D ( T ( x )) to be as high as possible, while D ( S ( x )) to be as low as possible, that is, log(1 - D ( S ( x ))) will also be high. Thus, the process of training D is the process of making the Loss reach the maximum value. Another part of the purpose of optimizing the loss is to make the feature extractor S generate feature maps that are realistic enough to pass off as real, that is, to make D ( S ( x )) produce a high score, that is, log(1 - D ( S ( x ))) becomes low, so the process of training the feature extractor S 、 T is the process of making the Loss reach the minimum value. In the embodiments of the present invention, D and S are alternately trained, T which can be summarized as the above min max loss function. Since the GAN Loss itself is calculated in the feature space, this loss is actually the Feature Loss of the overall network.
[0055] Finally, add the above real label loss GT Loss and the generative adversarial loss GAN Loss to obtain the total loss function: .
[0056] Optionally, after the training is completed according to the above steps, connected regions, that is, shrunk text regions, are obtained from the resulting binary map according to the post-processing steps of the DB algorithm. Then, the shrunk text regions are expanded according to the offset coefficient D' of the Vatti clipping algorithm to obtain the final text box. The calculation formula of D' is as follows: ; Among them, A' and L' are the area and perimeter of the shrinking region, and r' is set empirically. For example, it can be set to 1.5 (corresponding to the shrinking ratio r = 0.4).
[0057] Optionally, in order to evaluate the performance metrics of the training model of the text detection method, three metrics are used: Precision, Recall, and Hmean: Traverse all detected text boxes, calculate the intersection over union (IOU) between each of them and the labeled box respectively. If it is greater than a preset threshold (such as 0.5), it is recorded as a correctly detected text box. Count the number of all correctly detected boxes, and the ratio of the number of correctly detected text boxes to the number of predicted detected text boxes is recorded as the precision Precision; The ratio of the number of correctly detected text boxes to the number of labeled text boxes is the recall Recall; The calculation method of the harmonic mean metric is the same as that of the F1-score, and the formula is as follows: .
[0058] In summary, the text detection method provided by the present invention trains a student model and a teacher model. Only two models need to be trained to obtain a model with better performance, reducing the network training burden; by introducing the adversarial training loss into the overall loss, the teacher model can provide rich feature space information for the student model, enabling an effective optimization alignment method for the feature space of the visual recognition text detection model, narrowing the feature output of the two models, thereby improving the performance of the target network; integrating the adversarial generative network and the differentiable binarization algorithm network, enabling the efficient coupling of two different types of networks, without separate intermediate training products, and realizing end-to-end automation in the entire training process.
[0059] The text detection device provided by the embodiments of the present invention will be described below. The text detection device described below can be correspondingly referred to the text detection method described above.
[0060] Figure 4 is a schematic structural diagram of the text detection device provided by the present invention. As Figure 4 shown, the text detection device may include but is not limited to; A feature extraction module 410, configured to: input a target image into a pre-trained student model to obtain a group of feature maps output by the student model, the group of feature maps includes feature maps of multiple different scales, the student model is jointly trained with a teacher model, and the number of parameters of the student model is less than that of the teacher model; The feature fusion module 420 is configured to: perform fusion processing on multiple feature maps with different scales in the feature map group to obtain a fused feature map; The probability calculation module 430 is configured to: input the fused feature map into the head network to obtain a probability map and a threshold map output by the head network. The head network includes a first branch and a second branch. The first branch is used to generate the probability map, and the probability map reflects the probability that each pixel point in the target image belongs to the text region. The second branch is used to generate the threshold map, and the threshold map includes the threshold used by each pixel point in the target image during the binarization process; The binarization calculation module 440 is configured to: perform binarization calculation based on the probability map and the threshold map to obtain a binary map; The text region prediction module 450 is configured to: determine the text region based on the binary map.
[0061] In an optional embodiment, the feature fusion module 420 is specifically configured to: Input multiple feature maps with different scales in the feature map group into the feature pyramid network to obtain the fused feature map output by the feature pyramid network.
[0062] In an optional embodiment, the text detection device further includes a joint training module. The joint training module is configured to: use the student model and the feature pyramid network as the student feature extractor, and use the teacher model and the feature pyramid network as the teacher feature extractor to perform adversarial training on the student feature extractor, the teacher feature extractor, and the feature discriminator.
[0063] In an optional embodiment, the joint training module is specifically configured to: After loading the pre-trained model parameters into the student feature extractor and the teacher feature extractor respectively, freeze the model parameters of the student feature extractor and the teacher feature extractor, and perform multiple rounds of training on the feature discriminator; Freeze the model parameters of the trained feature discriminator, thaw the model parameters of the student feature extractor and the teacher feature extractor, and perform multiple rounds of training on the student feature extractor and the teacher feature extractor simultaneously; Perform cyclic iterative training on the feature discriminator, the student feature extractor, and the teacher feature extractor until a preset iteration stop condition is reached.
[0064] In an optional embodiment, the joint training module is further specifically configured to: When training the feature discriminator, label the fused feature map output by the teacher feature extractor as 1, and label the fused feature map output by the student feature extractor as 0; When training the student feature extractor and the teacher feature extractor, the fused feature maps output by the teacher feature extractor and the fused feature maps output by the student feature extractor are labeled as 1.
[0065] In an alternative embodiment, the joint training module is further specifically configured to: When performing adversarial training on the student feature extractor, the teacher feature extractor, and the feature discriminator, train based on a preset loss function; The preset loss function is determined based on the sum of a true label loss function and a generative adversarial loss function; the true label loss function is determined based on a probability map loss, a binary map loss, and a threshold map loss; the generative adversarial loss function is a cross-entropy loss function.
[0066] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute a text detection method, which includes: inputting a target image into a pre-trained student model to obtain a set of feature maps output by the student model, the set of feature maps including feature maps of multiple different scales, the student model being jointly trained with a teacher model, and the number of parameters of the student model being less than the number of parameters of the teacher model; Performing fusion processing on the feature maps of multiple different scales in the set of feature maps to obtain a fused feature map; Inputting the fused feature map into a head network to obtain a probability map and a threshold map output by the head network, the head network including a first branch and a second branch, the first branch being used to generate a probability map, the probability map reflecting the probability that each pixel point in the target image belongs to a text region, and the second branch being used to generate a threshold map, the threshold map including the threshold used by each pixel point in the target image during binarization; Performing binarization calculation based on the probability map and the threshold map to obtain a binary map; Determining a text region based on the binary map.
[0067] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0068] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the text detection method provided by the above-mentioned various methods. The method includes: inputting a target image into a pre-trained student model to obtain a set of feature maps output by the student model. The set of feature maps includes feature maps of multiple different scales. The student model is jointly trained with a teacher model, and the number of parameters of the student model is less than the number of parameters of the teacher model; Performing fusion processing on the feature maps of multiple different scales in the set of feature maps to obtain a fused feature map; Inputting the fused feature map into a head network to obtain a probability map and a threshold map output by the head network. The head network includes a first branch and a second branch. The first branch is used to generate the probability map, and the probability map reflects the probability that each pixel point in the target image belongs to the text region. The second branch is used to generate the threshold map, and the threshold map contains the threshold used by each pixel point in the target image during the binarization process; Based on the probability map and the threshold map, performing binarization calculation to obtain a binary map; Based on the binary map, determining the text region.
[0069] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the text detection method provided by the above-mentioned various methods. The method includes: inputting a target image into a pre-trained student model to obtain a set of feature maps output by the student model. The set of feature maps includes feature maps of multiple different scales. The student model is jointly trained with a teacher model, and the number of parameters of the student model is less than that of the teacher model; Performing fusion processing on the feature maps of multiple different scales in the set of feature maps to obtain a fused feature map; Inputting the fused feature map into a head network to obtain a probability map and a threshold map output by the head network. The head network includes a first branch and a second branch. The first branch is used to generate a probability map, and the probability map reflects the probability that each pixel point in the target image belongs to a text region. The second branch is used to generate a threshold map, and the threshold map contains the threshold used by each pixel point in the target image during binarization; Performing binarization calculation based on the probability map and the threshold map to obtain a binary map; Determining a text region based on the binary map.
[0070] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text detection method, characterized in that: include: Inputting a target image into a pre-trained student model to obtain a feature map group output by the student model, wherein the feature map group includes a plurality of feature maps of different scales, wherein the student model is obtained by joint training with a teacher model, and the number of parameters of the student model is smaller than the number of parameters of the teacher model; Fusing a plurality of feature maps of different scales in the feature map group to obtain a fused feature map; Input the fused feature map into a head network to obtain a probability map and a threshold map output by the head network, wherein the head network includes a first branch and a second branch, wherein the first branch is used to generate a probability map, wherein the probability map reflects the probability that each pixel in the target image belongs to a text area, and the second branch is used to generate a threshold map, wherein the threshold map includes a threshold used by each pixel in the target image during binarization; Based on the probability map and the threshold map, a binarization calculation is performed to obtain a binary map; Based on the binary image, a text area is determined.
2. The text detection method according to claim 1, characterized in that: The fusing the feature maps of multiple different scales in the feature map group to obtain a fused feature map includes: The feature maps of multiple different scales in the feature map group are input into a feature pyramid network to obtain a fused feature map output by the feature pyramid network.
3. The text detection method according to claim 2, characterized in that: The student model and the teacher model are jointly trained based on the following steps: The student model and the feature pyramid network are used as student feature extractors, the teacher model and the feature pyramid network are used as teacher feature extractors, and adversarial training is performed on the student feature extractor, the teacher feature extractor and the feature discriminator.
4. The text detection method according to claim 3, characterized in that: The adversarial training of the student feature extractor, the teacher feature extractor and the feature discriminator comprises: After loading the pre-trained model parameters into the student feature extractor and the teacher feature extractor respectively, freezing the model parameters of the student feature extractor and the teacher feature extractor, and performing multiple rounds of training on the feature discriminator; Freezing the model parameters of the trained feature discriminator, unfreezing the model parameters of the student feature extractor and the teacher feature extractor, and performing multiple rounds of training on the student feature extractor and the teacher feature extractor; The feature discriminator, the student feature extractor and the teacher feature extractor are trained in a cyclic iterative manner until a preset iteration stop condition is reached.
5. The text detection method according to claim 4, characterized in that: The text detection method also includes: When training the feature discriminator, marking the fused feature map output by the teacher feature extractor as 1, and marking the fused feature map output by the student feature extractor as 0; When the student feature extractor and the teacher feature extractor are trained, the fused feature map output by the teacher feature extractor and the fused feature map output by the student feature extractor are marked as 1.
6. The text detection method according to claim 4 or 5, characterized in that: When performing adversarial training on the student feature extractor, the teacher feature extractor and the feature discriminator, training is performed based on a preset loss function; The preset loss function is determined based on the sum of the true label loss function and the generative adversarial loss function; the true label loss function is determined based on the probability map loss, the binary map loss and the threshold map loss; the generative adversarial loss function is the cross entropy loss function.
7. A text detection device, characterized in that: include: A feature extraction module is used to: input a target image into a pre-trained student model to obtain a feature map group output by the student model, wherein the feature map group includes a plurality of feature maps of different scales, wherein the student model is obtained by joint training with a teacher model, and the parameter amount of the student model is less than the parameter amount of the teacher model; A feature fusion module is used to: fuse multiple feature maps of different scales in the feature map group to obtain a fused feature map; A probability calculation module, used to: input the fused feature map into the head network to obtain a probability map and a threshold map output by the head network, wherein the probability map reflects the probability that each pixel in the target image belongs to the text area, and the threshold map contains the threshold used by each pixel in the target image in the binarization process; A binary calculation module, used to: perform a binary calculation based on the probability map and the threshold map to obtain a binary map; The text region prediction module is used to determine the text region based on the binary image.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the text detection method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the text detection method according to any one of claims 1 to 6 is implemented.