Text detection model training method, text detection method, and electronic device
The text detection model training method enhances text detection accuracy by using a network architecture to learn text inter-difference features, improving the precision of text recognition in varied text styles.
Patent Information
- Application Number
- CN202210260489.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-16
AI Technical Summary
The existing text detection model lacks learning of the differential characteristics between different texts and cannot perform fine-grained text detection, resulting in different styles of text being classified as the same detection text box, affecting the accuracy of subsequent text recognition.
The feature extraction network is used to extract text feature maps, and the text backbone detection network and corner area detection network are trained respectively. The text detection model is trained in combination with the loss function, and the differential characteristics between texts are further mined through the corner area detection network.
It improves the accuracy of text detection, can distinguish different types of text in a more fine-grained manner, and improves the accuracy of text detection.
Smart Images

Figure CN114581925B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly provides a method for training a text detection model, a text detection method, and an electronic device. Background Art
[0002] The General Optical Character Recognition (General OCR) algorithm is a basic algorithm for carrying out various OCR services. Based on advanced deep learning technologies, it can recognize the text on different documents and bill pictures in various scenarios as editable text, thus greatly improving the information processing efficiency. Currently, the industry generally adopts a two-step General OCR algorithm strategy, that is, first, perform text detection on the input image to obtain the text positions; then, crop out the image slices containing only text according to the text positions, and send them into the text recognition model for recognition. Finally, the outputs of the two models are summarized to obtain the final result.
[0003] Among them, the general text detection task is a basic and important task in the General OCR algorithm. Its main goal is to obtain the positions of all texts from the input image (usually this position is represented by a quadrilateral that can contain the entire text and has the smallest area) as the input for the next text recognition model. Therefore, the accuracy of text detection will directly affect the overall effect of text recognition. However, due to the complex background and diverse scenarios of the input image, and the different styles and sizes of the texts it contains, how to quickly and accurately obtain the text detection result has become a challenging task.
[0004] Currently, the training of text detection models is based on the idea of semantic segmentation. That is, when performing text detection, the model will judge whether each pixel position in the picture belongs to text or background one by one, and then merge all the pixels judged as text. Take the minimum bounding rectangle for each merged text region to obtain the final detected text box. This method has a simple and clear idea and can achieve good results in most scenarios. However, its disadvantage is that it only simply learns the differences between text and background pixels in the picture, and lacks the learning of the differences between texts of different font types and different sizes. Therefore, it cannot accurately distinguish texts of different font types or different sizes in more detail, which leads to the phenomenon that texts of different styles are often output based on the same detected text box in the final detection result. When the style differences between the texts contained in the same detected text box are too large, it will cause a decrease in the accuracy of the subsequent text recognition model. Therefore, how to learn the differential features between different texts and perform text detection with higher granularity has become an urgent problem to be solved. Summary of the Invention
[0005] The present invention aims to solve the above technical problems, that is, to solve the problem that the existing text detection model lacks the learning of the differential features between different texts and cannot perform more fine-grained text detection.
[0006] In a first aspect, the present invention provides a method for training a text detection model. The text detection model includes a feature extraction network, a text backbone detection network, and a corner region detection network. The method includes:
[0007] Using the feature extraction network to extract the text feature map of the text training sample;
[0008] Taking the text feature map as input, training the text backbone detection network and the corner region detection network respectively to obtain the loss function of the text backbone detection network and the loss function of the corner region detection network;
[0009] Training the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain the trained text detection model.
[0010] In some embodiments, taking the text feature map as input and training the corner region detection network to obtain the loss function of the corner region detection network includes:
[0011] Using the corner region detection network to process the text feature map to obtain a corner region prediction map, and the corner region prediction map includes corner region sub-prediction maps corresponding to each corner region;
[0012] For each corner region sub-prediction map in the corner region prediction map, calculating the corner region loss function using the corner region sub-prediction map and the pre-stored corresponding corner region map;
[0013] Summing up the corner region loss functions of all the corner region sub-prediction maps in the corner region prediction map to obtain the loss function of the corner region detection network.
[0014] In some embodiments, the using the corner region detection network to process the text feature map to obtain a corner region prediction map includes:
[0015] For each text region in the text feature map, obtaining the center point coordinates of the text region and the midpoint coordinates of each border;
[0016] Using the center point coordinates of the text region and the midpoint coordinates of each border to divide the text region into a first azimuth corner region, a second azimuth corner region, a third azimuth corner region, and a fourth azimuth corner region;
[0017] Perform binarization processing on the first azimuth point region, the second azimuth point region, the third azimuth point region, and the fourth azimuth point region respectively to obtain the corner point region prediction map.
[0018] In some embodiments, using the text feature map as the input, training the text backbone detection network to obtain the loss function of the text backbone detection network, including:
[0019] Using the text backbone detection network to process the text feature map to obtain a text probability map, a boundary threshold prediction map, and a text binary prediction map;
[0020] Calculating a first sub-loss function using the text probability map and a pre-stored text binarization segmentation map, calculating a second sub-loss function using the boundary threshold prediction map and a pre-stored boundary threshold map, and calculating a third sub-loss function using the text binary prediction map and the pre-stored text binarization segmentation map, and obtaining the loss function of the text backbone detection network based on the first sub-loss function, the second sub-loss function, and the third sub-loss function;
[0021] Training the text detection model based on the loss function of the text backbone detection network and the loss function of the corner point region detection network to obtain the trained text detection model, including:
[0022] Calculating the loss function of the corner point region detection network using the corner point region prediction map and pre-stored corner point region maps;
[0023] Performing weighted summation on the first sub-loss function, the second sub-loss function, the third sub-loss function, and the loss function of the corner point region detection network to obtain a total loss function, and training the text detection model using the total loss function to obtain the trained text detection model.
[0024] In some embodiments, the using the text backbone detection network to process the text feature map to obtain a text probability map, a boundary threshold prediction map, and a text binary prediction map includes:
[0025] Performing shrinking processing and binarization processing on the text feature map to obtain the text probability map;
[0026] And, performing dilation processing on the text feature map and performing pixel normalization processing on the region other than the region corresponding to the text probability map in the dilated text feature map to obtain the boundary threshold prediction map;
[0027] Obtaining the text binary prediction map according to the text probability map and the boundary threshold prediction map.
[0028] In a second aspect, the present invention provides a text detection method, which performs text detection based on a pre-trained text detection model. The text detection model is trained by the text detection model training method described in any one of the above. The method includes:
[0029] Input the text image to be detected into the text detection model to obtain a text binary prediction map and a corner point region prediction map of the text image;
[0030] Determine and output the text detection result according to the number of corner point regions corresponding to each text region in the text binary prediction map in the corner point region prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner point region corresponding to the text region in the corner point region prediction map.
[0031] In some embodiments, the determining and outputting the text detection result according to the number of corner point regions corresponding to each text region in the text binary prediction map in the corner point region prediction map includes:
[0032] Judge whether the number of corner point regions corresponding to the same azimuth in the corner point region prediction map corresponding to each text region is greater than 1;
[0033] When the number of corner point regions corresponding to the same azimuth in the corner point region prediction map corresponding to each text region is greater than 1, divide the text region into multiple text sub-regions according to the corner point region prediction map corresponding to the text region; output the text detection result based on the multiple text sub-regions obtained by the division;
[0034] When the number of corner point regions corresponding to the same azimuth in the corner point region prediction map corresponding to each text region is less than or equal to 1, output the text detection result according to the text binary prediction map.
[0035] In some embodiments, the corner point region prediction map includes a corner point region sub-prediction map corresponding to each corner point region; the dividing the text region into multiple text sub-regions according to the corner point region prediction map corresponding to the text region includes:
[0036] For each corner point region in the text region, respectively obtain the text bounding box of the corner point region according to the corner point region sub-prediction map corresponding to the corner point region;
[0037] Calculate the center point coordinates of the corner point region according to the text bounding box;
[0038] According to the central point coordinates of each of the corner regions and the extension direction of the text region, for each of the corner regions in the same orientation, successively obtain the corner regions in other orientations that are within a preset range from the corner region in the current orientation;
[0039] Combine the corner region in the current orientation with the obtained corner regions in other orientations to obtain a set of corner regions corresponding to the corner region in the current orientation;
[0040] Use the sets of corner regions respectively corresponding to each of the corner regions in the same orientation to divide the text region into multiple text sub-regions, with each set of corner regions corresponding to one text sub-region.
[0041] In some embodiments, the determining and outputting the text detection result according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map includes:
[0042] Obtain the union region of the corner regions corresponding to the text region;
[0043] Calculate the intersection over union of the union region and the corresponding text region;
[0044] Determine whether the intersection over union of the union region and the corresponding text region in the binary prediction map is greater than a preset threshold;
[0045] When the intersection over union is greater than the preset threshold, output the text detection result based on the corner region prediction map corresponding to the text region;
[0046] When the intersection over union is less than or equal to the preset threshold, output the text detection result based on the text binary prediction map corresponding to the text region.
[0047] In some embodiments, the determining and outputting the text detection result according to the number of corner regions corresponding to each text region in the corner region prediction map and the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map includes:
[0048] Determine whether the number of corner regions corresponding to the same azimuth corner region in the corner region prediction map of each text region is greater than 1;
[0049] And, obtain the union region of the corner regions corresponding to the text region; calculate the intersection over union of the union region and the corresponding text region; determine whether the intersection over union of the union region and the corresponding text region in the binary prediction map is greater than a preset threshold;
[0050] When the number of corner region prediction maps corresponding to the same azimuth corner region in the corner region prediction map corresponding to each text region is greater than 1 and the intersection over union is greater than a preset threshold, a text detection result is output based on the corner region prediction map corresponding to the text region;
[0051] Otherwise, a text detection result is output based on the text binary prediction map corresponding to the text region.
[0052] In a third aspect, the present invention provides a text detection model training device. The text detection model includes a feature extraction network, a text backbone detection network, and a corner region detection network. The text detection model training device includes:
[0053] A feature extraction module, which is used to extract a text feature map of a text training sample by using the feature extraction network;
[0054] A text detection module, which is used to take the text feature map as an input, and train the text backbone detection network and the corner region detection network respectively to obtain a loss function of the text backbone detection network and a loss function of the corner region detection network;
[0055] A network training module, which is used to train the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain the trained text detection model.
[0056] In a fourth aspect, the present invention provides a text detection device. The text detection device performs text detection based on a pre-trained text detection model. The text detection model is trained by using the text detection model training method described in any one of the above. The text detection device includes:
[0057] A prediction module, which is used to input a text image to be detected into the text detection model to obtain a text binary prediction map and a corner region prediction map of the text image;
[0058] An output module, which is used to determine and output a text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map.
[0059] In a fifth aspect, the present invention provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the text detection model training method described in any one of the above or the text detection method described in any one of the above is implemented.
[0060] In a sixth aspect, the present invention provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory, and when the computer program is executed by the processor, the text detection model training method described in any one of the above or the text detection method described in any one of the above is implemented.
[0061] In the case of adopting the above technical solution, the present invention can use a feature extraction network to extract a text feature map of a text training sample; taking the text feature map as an input, training a text backbone detection network and a corner region detection network respectively to obtain a loss function of the text backbone detection network and a loss function of the corner region detection network; training a text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain a trained text detection model. By training the text detection model with the loss function of the text backbone detection network combined with the loss function of the corner region detection network, the differential features between different texts can be learned, and a more fine-grained text detection result can be provided, improving the accuracy of text detection.
[0062] On the other hand, the present invention can perform text detection based on the text detection model trained by using the above text detection model training method. By inputting a text image to be detected into the text detection model, a text binary prediction map and a corner region prediction map of the text image are obtained; according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map, a text detection result is determined and output. This method can effectively distinguish different types of texts and improve the accuracy of text detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The preferred embodiments of the present invention will be described below with reference to the drawings, in which:
[0064] Figure 1 is a schematic flowchart of a text detection model training method provided by an embodiment of the present invention;
[0065] Figure 2 is a schematic flowchart of a method for determining the loss function of a corner region detection network provided by an embodiment of the present invention;
[0066] Figure 3 is a schematic flowchart of a method for obtaining a corner region prediction map provided by an embodiment of the present invention;
[0067] Figure 4 is a schematic diagram of a text training sample with original labels provided by the present invention;
[0068] Figure 5It is a schematic flowchart of a method for obtaining a text probability map, a boundary threshold prediction map, and a text binary prediction map provided by an embodiment of the present invention;
[0069] Figure 6(1) is a schematic diagram of the text probability map provided by the present invention; Figure 6(2) is a schematic diagram of the boundary threshold prediction map provided by the present invention;
[0070] Figure 7 It is a schematic flowchart of a text detection method provided by an embodiment of the present invention;
[0071] Figure 8 It is a schematic flowchart of a method for outputting a text detection result according to the number of corner regions provided by an embodiment of the present invention;
[0072] Figure 9(1) is a corner region prediction map of a text region provided by a specific example of the present invention; Figure 9(2) is a text detection map output after dividing the text region into multiple text sub-regions provided by a specific example of the present invention;
[0073] Figure 10 It is a schematic flowchart of a method for outputting a text detection result according to the intersection over union provided by an embodiment of the present invention;
[0074] Figure 11 It is a schematic flowchart of a method for outputting a text detection result according to the number of corner regions and the intersection over union provided by an embodiment of the present invention;
[0075] Figure 12 It is a schematic structural diagram of a text detection model training device provided by an embodiment of the present invention;
[0076] Figure 13 It is a schematic structural diagram of a text detection device provided by an embodiment of the present invention;
[0077] Figure 14 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0078] The following describes some embodiments of the present invention with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0079] The General Optical Character Recognition (General OCR) algorithm is a fundamental algorithm for carrying out various OCR services. Based on cutting-edge deep learning technologies, it can recognize the text on different documents and bill images in various scenarios into editable text, thus significantly improving the information processing efficiency. Currently, the industry generally adopts a two-step General OCR algorithm strategy, that is, first perform text detection on the input image to obtain the text positions, then crop out the image slices containing only text according to the text positions, and then send them into the text recognition model for recognition. Finally, the outputs of the two models are summarized to obtain the final result.
[0080] Among them, the general text detection task is a basic and important task in the General OCR algorithm. Its main goal is to obtain the positions of all texts from the input image (usually this position is represented by a quadrilateral that can contain the entire text and has the smallest area) as the input for the next text recognition model. Therefore, the accuracy of text detection will directly affect the overall effect of text recognition. However, due to the complex background and diverse scenarios of the input image, and the different styles and sizes of the texts it contains, how to quickly and accurately obtain the text detection result has become a challenging task.
[0081] Currently, the training of text detection models is based on the idea of semantic segmentation. That is, when performing text detection, the model will judge whether each pixel position in the picture belongs to text or background one by one, and then merge all the pixels judged as text, and take the minimum bounding rectangle for each merged text region to obtain the final detected text box. This method has a simple and clear idea and can achieve good results in most scenarios. However, its disadvantage is that it only simply learns the differences between text and background pixels in the picture, and lacks the learning of the differences between texts of different font types and different sizes. Therefore, it cannot accurately distinguish texts of different font types or different sizes in more detail, which leads to the phenomenon that texts of different styles are often output based on the same detected text box in the final detection result. When the style differences between the texts contained in the same detected text box are too large, it will cause the accuracy of the subsequent text recognition model to decline. Therefore, how to learn the differential features between different texts and perform text detection with higher fine-grainedness has become an urgent problem to be solved.
[0082] In view of this, the present invention provides a method for training a text detection model, which extracts a text feature map of a text training sample by using a feature extraction network; uses the text feature map as an input to train a text backbone detection network and a corner region detection network respectively, to obtain a loss function of the text backbone detection network and a loss function of the corner region detection network; trains the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network, so as to obtain a trained text detection model. By further learning the text feature map by using the corner region detection network on the basis of obtaining a preliminary text detection result by using the text backbone detection network, fine-grained features between different text types are further mined, the differential features between different texts are learned, and the text detection model is trained by combining the loss function of the text backbone detection network with the loss function of the corner region detection network, which can improve the accuracy of the text detection model and achieve more fine-grained text detection.
[0083] See Figure 1 as described Figure 1 FIG. is a schematic flowchart of a method for training a text detection model provided by an embodiment of the present invention. The text detection model includes a feature extraction network, a text backbone detection network, and a corner region detection network. The method for training the text detection model may include:
[0084] Step S11: Extract a text feature map of a text training sample by using a feature extraction network;
[0085] Step S12: Use the text feature map as an input to train a text backbone detection network and a corner region detection network respectively, to obtain a loss function of the text backbone detection network and a loss function of the corner region detection network;
[0086] Step S13: Train the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network, so as to obtain a trained text detection model.
[0087] In some embodiments, the feature extraction network may adopt a convolutional neural network with a pyramid architecture.
[0088] In some embodiments, see Figure 2 as shown Figure 2 FIG. is a schematic flowchart of a method for determining the loss function of the corner region detection network provided by an embodiment of the present invention. In step S12, the corner region detection network is trained by using the text feature map as an input to obtain the loss function of the corner region detection network, which may include:
[0089] Step S120: Process the text feature map using a corner region detection network to obtain a corner region prediction map, where the corner region prediction map includes corner region sub-prediction maps corresponding to respective corner regions;
[0090] Step S121: For each corner region sub-prediction map in the corner region prediction map, calculate a corner region loss function using the corner region sub-prediction map and a pre-stored corresponding corner region map;
[0091] Step S122: Sum up the corner region loss functions of all corner region sub-prediction maps in the corner region prediction map to obtain the loss function of the corner region detection network.
[0092] Among them, the text feature map may include at least one text region. In some embodiments, for a text region in the corner region prediction map corner_pred_map, the text region may include a corner region sub-prediction map lt_corner_pred_map corresponding to a first azimuth corner region, a corner region sub-prediction map rt_corner_pred_map corresponding to a second azimuth corner region, a corner region sub-prediction map rb_corner_pred_map corresponding to a third azimuth corner region, and a corner region sub-prediction map lb_corner_pred_map corresponding to a fourth azimuth corner region.
[0093] In some embodiments, as shown in Figure 3 shown, Figure 3 is a schematic flowchart of a method for obtaining a corner region prediction map provided by an embodiment of the present invention; Step S120 may specifically include:
[0094] Step S1201: For each text region in the text feature map, obtain the center point coordinates of the text region and the midpoint coordinates of each border;
[0095] Step S1202: Use the center point coordinates of the text region and the midpoint coordinates of each border to divide the text region into a first azimuth corner region, a second azimuth corner region, a third azimuth corner region, and a fourth azimuth corner region;
[0096] Step S1203: Perform binarization processing on the first azimuth corner region, the second azimuth corner region, the third azimuth corner region, and the fourth azimuth corner region respectively to obtain a corner region prediction map.
[0097] In some embodiments, Step S1201 may specifically be: For each text region in the text feature map, use the four vertex coordinates of the text box G corresponding to the text region to calculate the center point coordinates of each text region and the midpoint coordinates of each border.
[0098] In some embodiments, step S1202 may be specifically: form a dividing line according to the center point coordinates and the midpoint coordinates of the upper and lower borders of the text area, form another dividing line according to the center point coordinates and the midpoint coordinates of the left and right borders of the text area, and divide the text area into a first azimuth point area, a second azimuth point area, a third azimuth point area, and a fourth azimuth point area based on the two dividing lines whose intersection points pass through the center point. As an example, after dividing the text area, four corner point areas corresponding to the upper left, upper right, lower left, and lower right can be generated.
[0099] In some embodiments, step S1203 may be specifically to perform binarization processing on the first azimuth point area, the second azimuth point area, the third azimuth point area, and the fourth azimuth point area respectively, fill the internal pixels of each text area with "1", corresponding to white, and fill the pixels of the area outside the text area in the text feature map with "0", corresponding to black.
[0100] In some embodiments, step S121 may be specifically to calculate the corner region Dice Loss loss function for each corner region sub-prediction map in the corner region prediction map by using the corner region sub-prediction map and the pre-stored corresponding corner region map.
[0101] As an example, use the corner region sub-prediction map lt_corner_pred_map corresponding to the first azimuth point area and the pre-stored corner region map lt_corner_map of the first azimuth to calculate the corner region loss function l of the first azimuth point area through the following expression corner_lt :
[0102]
[0103] where pred represents the corner region sub-prediction map corresponding to the first azimuth point area, label represents the corner region map of the first azimuth, pred∩label is the intersection of the pixels with "1" in the corner region sub-prediction map corresponding to the first azimuth point area and the corner region map of the first azimuth, and pred+label is the sum of the number of pixels with "1" in the corner region sub-prediction map corresponding to the first azimuth point area and the corner region map of the first azimuth.
[0104] Similarly, the corner region loss function l of the second azimuth point area can be obtained respectively corner_rt , the corner region loss function l of the third azimuth point area corner_lb and the corner region loss function l of the fourth azimuth point area corner_rb .
[0105] Finally, the loss function of the corner region detection network can be obtained through the following expression: l corner= l corner_lt + l corner_rt + l corner_rb + l corner_lb 。
[0106] It should be noted that the corresponding corner region maps stored in advance can be obtained through the following steps:
[0107] Select some text training samples, generate original labels for each text region in the selected text training samples. As an example, see Figure 4 As shown, for each text region, generate a quadrilateral text box G and record the coordinates of the four vertices of each text box G. Label the text box G according to the coordinates of the four vertices of the text box G to generate the original label of the text region;
[0108] Using the coordinates of the four vertices of the text box G corresponding to the text region, calculate the center point coordinates of each text region and the midpoint coordinates of each side frame; form a dividing line according to the center point coordinates and the midpoint coordinates of the upper and lower side frames of the text region, and form another dividing line according to the center point coordinates and the midpoint coordinates of the left and right side frames of the text region. Divide the text region into a first azimuth corner region, a second azimuth corner region, a third azimuth corner region, and a fourth azimuth corner region based on the two dividing lines whose intersection points pass through the center point. As an example, after dividing the text region, four corner regions corresponding to the upper left, upper right, lower left, and lower right can be generated;
[0109] Perform binarization processing on the first azimuth corner region, the second azimuth corner region, the third azimuth corner region, and the fourth azimuth corner region respectively. Fill the internal pixels of each text region with "1", corresponding to white, and fill the pixels of the region outside the text region in the text feature map with "0", corresponding to black, so as to obtain the corner region maps of each azimuth and store them.
[0110] In some embodiments, in step S12, using the text feature map as the input to train the text backbone detection network, the loss function obtained for the text backbone detection network may include:
[0111] Use the text backbone detection network to process the text feature map to obtain a text probability map, a boundary threshold prediction map, and a text binary prediction map;
[0112] Calculate the first sub-loss function using the text probability map and the pre-stored text binarization segmentation map, calculate the second sub-loss function using the boundary threshold prediction map and the pre-stored boundary threshold map, and calculate the third sub-loss function using the text binary prediction map and the pre-stored text binarization segmentation map. Obtain the loss function of the text backbone detection network based on the first sub-loss function, the second sub-loss function, and the third sub-loss function.
[0113] In some embodiments, referring to Figure 5 as shown Figure 5 FIG. 5 is a schematic flowchart of a method for obtaining a text probability map, a boundary threshold prediction map, and a text binary prediction map provided by an embodiment of the present invention. Using a text backbone detection network to process a text feature map to obtain a text probability map, a boundary threshold prediction map, and a text binary prediction map may include:
[0114] Step S123: Performing a shrinking process and a binarization process on the text feature map to obtain a text probability map;
[0115] And,
[0116] Step S124: Performing a dilation process on the text feature map and performing pixel normalization on regions other than the region corresponding to the text probability map in the dilated text feature map to obtain a boundary threshold prediction map;
[0117] Step S125: Obtaining a text binary prediction map according to the text probability map and the boundary threshold prediction map.
[0118] In some embodiments, step S123 may specifically be to perform a shrinking process on the text feature map by using the Vatti Clipping algorithm and setting an offset. For each text region in the text feature map, a corresponding shrunk text box G s is obtained, and the offset D can be set by using the following expression:
[0119]
[0120] where A represents the area of the text box G corresponding to the text region, L represents the perimeter of the text box G, and r is an empirical value. In some embodiments, the value of r may be 0.4.
[0121] Then, the internal region pixels of the shrunk text box G s are filled with "1", corresponding to white; the pixels in the region other than the text box G s in the shrunk text feature are filled with "0", corresponding to black, so as to obtain the text probability map preb_map corresponding to the text feature map. Referring to FIG. 6(1) as shown, FIG. 6(1) is a schematic diagram of the text probability map provided by the present invention.
[0122] In some embodiments, step S124 may specifically be to perform pixel normalization on each pixel in the region of the text feature map after dilation processing except for the region corresponding to the text probability map, to obtain a boundary threshold prediction map thresh_pred_map, as shown in FIG. 6(2). FIG. 6(2) is a schematic diagram of the boundary threshold prediction map provided by the present invention, where the region corresponding to the text probability map is the region after contraction processing of the text feature map.
[0123] In some embodiments, step S125 may specifically be to binarize the text probability map preb_map using the boundary threshold prediction map thresh_pred_map to obtain a text binary prediction map binary_pred_map, and the text binary prediction map may be used as a preliminary text detection result.
[0124] In some embodiments, the binary cross-entropy loss function of the text probability map pred_map and the pre-stored text binarized segmentation map prob_map may be calculated through the following expression as the first sub-loss function l s :
[0125]
[0126] where y i is the label of each pixel on the pre-stored text binarized segmentation map prob_map, x i is the predicted probability value of each pixel on the text probability map pred_map, and N is the number of pixels in the text feature map.
[0127] In some embodiments, the L1 norm loss function of the boundary threshold prediction map thresh_pred_map and the pre-stored boundary threshold map thresh_map may be calculated through the following expression as the second sub-loss function l t :
[0128]
[0129] where y i is the label of each pixel on the pre-stored boundary threshold map thresh_map, x i is the predicted probability value of each pixel on the boundary threshold prediction map thresh_pred_map.
[0130] In some embodiments, the binary cross-entropy loss function of the text binary prediction map binary_pred_map and the pre-stored text binarized segmentation map prob_map may be calculated using the following expression as the third sub-loss function l b :
[0131]
[0132] Among them, y i is the label of each pixel on the pre-stored text binarized segmentation map prob_map, and x i is the predicted probability value of each pixel on the text binary prediction map binary_pred_map.
[0133] It should be noted that the pre-stored text binarized segmentation map prob_map can be obtained by using a method similar to step S123 based on the text training samples with original labels shown in Figure 4 ; the pre-stored boundary threshold map thresh_map can be obtained by using a method similar to step S124 based on the text training samples with original labels shown in Figure 4 .
[0134] Based on the method for determining the loss function of the above corner region detection network and the method for determining the loss function of the text backbone detection network, in some embodiments, step S13 may specifically be:
[0135] Calculate the loss function of the corner region detection network by using the corner region prediction map and the pre-stored corner region maps;
[0136] Perform weighted summation on the first sub-loss function, the second sub-loss function, the third sub-loss function, and the loss function of the corner region detection network to obtain the total loss function, and use the total loss function to train the text detection model to obtain a trained text detection model.
[0137] As an example, the total loss function can be expressed as L = l s + α × l b + β × l t + γ × l corner , where α, β, and γ represent the weights corresponding to different loss functions and can be valued according to experience.
[0138] The above is a method for training a text detection model provided by an embodiment of the present invention. By using a feature extraction network to extract the text feature map of a text training sample; using the text feature map as input, training the text backbone detection network and the corner region detection network respectively to obtain the loss function of the text backbone detection network and the loss function of the corner region detection network; training the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain a trained text detection model. By further learning the text feature map using the corner region detection network on the basis of obtaining the preliminary text detection result using the text backbone detection network, the fine-grained features between different text types are further mined, the differential features between different texts are learned, and the text detection model is trained by combining the loss function of the text backbone detection network with the loss function of the corner region detection network, which can improve the accuracy of the text detection model and achieve more fine-grained text detection.
[0139] The present invention also provides a text detection method, which performs text detection based on a pre-trained text detection model, and the text detection model is trained based on the text detection model training method described in any one of the above embodiments.
[0140] See Figure 7 shown in Figure 7 is a schematic flowchart of the text detection method provided by an embodiment of the present invention, which may include:
[0141] Step S21: Input the text image to be detected into the text detection model to obtain the text binary prediction map and the corner region prediction map of the text image;
[0142] Step S22: Determine and output the text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map.
[0143] In some embodiments, when determining and outputting the text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map in step S22, the text detection result can be output according to whether the corresponding number of corner regions is greater than 4; when it is greater than 4, the text detection result is output based on the corner region prediction map; when it is less than or equal to 4, the text detection result is output based on the text binary prediction map.
[0144] In other embodiments, when determining and outputting the text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map in step S22, see Figure 8 shown in
[0145] Step S31: Determine whether the number of corner region areas corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1;
[0146] Step S32: When the number of corner region areas corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1, divide the text region into multiple text sub-regions according to the corner region prediction map corresponding to the text region; output a text detection result based on the multiple text sub-regions obtained by the division;
[0147] Step S33: When the number of corner region areas corresponding to the same azimuth in the corner region prediction map corresponding to each text region is less than or equal to 1, output a text detection result according to the text binary prediction map.
[0148] In some embodiments, step S31 may specifically be to match the pixel set corresponding to the text region in the text binary prediction map with the pixel sets corresponding to each corner region in the corner region prediction map, so as to determine the corner regions corresponding to each text region; determine the number of corner region areas with the same azimuth in the corner regions corresponding to the text region.
[0149] In some embodiments, the corner region prediction map includes corner region sub-prediction maps corresponding to each corner region, and step S32 may specifically be:
[0150] For each corner region in the text region, respectively obtain the text bounding box of the corner region according to the corner region sub-prediction map corresponding to the corner region;
[0151] According to the text bounding box, calculate the center point coordinates of the corner region;
[0152] According to the center point coordinates of each corner region and the extension direction of the text region, for each corner region with the same azimuth, sequentially obtain the corner regions in other azimuths whose distance from the corner region in the current azimuth is within a preset range;
[0153] Combine the corner region in the current azimuth with the obtained corner regions in other azimuths to obtain a corner region set corresponding to the corner region in the current azimuth;
[0154] Use the corner region sets corresponding to each corner region with the same azimuth to divide the text region into multiple text sub-regions, and each corner region set corresponds to a text sub-region.
[0155] As an example, according to the text binary prediction map, obtain the pixel set {G pred} where all pixels are "1", that is, it corresponds to all text regions in the text image to be detected; according to the corner region prediction map, obtain the pixel sets of the corner regions in four azimuths corresponding to all text regions {LTpred , RT pred , RB pred , LB pred}, through pixel matching, the corner regions included in each text region are determined. As shown in Fig. 9(1), the largest quadrilateral frame in Fig. 9(1) is the text region predicted in the text binary prediction map, and each relatively smaller quadrilateral frame is the multiple corner regions predicted corresponding to the current text region in the corner region prediction map, that is, the corner region prediction result corresponding to the current text region is {lt1, lt2, lt3, rt1, rt2, rt3, rb1, rb2, rb3, lb1, lb2, lb3}, where lt represents the upper left position, rt represents the upper right position, lb represents the lower left position, and rb represents the lower right position; thus, it can be determined that the number of corner regions in any one position is greater than 1.
[0156] For each corner region in the text region, the text bounding box of the corner region is obtained respectively according to the corner region sub-prediction map corresponding to the corner region, that is, the relatively smaller quadrilateral frame in Fig. 9(1); according to the text bounding box, the center point coordinates of the corner region are calculated; the abscissas in the center point coordinates of each corner region are sorted, and based on the corner regions in the same position after sorting, such as lt1, lt2, and lt3, the corner regions in the other three positions that are closest to the corner region of lt1, the corner regions in the other three positions that are closest to the corner region of lt2, and the corner regions in the other three positions that are closest to the corner region of lt3 are obtained in turn; the corner region in the current position is combined with the corner regions in the other obtained positions to obtain the corner region set corresponding to the corner region in the current position, such as obtaining the corner region set G1 = {lt1, rt1, lb1, rb1} corresponding to the corner region of lt1, the corner region set G2 = {lt2, rt2, lb2, rb2} corresponding to the corner region of lt2; the corner region set G3 = {lt3, rt3, lb3, rb3} corresponding to the corner region of lt3; furthermore, the text region can be divided into multiple text sub-regions {G1, G2, G3} by using the corner region sets corresponding to the corner regions of lt1, lt2, and lt3 respectively. As shown in Fig. 9(2), the text region can be divided into 3 text sub-regions.
[0157] In some embodiments, when determining and outputting the text detection result according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map in step S22, as shown in Figure 10 shown, it may specifically include:
[0158] Step S41: Obtain the union region of the corner regions corresponding to the text region;
[0159] Step S42: Calculate the intersection over union (IoU) between the union region and the corresponding text region;
[0160] Step S43: Determine whether the IoU between the union region and the corresponding text region is greater than a preset threshold;
[0161] Step S44: When the IoU is greater than the preset threshold, output the text detection result based on the corner region prediction map corresponding to the text region;
[0162] Step S45: When the IoU is less than or equal to the preset threshold, output the text detection result based on the text binary prediction map corresponding to the text region.
[0163] When the IoU is greater than the preset threshold, output the text detection result based on the corner region prediction map corresponding to the text region, so as to divide the text region into multiple text sub-regions, realize the fine-grained detection of the text region, identify different types of text, and improve the accuracy of text detection; when the IoU is less than or equal to the preset threshold, it indicates that the fine-grained features of the text region are insufficient, and there is no need to further divide the text region, and the text detection result can be directly output according to the text binary prediction map.
[0164] In some other embodiments, when determining and outputting the text detection result according to the number of corner regions corresponding to each text region in the corner region prediction map and the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map in step S22, refer to Figure 11 as shown, it may include:
[0165] Step S51: Determine whether the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1;
[0166] And, Step S52: Obtain the union region of the corner regions corresponding to the text region; calculate the IoU between the union region and the corresponding text region; determine whether the IoU between the union region and the text region is greater than the preset threshold;
[0167] Step S53: When the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1 and the IoU is greater than the preset threshold, output the text detection result based on the corner region prediction map corresponding to the text region; otherwise, execute Step S54;
[0168] Step S54: Output the text detection result based on the text binary prediction map corresponding to the text region.
[0169] Among them, step S51 can be implemented in the same way as step S31 described above, and step S52 can be implemented in the same way as steps S41 to S43 described above. Details are not described herein again. Please refer to the description in the above text.
[0170] In some embodiments, step S53 can be specifically that when the number of corner point regions corresponding to the same orientation in the text region is greater than 1 and the intersection over union is greater than a preset threshold, the text region is divided into multiple text sub-regions based on the corner point region prediction map corresponding to the text region; and a text detection result is output based on the multiple text sub-regions. The method of dividing the text region into multiple text sub-regions can be the same as that in step S32.
[0171] In some embodiments, step S54 can be specifically that when the number of corner point regions corresponding to the same orientation in the corner point region prediction map corresponding to each text region is less than or equal to 1, and / or the intersection over union is less than or equal to the preset threshold, a text detection result is output based on the text binary prediction map corresponding to the text region.
[0172] In the embodiments of the present invention, by jointly determining the output of the text detection result according to the number of corner point regions corresponding to each text region in the text binary prediction map in the corner point region prediction map and the intersection over union of each text region in the text binary prediction map and the corner point regions corresponding to the text region in the corner point region prediction map, the effectiveness and accuracy of text detection can be further improved.
[0173] It should be noted that in some embodiments, it is also possible to first determine whether the number of corner point regions is greater than 1 according to the number of corner point regions corresponding to each text region in the text binary prediction map; when the number of corner point regions is less than or equal to 1, the text detection result can be directly output according to the text binary prediction map; when the number of corner point regions is greater than 1, the text region is divided to obtain multiple text sub-regions, the intersection over union of the union region of the multiple text sub-regions and the corresponding text region is calculated, and it is determined whether the intersection over union of the union region and the text region is greater than the preset threshold. When it is greater than the threshold, the text detection result is output based on the multiple text sub-regions; when it is less than or equal to the threshold, the text detection result is output based on the text binary prediction map. Thus, while ensuring the accuracy of text detection, the speed of text detection can also be improved.
[0174] See Figure 12 as shown in Figure 12 is a schematic structural diagram of a text detection model training device provided by an embodiment of the present invention. The text detection model includes a feature extraction network, a text backbone detection network, and a corner point region detection network. The text detection model training device includes:
[0175] A feature extraction module 61, which is used to extract a text feature map of a text training sample by using the feature extraction network;
[0176] A text detection module 62, which is used to take a text feature map as an input, train a text backbone detection network and a corner region detection network respectively, and obtain a loss function of the text backbone detection network and a loss function of the corner region detection network;
[0177] A network training module 63, which is used to train a text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network, so as to obtain a trained text detection model.
[0178] The text detection model training device provided by the present invention can be used to execute the above-mentioned text detection model training method, and achieve the same beneficial effects as the text detection model training method in the above embodiments. Further, it should be understood that since the setting of each module is only for explaining the functional units of the device of the present invention, the physical devices corresponding to these modules can be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, Figure 12 the number of each module in [] is only illustrative. Those skilled in the art can understand that each module in the device can be adaptively split or combined. Such splitting or combining of specific modules will not cause the technical solution to deviate from the principle of the present invention. Therefore, the technical solutions after splitting or combining will all fall within the protection scope of the present invention.
[0179] See Figure 13 as shown in Figure 13 is a structural schematic diagram of a text detection device provided by an embodiment of the present invention. The text detection device performs text detection based on a pre-trained text detection model, and the text detection model is trained based on the text detection model training method described in any of the above embodiments. The text detection device includes:
[0180] A prediction module 71, which is used to input a text image to be detected into the text detection model to obtain a text binary prediction map and a corner region prediction map of the text image;
[0181] An output module 72, which is used to determine and output a text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map.
[0182] The text detection device provided by the present invention can be used to execute the above-mentioned text detection method, achieving the same beneficial effects as the text detection method in the above-mentioned embodiments. Further, it should be understood that since the setting of each module is only to illustrate the functional units of the device of the present invention, the physical devices corresponding to these modules can be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, Figure 13 the number of each module in
[0183] is only illustrative. Those skilled in the art can understand that the various modules in the device can be adaptively split or combined. Such splitting or combination of specific modules will not cause the technical solution to deviate from the principle of the present invention. Therefore, the technical solutions after splitting or combination will all fall within the protection scope of the present invention.
[0184] On the other hand, the present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it can implement the text detection model training method in any of the above embodiments or can implement the text detection method in any of the above embodiments. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present invention is a non-transitory computer-readable storage medium.
[0185] See Figure 14 as shown in Figure 14 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. It may include a memory 81 and a processor 82. A computer program is stored in the memory 81, and the computer program includes, but is not limited to, a program for executing the method in the above method embodiments. When the computer program is executed by the processor 82, it can implement the text detection model training method in any of the above embodiments or can implement the text detection method in any of the above embodiments.
[0186] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A method for training a text detection model, characterized in that, The text detection model includes a feature extraction network, a text backbone detection network, and a corner region detection network. The method includes: Using the feature extraction network to extract the text feature map of the text training sample; Taking the text feature map as the input, training the text backbone detection network and the corner region detection network respectively to obtain the loss function of the text backbone detection network and the loss function of the corner region detection network; Training the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain the trained text detection model; wherein, taking the text feature map as the input and training the corner region detection network to obtain the loss function of the corner region detection network includes: Using the corner region detection network to process the text feature map to obtain a corner region prediction map, where the corner region prediction map includes corner region sub-prediction maps corresponding to each corner region; For each corner region sub-prediction map in the corner region prediction map, calculating the corner region loss function using the corner region sub-prediction map and the pre-stored corresponding corner region map; Summing up the corner region loss functions of all the corner region sub-prediction maps in the corner region prediction map to obtain the loss function of the corner region detection network; The process of using the corner region detection network to process the text feature map to obtain the corner region prediction map includes: For each text region in the text feature map, obtaining the center point coordinates of the text region and the midpoint coordinates of each border; Using the center point coordinates of the text region and the midpoint coordinates of each border to divide the text region into a first azimuth corner region, a second azimuth corner region, a third azimuth corner region, and a fourth azimuth corner region; Performing binarization processing on the first azimuth corner region, the second azimuth corner region, the third azimuth corner region, and the fourth azimuth corner region respectively to obtain the corner region prediction map.
2. The method according to claim 1, characterized in that Taking the text feature map as the input and training the text backbone detection network to obtain the loss function of the text backbone detection network includes: Using the text backbone detection network to process the text feature map to obtain a text probability map, a boundary threshold prediction map, and a text binary prediction map; Calculating a first sub-loss function using the text probability map and the pre-stored text binarization segmentation map, calculating a second sub-loss function using the boundary threshold prediction map and the pre-stored boundary threshold map, and calculating a third sub-loss function using the text binary prediction map and the pre-stored text binarization segmentation map, and obtaining the loss function of the text backbone detection network based on the first sub-loss function, the second sub-loss function, and the third sub-loss function; Training the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain the trained text detection model includes: Calculate the loss function of the corner region detection network by using the corner region prediction map and each pre-stored corner region map; Perform weighted summation on the first sub-loss function, the second sub-loss function, the third sub-loss function and the loss function of the corner region detection network to obtain the total loss function, and use the total loss function to train the text detection model to obtain the trained text detection model.
3. The method according to claim 2, wherein The processing of the text feature map by using the text backbone detection network to obtain the text probability map, the boundary threshold prediction map and the text binary prediction map includes: Performing shrinking processing and binarization processing on the text feature map to obtain the text probability map; And, performing dilation processing on the text feature map and performing pixel normalization processing on the regions other than the regions corresponding to the text probability map in the dilated text feature map to obtain the boundary threshold prediction map; Obtain the text binary prediction map according to the text probability map and the boundary threshold prediction map.
4. A text detection method, characterized in that, The method performs text detection based on a pre-trained text detection model, and the text detection model is trained based on the text detection model training method described in any one of claims 1 to 3. The method includes: Input the text image to be detected into the text detection model to obtain the text binary prediction map and the corner region prediction map of the text image; Determine and output the text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map.
5. The method according to claim 4, wherein The determining and outputting the text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map includes: Judge whether the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1; When the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1, divide the text region into multiple text sub-regions according to the corner region prediction map corresponding to the text region; output the text detection result based on the multiple divided text sub-regions; When the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is less than or equal to 1, output the text detection result according to the text binary prediction map.
6. The method according to claim 5, characterized in that The corner region prediction map includes corner region sub-prediction maps corresponding to each corner region; the dividing the text region into multiple text sub-regions according to the corner region prediction map corresponding to the text region includes: For each corner region in the text region, respectively obtain the text bounding box of the corner region according to the corner region sub-prediction map corresponding to the corner region; According to the text bounding box, calculate the center point coordinates of the corner region; According to the center point coordinates of each of the corner regions and the extension direction of the text region, for each of the corner regions in the same orientation, successively obtain the corner regions in other orientations that are within a preset range from the corner region in the current orientation; Combine the corner region in the current orientation with the obtained corner regions in other orientations to obtain a set of corner regions corresponding to the corner region in the current orientation; Use the sets of corner regions respectively corresponding to each of the corner regions in the same orientation to divide the text region into multiple text sub-regions, with each set of corner regions corresponding to one text sub-region.
7. The method according to claim 4, characterized in that, The determining and outputting of the text detection result according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map includes: Obtain the union region of the corner regions corresponding to the text region; Calculate the intersection over union of the union region and the corresponding text region; Judge whether the intersection over union of the union region and the corresponding text region is greater than a preset threshold; When the intersection over union is greater than the preset threshold, output the text detection result based on the corner region prediction map corresponding to the text region; When the intersection over union is less than or equal to the preset threshold, output the text detection result based on the text binary prediction map corresponding to the text region.
8. The method according to claim 4, characterized in that, The determining and outputting of the text detection result according to the number of corner regions corresponding to each text region in the corner region prediction map and the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map includes: Judge whether the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1; And, obtain the union region of the corner regions corresponding to the text region; calculate the intersection over union of the union region and the corresponding text region; judge whether the intersection over union of the union region and the corresponding text region is greater than a preset threshold; When the number of corner regions corresponding to the same azimuth in the corner region prediction map corresponding to each text region is greater than 1 and the intersection over union is greater than the preset threshold, output the text detection result based on the corner region prediction map corresponding to the text region; Otherwise, output the text detection result based on the text binary prediction map corresponding to the text region.
9. A text detection model training device, characterized in that, The text detection model includes a feature extraction network, a text backbone detection network, and a corner region detection network. The text detection model training device includes: A feature extraction module, which is used to extract the text feature map of the text training sample by using the feature extraction network; A text detection module, which is used to use the text feature map as input to train the text backbone detection network and the corner region detection network respectively, to obtain the loss function of the text backbone detection network and the loss function of the corner region detection network; A network training module, which is used to train the text detection model based on the loss function of the text backbone detection network and the loss function of the corner region detection network to obtain the trained text detection model; wherein, using the text feature map as the input to train the corner region detection network to obtain the loss function of the corner region detection network includes: Processing the text feature map by using the corner region detection network to obtain a corner region prediction map, where the corner region prediction map includes corner region sub-prediction maps corresponding to each corner region; For each corner region sub-prediction map in the corner region prediction map, calculating a corner region loss function by using the corner region sub-prediction map and a pre-stored corresponding corner region map; Summing up the corner region loss functions of all the corner region sub-prediction maps in the corner region prediction map to obtain the loss function of the corner region detection network; the processing the text feature map by using the corner region detection network to obtain the corner region prediction map includes: For each text region in the text feature map, obtaining the center point coordinates of the text region and the midpoint coordinates of each border; Using the center point coordinates of the text region and the midpoint coordinates of each border to divide the text region into a first azimuth corner region, a second azimuth corner region, a third azimuth corner region, and a fourth azimuth corner region; Performing binarization processing on the first azimuth corner region, the second azimuth corner region, the third azimuth corner region, and the fourth azimuth corner region respectively to obtain the corner region prediction map.
10. A text detection device, characterized in that, The text detection device performs text detection based on a pre-trained text detection model, and the text detection model is trained based on the text detection model training method according to any one of claims 1 to 3. The text detection device includes: A prediction module, which is used to input a text image to be detected into the text detection model to obtain a text binary prediction map and a corner region prediction map of the text image; An output module, which is used to determine and output a text detection result according to the number of corner regions corresponding to each text region in the text binary prediction map in the corner region prediction map and / or according to the matching degree between each text region in the text binary prediction map and the corner region corresponding to the text region in the corner region prediction map.
11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it implements the text detection model training method according to any one of claims 1 to 3 or implements the text detection method according to any one of claims 5 to 9.
12. An electronic device, characterized in that, Including a memory and a processor, a computer program is stored in the memory, and when the computer program is executed by the processor, it implements the text detection model training method according to any one of claims 1 to 3 or implements the text detection method according to any one of claims 5 to 9.
Citation Information
Patent Citations
Text detection method and device for natural scene and storage medium
CN113033558A
Model training method, device, text detection method, device and lightweight network model
CN113780283A