Text positioning method fused with Bessel control point

By fusing the text positioning method of Bessel control points, using the feature extraction network with deep separable convolution and attention mechanisms, the accuracy and robustness of complex text detection in natural scenes are solved, and more efficient text area detection and recognition are achieved.

CN120356221APending Publication Date: 2025-07-22NORTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311708296.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Text detection in natural scenes In complex scenes, especially in the case of image proportion imbalance and irregular curved text, the detection accuracy and robustness are insufficient, and the existing technology is difficult to effectively improve.

Method used

The text positioning method that fuses Bezier control points is adopted, combined with a feature extraction network that can separate the depth convolution and attention mechanism, and through Bezier curve fitting and feature fusion, probability maps and coordinate regression of the text area are generated, and the shape and positioning of the bounding box of the text area are optimized.

Benefits of technology

It significantly improves the detection speed and accuracy of irregular texts, and improves the application level of character recognition in fields such as robot navigation, autonomous driving and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356221A_ABST
    Figure CN120356221A_ABST
Patent Text Reader

Abstract

The invention discloses a text positioning method fused with Bezier control points, which comprises the following steps of: performing curve fitting on a text bounding box in a training set to obtain control points of a third-order Bezier curve, extracting feature information in an image based on a lightweight feature extraction network, and fusing multi-scale features through up-sampling and a convolution attention module to obtain a text positioning result. A probability graph of a character region and coordinate regression of a Bessel control point are respectively generated through convolution operation, region loss is calculated according to the probability graph and a text bounding box truth value, coordinate loss is calculated according to a predicted coordinate point and a control point truth value, and the two losses are added according to weights to serve as an overall loss function for joint training. The method can effectively improve the detection speed and accuracy of irregular texts in a complex scene, and can further improve the application of character recognition in the aspects of robot navigation, automatic driving, augmented reality and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of deep learning and artificial intelligence, and particularly relates to a text localization method integrating Bezier control points. Background Art

[0002] Natural scenes contain rich visual and semantic information. If only relying on manual extraction and analysis of character information, it will consume a large amount of time and energy. Therefore, the automatic detection and recognition of text regions in natural scenes is of great significance for realizing intelligent image processing.

[0003] Traditional optical character detection and recognition often rely on manually designed text features, and have poor generalization ability and robustness, and are easily affected by illumination changes, blurring, noise, and text deformation, resulting in poor detection and recognition accuracy. In recent years, with the booming development of deep learning and artificial intelligence technologies in the field of optical character detection and recognition, deep learning models can learn image features through a large amount of training data and do not need to rely on manually designed feature extractors. They have strong adaptability to problems such as illumination changes, occlusion, and tilt, and have strong generalization ability and robustness.

[0004] However, text detection in natural scenes is still a challenging task. Especially in complex scenes such as unbalanced image ratios and irregularly curved texts, it will lead to a decrease in the detection accuracy of text regions. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a text localization method integrating Bezier control points.

[0006] The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] The present invention provides a text localization method integrating Bezier control points, including:

[0008] Obtaining an image to be detected;

[0009] Preprocessing the image to be detected to obtain a preprocessed image;

[0010] Inputting the preprocessed image into a trained text detection network to obtain a text region prediction result in the image to be detected; wherein, the trained text detection network is obtained by training an initial text detection network and the learnable weights using a sample image with an annotation file, the coordinates of Bezier curve control points extracted from the sample image, and a loss function with learnable weights; the initial text detection network includes: a feature extraction network with depthwise separable convolution, and a feature fusion network with an attention mechanism.

[0011] Advantages of the present invention compared with the prior art:

[0012] (1) Lightweight feature extraction network: By using depthwise separable convolutions, compared with traditional convolution operations, the number of parameters can be significantly reduced, effectively improving the real-time performance and response speed of the system, while reducing the complexity of the model (i.e., the text detection network), making the model more robust and better able to adapt to text regions of various different shapes.

[0013] (2) Feature fusion network integrating attention mechanism: Considering that the upsampling operation in the feature fusion network may lose deep detail information, the present invention introduces an attention mechanism to give the network a higher degree of focus and more accurate representation of the text regions in the image before the upsampling operation, improving the detection accuracy and robustness, and thus being able to adapt to text detection in complex scenarios with multi-scale changes and pose changes.

[0014] (3) Providing more refined text region bounding boxes: By integrating B-splines into the text region localization method, the present invention can significantly improve the text region detection accuracy and optimize the shape of the text region bounding box to make it more coherent and smooth, which helps to improve the recognition accuracy of the character recognition system.

[0015] (4) Improving the detection accuracy for curved and other irregular texts: In complex scenarios or under occlusion, existing text region detection methods often have problems of weak accuracy and poor robustness. The present invention combines B-splines to better fit and locate text regions of arbitrary shapes, providing stronger robustness and accuracy.

[0016] The following will further elaborate on the present invention in conjunction with the accompanying drawings. Description of the Drawings

[0017] Figure 1 is a schematic flowchart of a text localization method integrating B-spline control points provided by an embodiment of the present invention;

[0018] Figure 2 is a schematic architecture diagram of an exemplary i-th depthwise separable convolution sub-network provided by an embodiment of the present invention;

[0019] Figure 3 is a comparison diagram of annotation schematic diagrams obtained by using B-spline annotation and polygon annotation provided by an embodiment of the present invention;

[0020] Figure 4 is a schematic diagram of an exemplary fitted B-spline curve and four B-spline curve control points extracted according to the B-spline curve;

[0021] Figure 5 It is an architecture diagram of an exemplary text detection network provided by an embodiment of the present invention;

[0022] Figure 6 It is a comparison diagram of an exemplary fusion of Bezier control points and detection results without Bezier control points provided by an embodiment of the present invention;

[0023] Figure 7 It is another comparison diagram of an exemplary fusion of Bezier control points and detection results without Bezier control points provided by an embodiment of the present invention. Detailed implementation manners

[0024] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0025] Traditional text localization methods rely too much on the design of feature extractors, which include characteristics such as color, brightness, texture, and gradient of text regions in images. However, it also leads to a deterioration in the accuracy of text box generation when the image is interfered by factors such as extreme lighting and blurring. Text detection methods based on deep learning are much more accurate than traditional methods in complex scenarios. Their training process uses a large amount of data to fit the detection model to improve the generalization ability of the model in the inference stage. However, in complex scenarios such as image scale imbalance and irregularly curved text, their detection results are still not very satisfactory.

[0026] The inventors of this case found that Bezier curves can be applied to text region detection. Specifically, the text bounding boxes in the training set can be curve-fitted to obtain the control points of the cubic Bezier curve. At the same time, feature information in the image is extracted based on a lightweight feature extraction network, and multi-scale features are fused through upsampling and convolutional attention modules. After convolutional operations, a probability map of the character region and the coordinate regression of the Bezier control points are generated respectively. The region loss is calculated according to the probability map and the ground truth of the text bounding box, and the coordinate loss is calculated according to the predicted coordinate points and the ground truth of the control points. The two losses are added together according to weights as the overall loss function for joint training. In this way, the detection speed and accuracy of irregular text in complex scenarios can be effectively improved, and the application of character recognition in aspects such as robot navigation, autonomous driving, and augmented reality can be further enhanced.

[0027] Figure 1 It is a schematic flowchart of a text localization method for fusing Bezier control points provided by an embodiment of the present invention. The method includes:

[0028] S101. Obtain the image to be detected.

[0029] S102. Preprocess the image to be detected to obtain a preprocessed image.

[0030] Here, the preprocessing may include operations such as image data augmentation, image size adjustment, and image normalization. For example, when adjusting the image size, the image size can be uniformly adjusted to 512×512×3.

[0031] S103. Input the preprocessed image into the trained text detection network to obtain the prediction result of the text region in the image to be detected; among them, the trained text detection network is obtained by training the initial text detection network and the learnable weights using the sample image with annotation files, the coordinates of the Bessel curve control points extracted from the sample image, and the loss function with learnable weights; the initial text detection network includes: a feature extraction network with depthwise separable convolution, and a feature fusion network with an attention mechanism.

[0032] In some embodiments, the above S102 can be implemented through the following steps:

[0033] S1021. After inputting the preprocessed image into the trained text detection network, use the feature extraction network to extract features from the preprocessed image to obtain multiple feature maps of different scales.

[0034] Here, in the order from largest to smallest scale, the multiple feature maps of different scales may include: the initial feature map, the first feature map, the second feature map, and the third feature map.

[0035] S1022. Use the feature fusion network to perform feature fusion on the multiple feature maps of different scales to obtain a fused feature map.

[0036] S1023. Process the fused feature map using the convolution and regression network to obtain the Bessel control point coordinates of the upper and lower boundaries of each text region in the preprocessed image, and the coarse-grained text region probability map.

[0037] S1024. According to the coarse-grained text region probability map, the preset threshold, and the Bessel control point coordinates of the upper and lower boundaries of each text region, obtain the prediction result of the text region in the image to be detected.

[0038] In some embodiments, the above S1021 can be implemented through the following steps:

[0039] S1. Use the first convolutional sub-network to perform a convolutional operation on the preprocessed image to obtain the initial feature map.

[0040] Here, the first convolutional sub-network can be two cascaded convolutional layers, where the stride of the first convolutional layer is 2. For example, when the size of the preprocessed image is 512×512×3, the size of the initial feature map is 256×256×64.

[0041] S2. Process the initial feature map using the first depthwise separable convolutional sub-network to obtain a first feature map.

[0042] S3. Process the first feature map using the second depthwise separable convolutional sub-network to obtain a second feature map.

[0043] S4. Process the second feature map using the third depthwise separable convolutional sub-network to obtain a third feature map; wherein, there is a residual connection between the output ends of the first convolutional sub-network, the first depthwise separable convolutional sub-network, the second depthwise separable convolutional sub-network, and the third depthwise separable convolutional sub-network; the first depthwise separable convolutional sub-network, the second depthwise separable convolutional sub-network, and the third depthwise separable convolutional sub-network are sub-networks with the same structure.

[0044] Here, each depthwise separable convolutional sub-network includes: a convolutional layer, two depthwise separable convolutional layers, and a max-pooling layer; wherein, the two depthwise separable convolutional layers and the max-pooling layer are connected in series in sequence; the input of the convolutional layer is the same as the input of the first depthwise separable convolutional layer, and the output is added to the output of the max-pooling layer. Exemplarily, Figure 2 is the architecture diagram of the i-th depthwise separable convolutional sub-network, where the value of i ranges from 1 to 3. Here, Conv represents the convolutional layer, SepConv represents the depthwise separable convolutional layer, and Max-pool represents the max-pooling layer, represents addition; as Figure 2 shown, when the size of the feature map input to the first SepConv and Conv is then, the size of the output feature map is where H, W, and C are the sizes of the original image corresponding to the feature map. Based on this, the above S2 can be implemented through the following steps: perform convolutional processing on the initial feature map using Conv to obtain a first sub-feature map; use the first SepConv to perform channel convolution and pointwise convolution on the initial feature map to obtain a second sub-feature map; use the second SepConv to perform channel convolution and pointwise convolution on the second sub-feature map to obtain a third sub-feature map; use Max-pool to perform max-pooling on the third sub-feature map to obtain a fourth sub-feature map; use to add the first sub-feature map and the fourth sub-feature map to obtain a first feature map.

[0045] Specifically, the processing process of each SepConv for the input feature map is as follows:

[0046] The first step: Use the depthwise separable convolution kernel to perform convolution on the input feature map The feature map output by the depthwise separable convolution is This operation can be expressed as: W and H are respectively the width and height of the feature map output by the depthwise separable convolution, c is the number of channels of the feature map output by the depthwise separable convolution, n is the size of the depthwise separable convolution kernel, k and l are the position coordinates on the feature map output by the depthwise separable convolution, i and j are the convolution kernel indices, m is the channel position index, Y′ k,l,m is a specific value of, that is, it represents the value at the position where the number of channels is m, and the width and height are k and l respectively, K i,j,m is a specific value of.

[0047] Step 2: Perform pointwise convolution on the feature map output by the depthwise separable convolution and calculate the linear combination of the depthwise convolution to obtain the output feature map where c′ and c are the number of input and output channels respectively, Y k,l,h is a specific value of, that is, it represents the value at the position where the number of channels is h, and the width and height are k and l respectively, K′ 1,1,m,h is the convolution kernel of the pointwise convolution, and m and h respectively correspond to the indices of the number of input channels and the number of output channels, which can realize mapping the number of input channels from m to h.

[0048] Compared with the number of parameters of the traditional convolution (n×n×c×c′), the present invention adopts the depthwise separable convolution, so that the number of parameters of the whole operation can be reduced to (n×n×c + c×c′). Since the scene text image usually has a high resolution and the size of the text area is small, adopting the depthwise separable convolution can greatly reduce the number of parameters and improve the computing efficiency. When processing a large number of images, it can detect the text area faster, improve the real-time performance and response speed of the system. At the same time, the depthwise separable convolution can reduce the complexity of the model (i.e., the text detection network), reduce the possibility of overfitting, make the model have stronger robustness, and be able to better adapt to various text areas.

[0049] The processing principles of the above S3 and S4 are the same as those of the above S2, and will not be elaborated here.

[0050] In some embodiments, the above S1022 can be implemented through the following steps:

[0051] S11: Process the third feature map using a convolutional attention network to obtain the first convolutional attention feature map.

[0052] Specifically, when the third feature map is the feature map F (it should be noted that the feature map F here and the above feature map represent different feature maps respectively), the processing process of the convolutional attention network for the feature map F is as follows:

[0053] a) Perform max pooling and average pooling on the feature map F respectively, then merge them through a multi-layer perceptron, and finally activate them through the activation function σ to obtain the channel weights Specifically, the channel weights of the feature map F are as follows: where MLP represents a multi-layer perceptron, AvgPool represents average pooling, and MaxPool represents max pooling, and parameter sharing.

[0054] b) Multiply M C (F) by the feature map F to obtain the feature map F'.

[0055] c) Use max pooling and average pooling to compress the feature map F', and then, through convolution and activation operations, obtain the spatial weights The spatial weights M S (F) of the feature map F are as follows: where f7×7 is a 7×7 convolutional layer.

[0056] d) Multiply M S (F) by the feature map F' to obtain the feature map F'', and the feature map F'' is the first convolutional attention feature map.

[0057] S12. Upsample the first convolutional attention feature map and splice it with the second feature map, and use the second convolutional sub-network to reduce the number of channels of the spliced feature map to obtain the first low-channel feature map.

[0058] Here, the second convolutional sub-network can be two cascaded convolutional layers to further reduce the number of channels of the feature map.

[0059] S13. Process the first low-channel feature map using the convolutional attention network to obtain the second convolutional attention feature map.

[0060] Here, the processing principle of S13 is the same as that of S11 above.

[0061] S14. Upsample the second convolutional attention feature map and splice it with the first feature map, and use the second convolutional sub-network to reduce the number of channels of the spliced feature map to obtain the second low-channel feature map.

[0062] S15. Process the second low-channel feature map using the convolutional attention network to obtain the third convolutional attention feature map.

[0063] Here, the processing principle of S15 is the same as that of S11 above.

[0064] S16. Upsample the third convolutional attention feature map and concatenate it with the initial feature map, and use the second convolutional sub-network to reduce the number of channels of the concatenated feature map to obtain a fused feature map.

[0065] For example, the size of the fused feature map is 256×256×128.

[0066] In the present invention, the size of the feature map is reduced by the max-pooling operation. By introducing the attention mechanism, shallow information and deep information can be fused, thereby improving the network's ability to extract key information. Moreover, through the concatenation fusion operation, a multi-scale fusion feature map with higher richness can be obtained. By introducing the attention mechanism in text region detection, the network can pay more attention to the key features of the text region, improving the accuracy and robustness of detection, and better solving the problems such as inaccurate text region detection caused by scale changes and pose changes of text in images.

[0067] In some embodiments, the above S1023 can be implemented through the following steps:

[0068] S21. Perform convolutional operations on the fused feature map respectively to obtain a coarse-grained text region probability map and a Bessel control point coordinate feature map correspondingly.

[0069] S22. Use a fully connected layer to perform regression processing on the Bessel control point coordinate feature map to obtain the upper and lower boundary Bessel control point coordinates of each text region in the preprocessed image.

[0070] In some embodiments, the above S1024 can be implemented through the following steps:

[0071] S31. Generate a Bessel curve based on the upper and lower boundary Bessel control points of each text region as the text bounding box of the text region.

[0072] Here, the text bounding box is a text bounding frame.

[0073] S32. Perform weighted summation on the text bounding box of the text region and the coarse-grained text region probability map to obtain a weighted summation value.

[0074] S33. When the weighted summation value is greater than or equal to a preset threshold, it is considered that the region framed by the text bounding box is a text region.

[0075] In some embodiments, before the above S103, the method further includes:

[0076] S001. Obtain a training set composed of multiple sample images; each sample image corresponds to an annotation file; in the annotation file of each sample image, the annotation points of each text region of the sample image are included.

[0077] Here, the acquisition method of the annotation file corresponding to each sample image can be obtained by using the Bezier annotation method, or other annotation methods, such as the polygon annotation method. For example, Figure 3 is a comparison diagram of the annotation schematic diagrams obtained by using the Bezier annotation method and the polygon annotation method. Among them, Figure 3 Figure (a) in is the annotation schematic diagram obtained by using the Bezier annotation method, Figure 3 Figure (b) in is the annotation schematic diagram obtained by using the polygon annotation method.

[0078] S002. When training the text detection network obtained from the p-th training, select the training samples for the (p + 1)-th time from the training set; p is an integer greater than or equal to 0; when p is 0, the text detection network obtained from the p-th training is the initial text detection network.

[0079] Here, the training samples for the (p + 1)-th time can be one or more, and there is no limit to this.

[0080] S003. According to the annotation points in the annotation file of each training sample, extract the coordinates of the Bezier curve control points of the upper and lower boundaries of each text region of the training sample.

[0081] S004. Input the training samples for the (p + 1)-th time into the text detection network obtained from the p-th training, and predict the coordinates of the Bezier curve control points of the upper and lower boundaries of each text region of each training sample, and obtain the coarse-grained text region probability map of the training sample.

[0082] S005. According to the coordinates of the Bezier curve control points of the upper and lower boundaries extracted and predicted for all text regions in the training samples for the (p + 1)-th time, the coarse-grained text region probability map of the training samples for the (p + 1)-th time, and the learnable weights obtained from the p-th training, determine the loss for the (p + 1)-th time; when p is 0, the learnable weights obtained from the p-th training are preset values.

[0083] S006. According to the loss for the (p + 1)-th time, adjust the network parameters of the text detection network obtained from the p-th training and the learnable weights obtained from the p-th training, respectively, to obtain the text detection network obtained from the (p + 1)-th training and the learnable weights obtained from the (p + 1)-th training, and perform iterative training in this way until the trained text detection network and the trained weights are obtained.

[0084] Specifically, during training, it can be determined whether to stop training according to the loss obtained from each training. For example, when the loss continuously decreases and tends to be stable, training can be stopped, and the text detection network obtained from the last training is used as the trained text detection network.

[0085] In some embodiments, for the above S003, when the annotation file of the sample image is obtained by using the Bezier annotation method, for each training sample, the coordinates of the Bezier curve control points of the upper and lower boundaries of each text region of the training sample can be directly extracted from the annotation file of the training sample.

[0086] In some embodiments, for the above S003, when the annotation file of the sample image is obtained by using other annotation methods, the coordinates of the Bezier curve control points of the upper and lower boundaries of each text region of each training sample can be extracted through steps S41 to S43.

[0087] S41. According to the annotation file of each training sample, determine the number of annotation points of each text region of the training sample.

[0088] Here, for each text region in each training sample, the number of annotation points of the text region included in the annotation file of the training sample can be counted.

[0089] S42. When the number is greater than a preset value, according to the annotation points of the text region, use the least squares method to respectively fit the coordinates of the Bezier curve control points of the upper boundary of the text region and the coordinates of the Bezier curve control points of the lower boundary of the text region.

[0090] Exemplarily, the preset value can be 4. Thus, when the number of annotation points of the text region is greater than 4, according to the number of annotation points of the text region, use the least squares method to respectively fit the coordinates of the Bezier curve control points of the upper boundary of the text region and the coordinates of the Bezier curve control points of the lower boundary of the text region.

[0091] The specific calculation formula of the Bezier curve is as follows:

[0092]

[0093]

[0094] Among them, n in this formula is the order of the Bezier curve, b i is the i-th control point, B i,n (t) is the Bernstein polynomial, where is a binomial coefficient. During the process where the parameter t varies from 0 to 1, a complete Bessel curve is formed. Through comprehensive analysis of any existing scene text dataset with arbitrary shapes and annotation information, it is found that when n = 3 here, that is, when using a third-order Bessel curve, the computational complexity and fitting effect can be better balanced. For a specific text bounding box, the least squares fitting method is used for Bessel curve regression on the upper and lower boundaries of its text area in the row direction, and the short sides are directly connected. Therefore, for each polygon annotation area (i.e., text area), it can be transformed into the coordinates of 8 fixed control points, that is, the coordinates of four Bessel curve control points on the upper boundary of the text area are extracted respectively, and the coordinates of four Bessel curve control points on the lower boundary are extracted respectively.

[0095] Specifically, the formula for fitting the coordinates of four Bessel curve control points on the upper boundary of a text area using the least squares method is as follows:

[0096]

[0097] where m is the number of annotation points in the text area, t i′ is the ratio of the cumulative length of the i'-th annotation point in the annotation points of the text area to the total length of the polyline (the total length of the polyline is the total length of the line segments formed by sequentially connecting the m annotation points of the text area), and the value of i' ranges from 0 to m. is the coordinate of the j-th Bessel curve control point, and the value of j ranges from 0 to 3; where and are the coordinates of the Bessel curve control points at the two end positions in the upper boundary respectively. is the coordinate of the i'-th point on the fitted Bessel curve. For example, Figure 4 is a schematic diagram of a fitted Bessel curve and four Bessel curve control points extracted according to this Bessel curve. It should be noted that the principle of fitting the coordinates of four Bessel curve control points on the lower boundary of the text area using the least squares method is the same as the above principle, and will not be elaborated here.

[0098] S43. When the quantity is equal to the preset value, along the long side direction of the text area, the text area is divided into three equal parts, and the annotation points at the two end positions in the upper boundary of the text area, as well as the equally divided points in the upper boundary, are used as the coordinates of the Bessel curve control points of the upper boundary of the text area, and the annotation points at the two end positions in the lower boundary of the text area, as well as the equally divided points in the lower boundary, are used as the coordinates of the Bessel curve control points of the lower boundary of the text area.

[0099] For example, when the number of marked points in a text region is 4 (i.e., the text region is a rectangular region composed of four marked points), the text region can be trisected along the long side direction of the text region. In this way, two Bezier curve control points are added to the upper side of the rectangular region, and two Bezier curve control points are also added to the lower side of the rectangular region. Then, the two marked points that form the upper side of the rectangular region and the two Bezier curve control points added to the upper side of the rectangular region are used as the Bezier curve control points of the upper boundary of the text region; the two marked points that form the lower side of the rectangular region and the two Bezier curve control points added to the lower side of the rectangular region are used as the Bezier curve control points of the lower boundary of the text region.

[0100] In some embodiments, the above S005 can be implemented through the following steps:

[0101] S51. For each text region of each training sample, calculate the sum of the Euclidean distances between the coordinates of the extracted Bezier curve control points of the upper and lower boundaries of the text region and the coordinates of the predicted Bezier curve control points of the upper and lower boundaries, and perform an exponential transformation on the sum of the Euclidean distances of the text region to obtain the coordinate loss value of the text region.

[0102] For example, for a text region in a training sample, the calculation formula for the sum of the corresponding Euclidean distances is: Here, n represents the sum of the number of Bezier curve control points of the upper and lower boundaries of the text region, p i represents the predicted i-th Bezier curve control point, represents the extracted i-th Bezier curve control point corresponding to p i . The formula for performing an exponential transformation on the sum of the Euclidean distances of the text region is:

[0103] S52. Take the sum of the coordinate loss values of all text regions in the (p + 1)-th training sample as the coordinate loss of the (p + 1)-th time.

[0104] S53. For each text region in each training sample, calculate the GIoU between the text region and the coarse-grained text region probability map of the training sample, and use it as the region loss value of the text region.

[0105] Specifically, the calculation formula for GIoU is: A c is the area of the minimum bounding rectangle between the text region and the coarse-grained text region probability map, and U is the merged part of the text region and the coarse-grained text region probability map.

[0106] When the traditional IoU is used as the loss function, if there is no overlap between two regions, the IoU is 0, which cannot reflect the similarity between the two regions. In the present invention, by adopting GIoU, the minimum bounding rectangle of the two regions can be introduced, and its area can well measure the similarity of the non-overlapping regions between the regions.

[0107] S54. In the training samples of the (p + 1)-th time, determine the region loss of the (p + 1)-th time according to the sum of the region loss values of all text regions.

[0108] For example, when the expression of the region loss of the (p + 1)-th time is: In this formula, GIoU represents the sum of the region loss values of all text regions in the training samples of the (p + 1)-th time.

[0109] S55. Use the learnable weight obtained from the p-th training to perform weighted summation on the coordinate loss of the (p + 1)-th time and the region loss of the (p + 1)-th time to obtain the loss of the (p + 1)-th time.

[0110] For example, the calculation formula of the loss of the (p + 1)-th time is: In this formula, represents the coordinate loss of the (p + 1)-th time, and β represents the learnable weight obtained from the p-th training.

[0111] Exemplarily, Figure 5 is an architecture diagram of a text detection network, as Figure 5 shown, the text detection network sequentially includes: the above-mentioned feature extraction network composed of a convolutional layer Conv with a stride of 2, a convolutional layer Conv, Stage1, Stage2, and Stage3, the above-mentioned feature fusion network composed of CBAM, an upsampling layer, and a splicing layer, and the above-mentioned convolutional and regression network composed of a convolutional layer Conv and a fully connected layer FC; wherein, Stage1, Stage2, and Stage3 respectively represent the above-mentioned first depthwise separable convolutional sub-network, the above-mentioned second depthwise separable convolutional sub-network, and the above-mentioned third depthwise separable convolutional sub-network, and the convolutional layer Conv with a stride of 2 and the convolutional layer Conv constitute the above-mentioned first convolutional sub-network. As Figure 4 shown, when Input inputs into the text detection network, after passing through the feature extraction network and the feature fusion network in sequence, a fused feature F is obtained. After the fused feature F passes through the convolutional and regression network, a Bessel control point coordinate feature map Control point and a coarse-grained text region probability map score map are obtained. Then, according to Control point and score map, text region prediction is performed, and the final prediction result Fine grained text area is output; when training, the loss can be calculated respectively according to Control point and score map and losses for backpropagation.

[0112] To intuitively illustrate the effect of the text localization method proposed in this case, Figure 6 and Figure 7 , Figure 6 and Figure 7 are both comparisons of the detection results integrating Bessel control points and those without Bessel control points; among them, Figure 6 and Figure 7 the (a) figures in both are the detection results integrating Bessel control points, Figure 6 and Figure 7 the (b) figures in both are the detection results without Bessel control points. Obviously, the detection results integrating Bessel control points are more accurate.

[0113] A text localization method integrating Bessel control points provided by the present invention can better capture the curve features of the text and adapt to the shape changes of the text area when dealing with irregular or curved texts, with strong flexibility and adaptability. It effectively improves the problems such as weak text area localization accuracy and poor robustness in complex scenarios, and improves the generalization ability of the system and the accuracy of text recognition. It enhances the application level of character detection and recognition in aspects such as robot navigation, autonomous driving, and augmented reality.

[0114] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0115] In the description of this specification, the reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0116] In the specification, the term "including" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. Certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0117] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A text localization method integrating Bezier control points, characterized in that, Including: Obtain the image to be detected; Preprocess the image to be detected to obtain a preprocessed image; Input the preprocessed image into the trained text detection network to obtain the text region prediction result in the image to be detected; wherein, the trained text detection network is trained by using a sample image with an annotation file, the coordinates of the Bessel curve control points extracted from the sample image, and a loss function with learnable weights for the initial text detection network and the learnable weights; the initial text detection network includes: a feature extraction network with depthwise separable convolutions, and a feature fusion network with an attention mechanism.

2. The text localization method integrating Bessel control points according to claim 1, wherein The step of inputting the preprocessed image into the trained text detection network to obtain the text region prediction result in the image to be detected includes: After inputting the preprocessed image into the trained text detection network, use the feature extraction network to extract features from the preprocessed image to obtain multiple feature maps of different scales; Use the feature fusion network to perform feature fusion on the multiple feature maps of different scales to obtain a fused feature map; Use a convolution and regression network to process the fused feature map to obtain the upper and lower boundary Bessel control point coordinates of each text region in the preprocessed image, and a coarse-grained text region probability map; According to the coarse-grained text region probability map, a preset threshold, and the upper and lower boundary Bessel control point coordinates of each text region, obtain the text region prediction result in the image to be detected.

3. The text positioning method for fusing Bessel control points according to claim 2, characterized in that, The multiple feature maps of different scales include: an initial feature map, a first feature map, a second feature map, and a third feature map; the step of using the feature extraction network to extract features from the preprocessed image to obtain multiple feature maps of different scales includes: Use a first convolutional sub-network to perform a convolution operation on the preprocessed image to obtain the initial feature map; Use a first depthwise separable convolutional sub-network to process the initial feature map to obtain the first feature map; Use a second depthwise separable convolutional sub-network to process the first feature map to obtain the second feature map; Use a third depthwise separable convolutional sub-network to process the second feature map to obtain the third feature map; Wherein, there are residual connections between the output ends of the first convolutional sub-network, the first depthwise separable convolutional sub-network, the second depthwise separable convolutional sub-network, and the third depthwise separable convolutional sub-network; the first depthwise separable convolutional sub-network, the second depthwise separable convolutional sub-network, and the third depthwise separable convolutional sub-network are sub-networks with the same structure.

4. The text positioning method for fusing Bessel control points according to claim 3, characterized in that The step of using the first depthwise separable convolutional sub-network to process the initial feature map to obtain the first feature map includes: Use a convolutional layer to perform a convolution operation on the initial feature map to obtain a first sub-feature map; Use a first depthwise separable convolutional layer to perform channel convolution and pointwise convolution on the initial feature map to obtain a second sub-feature map; Use a second depthwise separable convolutional layer to perform channel convolution and pointwise convolution on the second sub-feature map to obtain a third sub-feature map; Perform max pooling on the third sub-feature map using a max pooling layer to obtain a fourth sub-feature map; Add the first sub-feature map and the fourth sub-feature map to obtain the first feature map.

5. The text localization method for fusing Bessel control points according to claim 2, wherein In the order from largest to smallest scale, the multiple feature maps of different scales include: an initial feature map, a first feature map, a second feature map, and a third feature map; the using the feature fusion network to perform feature fusion on the multiple feature maps of different scales to obtain a fused feature map includes: Process the third feature map using a convolutional attention network to obtain a first convolutional attention feature map; Upsample the first convolutional attention feature map and splice it with the second feature map, and use a second convolutional sub-network to reduce the number of channels of the spliced feature map to obtain a first low-channel feature map; Process the first low-channel feature map using the convolutional attention network to obtain a second convolutional attention feature map; Upsample the second convolutional attention feature map and splice it with the first feature map, and use the second convolutional sub-network to reduce the number of channels of the spliced feature map to obtain a second low-channel feature map; Process the second low-channel feature map using the convolutional attention network to obtain a third convolutional attention feature map; Upsample the third convolutional attention feature map and splice it with the initial feature map, and use the second convolutional sub-network to reduce the number of channels of the spliced feature map to obtain the fused feature map.

6. The text positioning method for fusing Bessel control points according to claim 5, wherein The using the convolutional attention network to process the third feature map to obtain a first convolutional attention feature map includes: Use a max pooling layer and an average pooling layer to perform max pooling and average pooling on the feature map F respectively, and then use a multi-layer perceptron to merge the feature maps after max pooling and average pooling to obtain a merged feature map; the feature map F is the third feature map; The merged feature map is activated by the activation function σ to obtain the channel weight M C (F), and the channel weight M C (F) is multiplied by the feature map F to obtain the feature map F'; Use a max pooling layer and an average pooling layer to perform max pooling and average pooling on the feature map F' respectively, and then use a convolutional layer to perform convolution on the feature maps after max pooling and average pooling to obtain a convolutional feature map; The convolution feature map is activated using the activation function σ to obtain the spatial weight M S (F), multiply the spatial weight M S (F) by the feature map F' to obtain the feature map F”, and use the feature map F” as the first convolutional attention feature map.

7. The text positioning method for fusing Bessel control points according to claim 2, characterized in that The obtaining the text region prediction result in the to-be-detected image according to the coarse-grained text region probability map, a preset threshold, and the upper and lower boundary Bezier control point coordinates of each text region includes: Connect the upper and lower boundary Bezier control points of each text region to obtain a text bounding box of this text region; Perform weighted summation on the text bounding box of this text region and the coarse-grained text region probability map to obtain a weighted summation value; When the weighted summation value is greater than or equal to the preset threshold, it is considered that the region framed by this text bounding box is a text region.

8. The text location method for fusing Bessel control points according to claim 1, wherein Before obtaining the text region prediction result in the to-be-detected image by inputting the preprocessed image into the trained text detection network, the method further includes: Obtain a training set composed of multiple sample images; each sample image corresponds to an annotation file; in the annotation file of each sample image, the annotation points of each text region of this sample image are included; When training the text detection network obtained from the p-th training, select the training samples for the (p + 1)-th time from the training set; p is an integer greater than or equal to 0; when p is 0, the text detection network obtained from the p-th training is the initial text detection network; According to the labeled points in the annotation file of each training sample, extract the Bessel curve control point coordinates of the upper and lower boundaries of each text region of the training sample; Input the training samples for the (p + 1)-th time into the text detection network obtained from the p-th training, predict the Bessel curve control point coordinates of the upper and lower boundaries of each text region of each training sample, and obtain the coarse-grained text region probability map of the training sample; Determine the loss for the (p + 1)-th time according to the Bessel curve control point coordinates of the upper and lower boundaries, both extracted and predicted, of all text regions in the training samples for the (p + 1)-th time, the coarse-grained text region probability map of the training samples for the (p + 1)-th time, and the learnable weights obtained from the p-th training; when p is 0, the learnable weights obtained from the p-th training are preset values; Adjust the network parameters of the text detection network obtained from the p-th training and the learnable weights obtained from the p-th training according to the loss for the (p + 1)-th time, to obtain the text detection network obtained from the (p + 1)-th training, and iterate the training in this way until the trained text detection network is obtained.

9. The text positioning method for fusing Bessel control points according to claim 8, characterized in that, The step of extracting the Bessel curve control point coordinates of the upper and lower boundaries of each text region of the training sample according to the labeled points in the annotation file of each training sample includes: According to the annotation file of each training sample, determine the number of labeled points of each text region of the training sample; When the number is greater than the preset value, respectively fit the Bessel curve control point coordinates of the upper boundary and the Bessel curve control point coordinates of the lower boundary of the text region by using the least squares method according to the labeled points of the text region; When the number is equal to the preset value, divide the text region into three equal parts along the long side direction of the text region, and use the labeled points located at the two end positions of the upper boundary of the text region and the equal division points in the upper boundary as the Bessel curve control point coordinates of the upper boundary of the text region, and use the labeled points located at the two end positions of the lower boundary of the text region and the equal division points in the lower boundary as the Bessel curve control point coordinates of the lower boundary of the text region.

10. The text positioning method for fusing Bessel control points according to claim 8, characterized in that, The step of determining the loss for the (p + 1)-th time according to the Bessel curve control point coordinates of the upper and lower boundaries, both extracted and predicted, of all text regions in the training samples for the (p + 1)-th time, the coarse-grained text region probability map of the training samples for the (p + 1)-th time, and the learnable weights obtained from the p-th training includes: For each text region of each training sample, calculate the sum of the Euclidean distances between the coordinates of the Bézier curve control points of the extracted upper and lower boundaries of the text region and the coordinates of the Bézier curve control points of the predicted upper and lower boundaries, and perform an exponential transformation on the sum of the Euclidean distances of the text region to obtain the coordinate loss value of the text region; Take the sum of the coordinate loss values of all text regions in the (p + 1)-th training sample as the coordinate loss in the (p + 1)-th iteration; For each text region in each training sample, calculate the GIoU between the text region and the coarse-grained text region probability map of the training sample, and take 1 - GIoU as the region loss value of the text region; Based on the sum of the region loss values of all text regions in the (p + 1)-th training sample, determine the region loss in the (p + 1)-th iteration; Use the learnable weights obtained from the p-th training to perform a weighted sum of the coordinate loss in the (p + 1)-th iteration and the region loss in the (p + 1)-th iteration to obtain the loss in the (p + 1)-th iteration.