Method and apparatus for detecting text
The corner points of the text area are determined through the probability information output from the neural network structure, and the corner points are connected to obtain a smooth detection box, which solves the problems of uneven boundaries of text area and large amounts of parameters in the prior art, and realizes efficient text detection.
Patent Information
- Application Number
- CN202210328626.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-30
AI Technical Summary
In the prior art, when detecting arbitrary shape text, it is difficult to effectively deal with the problem of unevenness in the bounding area of text, and the method based on bounding box regression is large and time-consuming.
By designing a neural network structure, input the image of the text to be detected, output the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability of each pixel point. Based on these probabilities, the label of the pixel point is determined, and the upper boundary corner point and the lower boundary corner point of the text area are determined, and these corner points are connected to obtain a smooth line text area detection box.
It realizes that text can be detected without careful classification, avoids glitch problems on the boundary of the detection box, reduces calculation parameters and time, and improves calculation speed.
Smart Images

Figure CN114663456B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method and device for detecting text. Background Art
[0002] Scene text detection is an important direction in the field of computer vision. In recent years, scene text detection technology has made great progress driven by deep learning. In the past, most scene text detection methods could only process horizontal or directional text, but there are actually many texts of arbitrary shapes in natural links, such as curved trademarks, wavy artistic characters, etc. Therefore, how to realize arbitrary shape text detection has become the current research focus of scene text detection. Common methods are divided into two categories. One is a method based on semantic segmentation, which separates the text area from the background in the image by pixel-level binary classification of the image. However, the text area boundary obtained by direct pixel-level segmentation usually has many uneven burrs, which is not conducive to the next step of text recognition. The other is a method based on bounding box regression, which regresses the corner points on the bounding box of the text area and connects the corner points into polygons to represent them. Texts of different shapes can be accurately represented by different numbers of corner points, but this method usually requires iterative update and prediction of two network structures, stop label classification and coordinate regression, which has a large number of parameters and is very time-consuming. Summary of the invention
[0003] An object of the embodiments of the present invention is to provide a method and apparatus for detecting text, which can solve or at least partially solve the above problems.
[0004] In order to achieve the above-mentioned purpose, one aspect of an embodiment of the present invention provides a method for detecting text, the method comprising: inputting an image of text to be detected into a neural network structure to obtain pixel point information of each pixel point of a plurality of pixel points of the image, wherein, for any pixel point of the plurality of pixel points, the pixel point information includes an upper probability, a lower probability, a first intermediate probability, a left probability, a right probability and a second intermediate probability of the pixel point, wherein the upper probability and the lower probability respectively represent the probabilities that the pixel point is located at the upper boundary and the lower boundary of a single-line text area in the text area to be detected in the image, the first intermediate probability represents the probability that the pixel point is located at a first intermediate position of the single-line text area other than the upper boundary and the lower boundary, the left probability and the right probability respectively represent the probabilities that the pixel point is located at the left boundary and the right boundary of the single-line text area, and the second intermediate probability The probability characterizes the probability that the pixel point is located at the second middle position of the single-line text area except the left boundary and the right boundary; for any pixel point among the multiple pixel points, based on the upper probability, the lower probability, the first middle probability, the left probability, the right probability and the second middle probability, determine the first label and the second label of the pixel point, wherein the first label indicates whether the pixel point is an upper boundary corner point, a lower boundary corner point and a first middle corner point of the single-line text area, and the second label indicates whether the pixel point is a starting corner point, a second middle corner point and an ending corner point of the single-line text area; according to the first labels and the second labels of the multiple pixel points, determine the upper boundary corner point and the lower boundary corner point in the same line of text; and connect the upper boundary corner point and the lower boundary corner point in the same line of text to obtain a line of text area detection box.
[0005] Optionally, for any pixel point among the multiple pixel points, the first label and the second label of the pixel point are determined based on the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability, including: comparing the upper probability, the lower probability and the first intermediate probability and comparing the left probability, the right probability and the second intermediate probability to determine the first maximum among the upper probability, the lower probability and the first intermediate probability and the second maximum among the left probability, the right probability and the second intermediate probability; and determining the first label based on the first maximum and determining the second label based on the second maximum.
[0006] Optionally, determining the upper boundary corner point and the lower boundary corner point in the same line of text according to the first label and the second label of the multiple pixel points includes: for any upper boundary starting corner point, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point; determining the lower boundary starting corner point that is closest to the upper boundary starting corner point; and determining the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point, wherein the upper boundary starting corner point is the pixel point indicated as the upper boundary corner point by the first label and as the starting corner point by the second label, and the lower boundary starting corner point is the pixel point indicated as the lower boundary corner point by the first label and as the starting corner point by the second label.
[0007] Optionally, for any of the upper boundary starting corner points, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point that is closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point includes: searching for the upper boundary corner point that is closest to the upper boundary starting corner point, and when the found upper boundary corner point is the second intermediate corner point, continuing to search for the upper boundary corner point that is closest to the found upper boundary corner point, and continuously looping the searching process. Until the upper boundary corner point found is the ending corner point, wherein all the upper boundary corner points found are the upper boundary corner points in the same line of text as the upper boundary starting corner point; find the lower boundary starting corner point that is closest to the upper boundary starting corner point; and based on the content of the upper boundary corner point found in the same line of text as the upper boundary starting corner point, find the lower boundary corner point in the same line of text as the found lower boundary starting corner point, so as to obtain the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point.
[0008] Optionally, for any of the pixel points, the pixel point information also includes a coordinate position, and for any of the upper boundary starting corner points, at least one of the following uses Euclidean measurement: finding the upper boundary corner point that is closest, finding the lower boundary starting corner point that is closest, and finding the lower boundary corner point that is closest.
[0009] Optionally, the output layer of the neural network structure is a convolutional layer, and the feature map output by the output layer corresponds to the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability; preferably, the output layer of the neural network structure includes a first convolutional layer and a second convolutional layer, wherein the feature map output by the first convolutional layer corresponds to the upper probability, the lower probability and the first intermediate probability, and the feature map output by the second convolutional layer corresponds to the left probability, the right probability and the second intermediate probability.
[0010] Optionally, the loss function used when the neural network structure is trained is: Wherein, L1 is the loss function of the first convolutional layer, L2 is the loss function of the second convolutional layer, f1 corresponds to the feature map output by the first convolutional layer, f2 corresponds to the feature map output by the second convolutional layer, k represents the feature dimension, and f 1ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, P ijk represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column of the feature map output by the first convolutional layer, f 2ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the second convolutional layer, Q ijk Represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column in the feature map output by the second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer respectively have H*W pixels.
[0011] Optionally, the neural network structure also satisfies the following conditions: the input image is first downsampled and then upsampled, and feature maps with the same feature dimensions in the downsampling stage and the upsampling stage are fused; and / or the hidden layer for downsampling is a resnet structure.
[0012] Correspondingly, another aspect of an embodiment of the present invention provides a device for detecting text, the device comprising: a pixel information acquisition module, used to input an image of text to be detected into a neural network structure to obtain pixel information of each pixel in the multiple pixels of the image, wherein, for any pixel in the multiple pixels, the pixel information includes an upper probability, a lower probability, a first intermediate probability, a left probability, a right probability and a second intermediate probability of the pixel, wherein the upper probability and the lower probability respectively represent the probability that the pixel is located at the upper boundary and the lower boundary of a single-line text area in the text area to be detected in the image, the first intermediate probability represents the probability that the pixel is located at a first intermediate position of the single-line text area other than the upper boundary and the lower boundary, the left probability and the right probability respectively represent the probability that the pixel is located at the left boundary and the right boundary of the single-line text area, and the second intermediate probability represents the probability that the pixel is located at the a probability of a second middle position of a single-line text area excluding the left boundary and the right boundary; a label determination module, for determining, for any pixel point among the multiple pixel points, a first label and a second label of the pixel point based on the upper probability, the lower probability, the first middle probability, the left probability, the right probability and the second middle probability, wherein the first label indicates whether the pixel point is an upper boundary corner point, a lower boundary corner point and a first middle corner point of the single-line text area, and the second label indicates whether the pixel point is a starting corner point, a second middle corner point and an ending corner point of the single-line text area; a corner point determination module, for determining, according to the first label and the second label of the multiple pixel points, the upper boundary corner point and the lower boundary corner point in the same line of text; and a line text area detection frame determination module, for connecting the upper boundary corner point and the lower boundary corner point in the same line of text to obtain a line text area detection frame.
[0013] Optionally, the label determination module determines the first label and the second label of any pixel point among the multiple pixel points based on the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability, including: comparing the upper probability, the lower probability and the first intermediate probability and comparing the left probability, the right probability and the second intermediate probability to determine the first maximum among the upper probability, the lower probability and the first intermediate probability and the second maximum among the left probability, the right probability and the second intermediate probability; and determining the first label based on the first maximum and determining the second label based on the second maximum.
[0014] Optionally, the corner point determination module determines the upper boundary corner point and the lower boundary corner point in the same line of text according to the first label and the second label of the multiple pixel points, including: for any upper boundary starting corner point, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point; determining the lower boundary starting corner point that is closest to the upper boundary starting corner point; and determining the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point, wherein the upper boundary starting corner point is the pixel point indicated as the upper boundary corner point by the first label and as the starting corner point by the second label, and the lower boundary starting corner point is the pixel point indicated as the lower boundary corner point by the first label and as the starting corner point by the second label.
[0015] Optionally, for any of the upper boundary starting corner points, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point that is closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point includes: searching for the upper boundary corner point that is closest to the upper boundary starting corner point, and when the found upper boundary corner point is the second intermediate corner point, continuing to search for the upper boundary corner point that is closest to the found upper boundary corner point, and continuously looping the searching process. Until the upper boundary corner point found is the ending corner point, wherein all the upper boundary corner points found are the upper boundary corner points in the same line of text as the upper boundary starting corner point; find the lower boundary starting corner point that is closest to the upper boundary starting corner point; and based on the content of the upper boundary corner point found in the same line of text as the upper boundary starting corner point, find the lower boundary corner point in the same line of text as the found lower boundary starting corner point, so as to obtain the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point.
[0016] Optionally, for any of the pixel points, the pixel point information also includes a coordinate position, and for any of the upper boundary starting corner points, at least one of the following uses Euclidean measurement: finding the upper boundary corner point that is closest, finding the lower boundary starting corner point that is closest, and finding the lower boundary corner point that is closest.
[0017] Optionally, the output layer of the neural network structure is a convolutional layer, and the feature map output by the output layer corresponds to the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability; preferably, the output layer of the neural network structure includes a first convolutional layer and a second convolutional layer, wherein the feature map output by the first convolutional layer corresponds to the upper probability, the lower probability and the first intermediate probability, and the feature map output by the second convolutional layer corresponds to the left probability, the right probability and the second intermediate probability.
[0018] Optionally, the loss function used when the neural network structure is trained is: Wherein, L1 is the loss function of the first convolutional layer, L2 is the loss function of the second convolutional layer, f1 corresponds to the feature map output by the first convolutional layer, f2 corresponds to the feature map output by the second convolutional layer, k represents the feature dimension, and f 1ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, P ijk represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column of the feature map output by the first convolutional layer, f 2ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the second convolutional layer, Q ijk Represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column in the feature map output by the second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer respectively have H*W pixels.
[0019] Optionally, the neural network structure also satisfies the following conditions: the input image is first downsampled and then upsampled, and feature maps with the same feature dimensions in the downsampling stage and the upsampling stage are fused; and / or the hidden layer for downsampling is a resnet structure.
[0020] In addition, another aspect of an embodiment of the present invention further provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to enable a machine to execute the above method.
[0021] In addition, another aspect of an embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the above method when executed by a processor.
[0022] Through the above technical solution, text detection is performed based on the neural network structure, and there is no need to classify the image in detail, and no burrs will appear on the boundary of the obtained line text area detection frame, which solves the problem that the text area boundary in the prior art usually has many uneven burrs. In addition, only one neural network structure is needed, and the line text area detection frame can be obtained according to the upper probability, lower probability, first middle probability, left probability, right probability and second middle probability obtained by the neural network structure to detect text. Compared with the use of two network structures in the prior art, the parameters that need to be calculated are reduced, the amount of calculation is reduced, time is saved, time-consuming phenomena are alleviated, and the calculation speed is improved.
[0023] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0025] Figure 1 is a flow chart of a method for detecting text provided by an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a line text area detection frame provided by another embodiment of the present invention;
[0027] Figure 3 is a schematic diagram of a neural network structure provided by another embodiment of the present invention; and
[0028] Figure 4 It is a structural block diagram of a device for detecting text provided by another embodiment of the present invention.
[0029] Description of Reference Numerals
[0030] 1 Pixel information acquisition module 2 Label determination module
[0031] 3 Corner point determination module 4 Line text area detection box determination module
[0032] 5 Single-line text area 6 Top border
[0033] 7 bottom border 8 left border
[0034] 9 Right border DETAILED DESCRIPTION
[0035] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.
[0036] In the present invention, unless otherwise specified, the directional words used, such as "up, down, left, right", generally refer to divisions based on the center point of the image. The left boundary refers to the boundary on the left side of the center point, the right boundary refers to the boundary on the right side of the center point, the upper boundary refers to the boundary above the center point, and the lower boundary refers to the boundary below the center point.
[0037] One aspect of an embodiment of the present invention provides a method for detecting text.
[0038] Figure 1 FIG. 1 is a flowchart of a method for detecting text provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following contents.
[0039] In step S10, the image of the text to be detected is input into the neural network structure to obtain pixel information of each pixel in the multiple pixels of the image, wherein for any pixel in the multiple pixels, the pixel information includes the upper probability, lower probability, first intermediate probability, left probability, right probability and second intermediate probability of the pixel, wherein the upper probability and the lower probability respectively represent the probability that the pixel is located at the upper boundary and the lower boundary of the single-line text area in the text area to be detected in the image, the first intermediate probability represents the probability that the pixel is located at the first intermediate position of the single-line text area except the upper boundary and the lower boundary, the left probability and the right probability respectively represent the probability that the pixel is located at the left boundary and the right boundary of the single-line text area, and the second intermediate probability represents the probability that the pixel is located at the second intermediate position of the single-line text area except the left boundary and the right boundary. For example, if the image is H*W, it means that the image includes H*W pixels; in this step, the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability of each pixel in the H*W pixels are determined. Specifically, in the content output by the neural network structure, the image is divided into multiple pixels and includes pixel information of each pixel.
[0040] In step S11, for any pixel point among the multiple pixels, the first label and the second label of the pixel point are determined based on the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability, wherein the first label indicates whether the pixel point is the upper boundary corner point, the lower boundary corner point and the first intermediate corner point of the single-line text area, and the second label indicates whether the pixel point is the starting corner point, the second intermediate corner point and the ending corner point of the single-line text area. Among them, when the pixel point is located at the upper boundary, it is an upper boundary corner point, and the first label indicates that the pixel point is an upper boundary corner point; when the pixel point is located at the lower boundary, it is a lower boundary corner point, and the first label indicates that the pixel point is a lower boundary corner point; when the pixel point is located at the first middle position, it is a first middle corner point, and the first label indicates that the pixel point is the first middle corner point; when the pixel point is located at the left boundary, it is a starting corner point, and the second label indicates that the pixel point is a starting corner point; when the pixel point is located at the right boundary, it is an ending corner point, and the second label indicates that the pixel point is an ending corner point; when the pixel point is located at the second middle position, it is a second middle corner point, and the second label indicates that the pixel point is a second middle corner point. In addition, if Figure 2As shown, the text area to be detected includes multiple single-line text areas 5, and any single-line text area 5 includes an upper boundary 6, a lower boundary 7, a left boundary 8, and a right boundary 9. When the pixel point is located at the upper boundary 6, it is an upper boundary corner point; when the pixel point is located at the lower boundary 7, it is a lower boundary corner point; when the pixel point is located at the first intermediate position, it is a first intermediate corner point, wherein the first intermediate position is a position in the single-line text area 5 excluding the upper boundary 6 and the lower boundary 7; when the pixel point is located at the left boundary 8, it is a starting corner point; when the pixel point is located at the right boundary 9, it is an ending corner point; when the pixel point is located at the second intermediate position, it is a second intermediate corner point, wherein the second intermediate position is a position in the single-line text area 5 excluding the left boundary 8 and the right boundary 9. Optionally, the first label and the second label can be determined according to the following content. The upper probability, the lower probability, and the first intermediate probability are compared, and the left probability, the right probability, and the second intermediate probability are compared to determine the first largest one among the upper probability, the lower probability, and the first intermediate probability, and the second largest one among the left probability, the right probability, and the second intermediate probability. The first label is determined based on the first maximum and the second label is determined based on the second maximum. For example, for a certain pixel point, among the upper probability, the lower probability and the first intermediate probability, the first maximum is the upper probability, and the upper probability represents the probability that the pixel point is located at the upper boundary of the single-line text area. If the first maximum is the upper probability, it means that the probability that the pixel point is located at the upper boundary is the largest, then the pixel point is the upper boundary corner point, and the first label indicates that the pixel point is the upper boundary corner point. For another example, for a certain pixel point, among the left probability, the right probability and the second intermediate probability, the second maximum is the right probability, and the right probability represents the probability that the pixel point is located at the right boundary of the single-line text area. If the second maximum is the right probability, it means that the probability that the pixel point is located at the right boundary is the largest, then the pixel point is the termination corner point, and the second label indicates that the pixel point is the termination corner point. In addition, the neural network structure has been trained to recognize blank parts during training. Among them, for the training of the blank part of the page, there is no connection between the pixels, and the pixels are only arranged linearly. In fact, the image is locally connected in space, and each pixel in it jointly determines the characteristics of this local area. Therefore, the image should be divided into small blocks, and each small block is jointly represented as a feature, and then input into the neuron. Each neuron does not need to sense the entire image, but only needs to sense the local features. Whether it is the text part or the non-text part in the image, it is recognized through the receptive field, and then these different local neurons that are sensed are combined at a higher level to obtain global information.
[0041] In step S12, the upper boundary corner points and the lower boundary corner points in the same line of text are determined according to the first labels and the second labels of the plurality of pixel points.
[0042] In step S13, the upper boundary corner points and the lower boundary corner points in the same line of text are connected to obtain a line of text area detection box.
[0043] Through the above technical solution, text detection is performed based on the neural network structure, and there is no need to classify the image in detail, and no burrs will appear on the boundary of the obtained line text area detection frame, which solves the problem that the text area boundary in the prior art usually has many uneven burrs. In addition, only one neural network structure is needed, and the line text area detection frame can be obtained according to the upper probability, lower probability, first middle probability, left probability, right probability and second middle probability obtained by the neural network structure to detect text. Compared with the use of two network structures in the prior art, the parameters that need to be calculated are reduced, the amount of calculation is reduced, time is saved, time-consuming phenomena are alleviated, and the calculation speed is improved.
[0044] Optionally, in an embodiment of the present invention, determining the upper boundary corner point and the lower boundary corner point in the same line of text based on the first label and the second label of multiple pixel points may include the following contents. For any upper boundary starting corner point, determine the upper boundary corner point in a line of text with the upper boundary starting corner point; determine the lower boundary starting corner point that is closest to the upper boundary starting corner point; determine the lower boundary corner point in the same line of text with the lower boundary starting corner point that is closest to the upper boundary starting corner point, wherein the upper boundary starting corner point is the pixel point indicated as the upper boundary corner point by the first label and as the starting corner point by the second label, and the lower boundary starting corner point is the pixel point indicated as the lower boundary corner point by the first label and as the starting corner point by the second label. For example, for any upper boundary starting corner point, after determining the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point that is closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point, the upper boundary starting corner point, the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point that is closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point are connected in sequence to obtain a line text area detection box. Among them, the upper boundary starting corner point and the upper corner point in the same line of text as the upper boundary starting corner point are connected to obtain the upper boundary of the line text area detection box, the lower boundary starting corner point closest to the upper boundary starting corner point and the lower boundary corner point in the same line of text as the lower boundary starting corner point closest to the upper boundary starting corner point are connected to obtain the lower boundary of the line text area detection box, the upper boundary starting corner point and the lower boundary starting corner point closest to it are connected to obtain the left boundary of the line text area detection box, the terminating corner point in the upper corner point in a line of text with the upper boundary starting corner point and the terminating corner point in the lower boundary corner point in a line of text with the lower boundary starting corner point closest to the upper boundary starting corner point are connected to obtain the right boundary of the line text area detection box.
[0045] Optionally, in an embodiment of the present invention, for any upper boundary starting corner point, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point closest to the upper boundary starting corner point can be performed according to the following content. Find the upper boundary corner point closest to the upper boundary starting corner point. For example, knowing the coordinate position of the pixel point, the Euclidean metric can be used to find the upper boundary corner point closest to the upper boundary starting corner point. In the case where the upper boundary corner point closest to the upper boundary starting corner point is the termination corner point, the search process is terminated, and the upper boundary corner point found is the upper boundary corner point in the same line of text as the upper boundary starting corner point. In the case where the upper boundary corner point closest to the upper boundary starting corner point is the second intermediate corner point, continue to find the upper boundary corner point closest to the found upper boundary corner point, and continuously loop the search process until the found upper boundary corner point is the termination corner point. Specifically, when the upper boundary corner point found in the continued search process is the second middle corner point, the upper boundary corner point closest to the upper boundary corner point found in the last continued search process is searched again. As long as the upper boundary corner point found in the current search is not the termination corner point, the upper boundary corner point closest to the upper boundary corner point found in the current search is searched again until the upper boundary corner point found in the current search is the termination corner point. In this way, the search process is continuously looped until the upper boundary corner point found is the termination corner point. Among them, all the upper boundary corner points found are the upper boundary corner points in the same line of text as the upper boundary starting corner point. Optionally, in an embodiment of the present invention, all the upper boundary corner points with the closest search distance can be searched using the Euclidean metric. Find the lower boundary starting corner point closest to the upper boundary starting corner point. For example, the Euclidean metric can be used for searching. According to the content of searching for the upper boundary corner point in the same line of text as the upper boundary starting corner point, searching for the lower boundary corner point in the same line of text as the found lower boundary starting corner point, so as to obtain the lower boundary corner point in the same line of text as the lower boundary starting corner point closest to the upper boundary starting corner point. Optionally, in the embodiment of the present invention, all the lower boundary corner points with the closest distance can be searched using the Euclidean metric. Optionally, in the embodiment of the present invention, a coordinate system is established based on an image, and the coordinate position of a pixel point is the coordinate position of the pixel point in the image. For example, an image includes H*W pixels, that is, the image includes H rows of pixels and W columns of pixels, and a coordinate system is established based on the image, and a row and a column determine a specific pixel point. Therefore, in the embodiment of the present invention, a row is equivalent to the horizontal coordinate of a pixel point, and a column is equivalent to the vertical coordinate of a pixel point, and a row and a column represent the coordinate position of a pixel point. Optionally, the output layer of the preset convolutional neural network can be set as a convolution layer, and the coordinate position of a pixel point is obtained through a feature map output by the convolution layer.For example, the feature map output by the convolution layer is an H*W*3 feature map, which includes H rows of pixels and W columns of pixels. The rows and columns determine a specific pixel. Knowing the row and column corresponding to a pixel can tell us the coordinate position of the pixel.
[0046] Optionally, in an embodiment of the present invention, the output layer of the neural network structure is a convolutional layer, and the feature map output by the output layer corresponds to the upper probability, the lower probability, the first middle probability, the left probability, the right probability and the second middle probability.
[0047] Optionally, in an embodiment of the present invention, the output layer of the neural network structure includes a first convolutional layer and a second convolutional layer, wherein the feature map output by the first convolutional layer corresponds to the upper probability, the lower probability and the first intermediate probability, and the feature map output by the second convolutional layer corresponds to the left probability, the right probability and the second intermediate probability.
[0048] Optionally, in an embodiment of the present invention, the loss function used when the neural network structure is trained is: Among them, L1 is the loss function of the first convolutional layer, L2 is the loss function of the second convolutional layer, f1 corresponds to the feature map output by the first convolutional layer, f2 corresponds to the feature map output by the second convolutional layer, k represents the feature dimension, and f 1ijk represents the probability of the kth dimension of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, P ijk represents the preset probability of the kth dimension of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, f 2ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the second convolutional layer, Q ijk Represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column in the feature map output by the second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer have H*W pixels respectively.
[0049] Optionally, in the embodiment of the present invention, the neural network structure further satisfies the following conditions: the input image is first downsampled and then upsampled, and the feature maps with the same feature dimensions in the downsampling stage and the upsampling stage are fused; and / or the hidden layer for downsampling is a resnet structure. The number of downsampling and upsampling can be determined according to the specific situation.
[0050] Optionally, in the embodiment of the present invention, the neural network structure may be a unet neural network structure. Compared with the traditional unet neural network structure, the unet neural network structure used in the embodiment of the present invention has made some improvements, see Figure 3As shown in the diagram, conv 3x3 means convolution 3x3; ReLU is a linear rectification function; copy and crop means copy and concatenate, which is equivalent to concat (fusion); max pool 2x2 means pooling layer 2x2; up-cpnv 2x2 means up-convolution 2x2; conv 1x1 means convolution 1x1. Figure 3 The technical solution provided by the embodiment of the present invention is introduced as an example. In this embodiment, the deficiencies in the prior art are overcome, and a complex scene arbitrary shape text detection method based on the unet neural network structure is proposed, which mainly solves the problem of uneven text area boundaries obtained by the semantic segmentation method and the problem of large number of parameters and time-consumingness of the bounding box regression method.
[0051] The input image is recorded as x, and the input image includes H*W pixels. x is input into the unet neural network structure. Compared with the traditional unet neural network structure, the unet neural network structure used in the embodiment of the present invention replaces the down-sampling convolution layer with the resnet structure, so that the probability obtained by the unet neural network structure can be more accurate. In the unet neural network structure used in the embodiment of the present invention, the input image is down-sampled and up-sampled 4 times, and the feature maps with the same feature dimensions in the up-sampling and down-sampling stages are concat (fused). The calculation speed of the unet neural network structure can be improved by upsampling, downsampling and fusion; in addition, fusion can also make the probability obtained by the unet neural network structure more accurate. The last output layer of the unet neural network structure used in the embodiment of the present invention includes a first convolutional layer and a second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer are both H*W*3 feature maps. Among them, the feature dimensions in the H*W*3 feature map output by the first convolutional layer correspond to the upper probability, lower probability and first intermediate probability of the pixel point; the feature dimensions in the H*W*3 feature map output by the second convolutional layer correspond to the left probability, right probability and second intermediate probability of the pixel point. The sum of the upper probability, the lower probability and the first intermediate probability is 1. The sum of the left probability, the right probability and the second intermediate probability is 1. The feature graph output by the first convolution layer and the second convolution layer is processed by softmax regression, and the loss function of the first convolution layer and the second convolution layer uses the cross entropy loss function. The loss function used when the unet neural network structure used in the embodiment of the present invention is trained is as follows: Among them, L1 is the loss function of the first convolutional layer, L2 is the loss function of the second convolutional layer, f1 corresponds to the feature map output by the first convolutional layer, f2 corresponds to the feature map output by the second convolutional layer, k represents the feature dimension, and f 1ijk represents the probability of the kth dimension of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, Pijk represents the preset probability of the kth dimension of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, f 2ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the second convolutional layer, Q ijk Represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column in the feature map output by the second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer have H*W pixels respectively. Using the H*W*3 feature map output by the first convolutional layer in the output layer of the unet neural network structure, the upper boundary corner points and the lower boundary corner points of all text areas on the image can be obtained; using the H*W*3 feature map output by the second convolutional layer in the output layer of the unet neural network structure, it can be determined whether each pixel point belongs to the starting corner point, the second intermediate corner point or the ending corner point. For each upper boundary starting corner point, the Euclidean metric is used to find the upper boundary corner point closest to the upper boundary starting corner point. When using the Euclidean metric to find the nearest upper boundary corner point, the search is performed based on the coordinate position of the upper boundary starting corner point (that is, the row and column in the feature map) and the coordinate positions of the remaining corner points (that is, the row and column in the feature map). If the nearest upper boundary corner point belongs to the second middle corner point, continue to search for the upper boundary corner point closest to the found upper boundary corner point until the nearest upper boundary ending corner point is found. If the upper boundary corner point closest to the upper boundary starting corner point is the ending corner point, terminate directly. Connect all the corner points obtained in the above steps to get the upper boundary of the text area. The method for finding the lower boundary is similar. Connect the upper boundary starting corner point with the nearest lower boundary starting corner point below it, and connect the upper boundary ending corner point with the nearest lower boundary ending corner point below it, and you can get a complete line of text area, that is, get the line text area detection box described in the embodiment of the present invention.
[0052] In the technical solution provided by the embodiment of the present invention, the advantages of the semantic segmentation method and the bounding box regression method are combined, and a method for detecting text based on the unet neural network structure is proposed. The pixel-level detection box regression is realized by the loss function and the unet neural network structure, which solves the problem of uneven boundaries of the bounding box area in the semantic segmentation method, and the unet neural network structure does not need to be divided into two stages for iterative update, but directly outputs the confidence that each pixel position on the image belongs to the corner point on the text area bounding box, which greatly improves the detection speed. In addition, in the embodiment of the present invention, the unet neural network structure is used as the skeleton. The unet neural network structure is characterized by a U-shaped structure, that is, downsampling is first performed and then upsampling is performed to obtain a feature map of the same size as the input image. The confidence that each pixel belongs to a certain category of corner points can be obtained through the feature map. The technical solution provided by the embodiment of the present invention improves the traditional unet neural network structure, so that the output of the unet neural network structure used in the embodiment of the present invention is directly used to predict the corner points of the upper and lower boundaries of the text area, and a smooth text detection box can be obtained by connecting multiple corner points belonging to the same text area.
[0053] Correspondingly, another aspect of an embodiment of the present invention further provides an apparatus for detecting text.
[0054] Figure 4It is a structural block diagram of a device for detecting text provided by another embodiment of the present invention. The device includes a pixel information acquisition module 1, a label determination module 2, a corner point determination module 3 and a line text area detection frame determination module 4. Among them, the pixel information acquisition module 1 is used to input the image of the text to be detected into the neural network structure to obtain the pixel information of each pixel in the multiple pixels of the image, wherein, for any pixel in the multiple pixels, the pixel information includes the upper probability, lower probability, first intermediate probability, left probability, right probability and second intermediate probability of the pixel, wherein the upper probability and the lower probability respectively represent the probability that the pixel is located at the upper boundary and the lower boundary of the single-line text area in the text area to be detected in the image, the first intermediate probability represents the probability that the pixel is located at the first intermediate position of the single-line text area except the upper boundary and the lower boundary, the left probability and the right probability respectively represent the probability that the pixel is located at the left boundary and the right boundary of the single-line text area, and the second intermediate probability represents the probability that the pixel is located at the left boundary and the right boundary of the single-line text area except the left boundary and the probability of the second middle position outside the right boundary; the label determination module 2 is used to determine the first label and the second label of the pixel point for any pixel point among the multiple pixel points based on the upper probability, the lower probability, the first middle probability, the left probability, the right probability and the second middle probability, wherein the first label indicates whether the pixel point is the upper boundary corner point, the lower boundary corner point and the first middle corner point of the single-line text area, and the second label indicates whether the pixel point is the starting corner point, the second middle corner point and the ending corner point of the single-line text area; the corner point determination module 3 is used to determine the upper boundary corner point and the lower boundary corner point in the same line of text according to the first labels and the second labels of the multiple pixel points; the line text area detection frame determination module 4 is used to connect the upper boundary corner point and the lower boundary corner point in the same line of text to obtain a line of text area detection frame.
[0055] Optionally, in an embodiment of the present invention, the label determination module determines the first label and the second label of the pixel point for any pixel point among the multiple pixels based on the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability, including: comparing the upper probability, the lower probability and the first intermediate probability and comparing the left probability, the right probability and the second intermediate probability to determine the first maximum one among the upper probability, the lower probability and the first intermediate probability and the second maximum one among the left probability, the right probability and the second intermediate probability; and determining the first label based on the first maximum one and determining the second label based on the second maximum one.
[0056] Optionally, in an embodiment of the present invention, the corner point determination module determines the upper boundary corner point and the lower boundary corner point in the same line of text according to the first label and the second label of multiple pixel points, including: for any upper boundary starting corner point, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point; determining the lower boundary starting corner point that is closest to the upper boundary starting corner point; and determining the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point, wherein the upper boundary starting corner point is the pixel point indicated as the upper boundary corner point by the first label and as the starting corner point by the second label, and the lower boundary starting corner point is the pixel point indicated as the lower boundary corner point by the first label and as the starting corner point by the second label.
[0057] Optionally, in an embodiment of the present invention, for any upper boundary starting corner point, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point that is closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point includes: searching for the upper boundary corner point that is closest to the upper boundary starting corner point, and when the found upper boundary corner point is the second intermediate corner point, continuing to search for the upper boundary corner point that is closest to the found upper boundary corner point, and continuously looping the search process until the found upper boundary corner point is the termination corner point, wherein all the found upper boundary corner points are the upper boundary corner points in the same line of text as the upper boundary starting corner point; searching for the lower boundary starting corner point that is closest to the upper boundary starting corner point; and searching for the lower boundary corner point in the same line of text as the found lower boundary starting corner point based on the content of searching for the upper boundary corner point in the same line of text as the upper boundary starting corner point, so as to obtain the lower boundary corner point in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point.
[0058] Optionally, in an embodiment of the present invention, for any pixel point, the pixel point information also includes a coordinate position, and for any upper boundary starting corner point, at least one of the following uses Euclidean measurement: finding the nearest upper boundary corner point, finding the nearest lower boundary starting corner point, and finding the nearest lower boundary corner point.
[0059] Optionally, in an embodiment of the present invention, the output layer of the neural network structure is a convolutional layer, and the feature map output by the output layer corresponds to the upper probability, the lower probability, the first intermediate probability, the left probability, the right probability and the second intermediate probability; preferably, the output layer of the neural network structure includes a first convolutional layer and a second convolutional layer, wherein the feature map output by the first convolutional layer corresponds to the upper probability, the lower probability and the first intermediate probability, and the feature map output by the second convolutional layer corresponds to the left probability, the right probability and the second intermediate probability.
[0060] Optionally, in an embodiment of the present invention, the loss function used when the neural network structure is trained is: Among them, L1 is the loss function of the first convolutional layer, L2 is the loss function of the second convolutional layer, f1 corresponds to the feature map output by the first convolutional layer, f2 corresponds to the feature map output by the second convolutional layer, k represents the feature dimension, and f 1ijk represents the probability of the kth dimension of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, P ijk represents the preset probability of the kth dimension of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, f 2ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the second convolutional layer, Q ijk Represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column in the feature map output by the second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer have H*W pixels respectively.
[0061] Optionally, in an embodiment of the present invention, the neural network structure also satisfies the following conditions: the input image is first downsampled and then upsampled, and feature maps with the same feature dimensions in the downsampling stage and the upsampling stage are fused; and / or the hidden layer for downsampling is a resnet structure.
[0062] The specific working principle and benefits of the device for detecting text provided by the embodiment of the present invention are similar to the specific working principle and benefits of the method for detecting text provided by the embodiment of the present invention, and will not be repeated here.
[0063] The device for detecting text includes a processor and a memory. The above-mentioned pixel information acquisition module, corner point determination module and line text area detection box determination module are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.
[0064] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the problem of many uneven burrs on the boundaries of text areas in the prior art can be solved by adjusting kernel parameters, reducing time consumption and improving calculation speed.
[0065] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0066] An embodiment of the present invention provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to enable a machine to execute the method described in the above embodiment.
[0067] Another aspect of the embodiments of the present invention further provides a processor, which is used to run a program, wherein the method described in the above embodiments is executed when the program is run.
[0068] Another aspect of the embodiments of the present invention further provides a device, the device includes a processor, a memory, and a program stored in the memory and executable on the processor, and the processor implements the method described in the above embodiments when executing the program. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0069] Another aspect of the embodiments of the present invention further provides a computer program product, including a computer program / instruction, which implements the method described in the above embodiments when executed by a processor.
[0070] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0071] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0072] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0074] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0075] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0076] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0077] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0078] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for detecting text, characterized in that The method includes: Inputting an image of text to be detected into a neural network structure to obtain pixel information of each pixel of a plurality of pixels of the image, wherein, for any pixel of the plurality of pixels, the pixel information includes an upper probability, a lower probability, a first intermediate probability, a left probability, a right probability and a second intermediate probability of the pixel, wherein the upper probability and the lower probability respectively represent the probability that the pixel is located at the upper boundary and the lower boundary of a single-line text area in the text area to be detected in the image, the first intermediate probability represents the probability that the pixel is located at a first intermediate position of the single-line text area excluding the upper boundary and the lower boundary, the left probability and the right probability respectively represent the probability that the pixel is located at the left boundary and the right boundary of the single-line text area, and the second intermediate probability represents the probability that the pixel is located at a second intermediate position of the single-line text area excluding the left boundary and the right boundary; For any pixel point among the multiple pixel points, based on the upper probability, the lower probability, the first middle probability, the left probability, the right probability and the second middle probability, determine a first label and a second label of the pixel point, wherein the first label indicates whether the pixel point is an upper boundary corner point, a lower boundary corner point and a first middle corner point of the single-line text area, and the second label indicates whether the pixel point is a starting corner point, a second middle corner point and an ending corner point of the single-line text area; Determining the upper boundary corner point and the lower boundary corner point in the same line of text according to the first labels and the second labels of the plurality of pixel points; and The upper boundary corner point and the lower boundary corner point in the same line of text are connected to obtain a line of text area detection box.
2. The method according to claim 1, characterized in that For any pixel point among the multiple pixels, determining a first label and a second label of the pixel point based on the upper probability, the lower probability, the first middle probability, the left probability, the right probability, and the second middle probability, includes: Compare the upper probability, the lower probability and the first intermediate probability and compare the left probability, the right probability and the second intermediate probability to determine a first maximum among the upper probability, the lower probability and the first intermediate probability and a second maximum among the left probability, the right probability and the second intermediate probability; and The first label is determined based on the first largest and the second label is determined based on the second largest.
3. The method according to claim 1, characterized in that Determining the upper boundary corner point and the lower boundary corner point in the same line of text according to the first label and the second label of the plurality of pixel points comprises: For any upper boundary starting corner point, Determine the upper boundary corner point that is in the same line of text as the upper boundary starting corner point; Determine the lower boundary starting corner point that is closest to the upper boundary starting corner point; and Determine the lower boundary corner point that is in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point, wherein the upper boundary starting corner point is the pixel point indicated as the upper boundary corner point by the first label and as the starting corner point by the second label, and the lower boundary starting corner point is the pixel point indicated as the lower boundary corner point by the first label and as the starting corner point by the second label.
4. The method according to claim 3, characterized in that For any of the upper boundary starting corner points, determining the upper boundary corner point in the same line of text as the upper boundary starting corner point, the lower boundary starting corner point closest to the upper boundary starting corner point, and the lower boundary corner point in the same line of text as the lower boundary starting corner point closest to the upper boundary starting corner point comprises: Find the upper boundary corner point that is closest to the upper boundary starting corner point, and when the found upper boundary corner point is the second intermediate corner point, continue to find the upper boundary corner point that is closest to the found upper boundary corner point, and continuously loop the search process until the found upper boundary corner point is the ending corner point, wherein all the found upper boundary corner points are the upper boundary corner points that are in the same line of text as the upper boundary starting corner point; Find the lower boundary starting corner point that is closest to the upper boundary starting corner point; and According to the content of the upper boundary corner point that is in the same line of text as the upper boundary starting corner point, the lower boundary corner point that is in the same line of text as the found lower boundary starting corner point is searched to obtain the lower boundary corner point that is in the same line of text as the lower boundary starting corner point that is closest to the upper boundary starting corner point.
5. The method according to claim 4, characterized in that For any of the pixel points, the pixel point information also includes a coordinate position, For any of the upper boundary starting corner points, at least one of the following uses the Euclidean metric: finding the upper boundary corner point that is closest to the upper boundary, finding the lower boundary starting corner point that is closest to the lower boundary, and finding the lower boundary corner point that is closest to the lower boundary.
6. The method according to claim 1, characterized in that The output layer of the neural network structure is a convolutional layer, and the feature map output by the output layer corresponds to the upper probability, the lower probability, the first middle probability, the left probability, the right probability and the second middle probability. The output layer of the neural network structure includes a first convolutional layer and a second convolutional layer, wherein the feature map output by the first convolutional layer corresponds to the upper probability, the lower probability and the first middle probability, and the feature map output by the second convolutional layer corresponds to the left probability, the right probability and the second middle probability.
7. The method according to claim 6, characterized in that The loss function used when the neural network structure is trained is: Wherein, L1 is the loss function of the first convolutional layer, L2 is the loss function of the second convolutional layer, f1 corresponds to the feature map output by the first convolutional layer, f2 corresponds to the feature map output by the second convolutional layer, k represents the feature dimension, and f 1ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the first convolutional layer, P ijk represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column of the feature map output by the first convolutional layer, f 2ijk represents the k-th dimension probability of the pixel in the i-th row and j-th column of the feature map output by the second convolutional layer, Q ijk Represents the preset probability of the kth dimension of the pixel point in the i-th row and j-th column in the feature map output by the second convolutional layer, and the feature maps output by the first convolutional layer and the second convolutional layer respectively have H*W pixels.
8. The method according to claim 6 or 7, characterized in that: The neural network structure also satisfies the following conditions: the input image is first downsampled and then upsampled, and feature maps with the same feature dimensions in the downsampling stage and the upsampling stage are fused; and / or the hidden layer for downsampling is a resnet structure.
9. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores instructions, which are used to enable a machine to execute the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multidirectional text region detection method and device based on boundary prediction
CN112580624A
Method and apparatus for neural network training and construction and method and apparatus for object detection
US20180032840A1