Text recognition method, device, storage medium and electronic device

By deformation processing on non-horizontal text lines, horizontal text areas are generated, and the problem of low recognition accuracy caused by incomplete text lines in the prior art is solved, and a higher text recognition accuracy is achieved.

CN114973268BActive Publication Date: 2025-05-13BEIJING CHUANGYING ORIENTAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210475607.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-05-13
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

When existing text recognition methods deal with non-horizontal text lines, it is difficult to maintain the integrity of text lines, resulting in low recognition accuracy.

Method used

By obtaining the initial text area of ​​the image to be detected, and when it is determined that it is a non-horizontal state, deformation processing is performed to obtain the horizontal text area, and then text recognition is performed.

Benefits of technology

The horizontal text area after deformation processing is recognized, which avoids the problem of text lines being truncated and improves the accuracy of text recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973268B_ABST
    Figure CN114973268B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a text recognition method, device, storage medium and electronic device. The method obtains an initial text area corresponding to an image to be detected; when it is determined that the initial text area is in a non-horizontal state, the initial text area is deformed to obtain a horizontal text area corresponding to the initial text area; and the text in the image to be detected is determined based on the horizontal text area. That is, when it is determined that the initial text area corresponding to the image to be detected is in a non-horizontal state, the present disclosure first deforms the initial text area, and then performs text recognition based on the deformed horizontal text area. Since the shape of the horizontal text area is relatively regular, its contour will not fit the text line too closely, so that the text in the text line recognized based on the horizontal text area will not be truncated, and the text line will be more complete, thereby improving the accuracy of text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a text recognition method, device, storage medium and electronic device. Background Art

[0002] Common text recognition methods can be divided into printed text recognition and handwritten text recognition. In addition to facing various problems in printed text recognition, handwritten text recognition is also affected by writing style. Especially in educational scenarios, it is difficult for primary school students to ensure that the content of the same line is horizontal and vertical when answering questions. It is easy for the answer text line to have various curved shapes such as arcs and waves. Based on this, arbitrary-shaped text lines are generated at the source of text line detection needs.

[0003] In the related art, the outline of a text line can be predicted through a neural network, and then the text in the outline can be recognized. However, due to the irregular shape of the text line, the predicted outline fits the text line too closely, causing the text to be easily truncated, resulting in an incomplete text line, which makes the text recognition accuracy relatively low. Summary of the invention

[0004] In order to solve the above problems, the present disclosure provides a text recognition method, device, storage medium and electronic device.

[0005] In a first aspect, the present disclosure provides a text recognition method, the method comprising:

[0006] Get the initial text area corresponding to the image to be detected;

[0007] In the case where it is determined that the initial text area is in a non-horizontal state, deforming the initial text area to obtain a horizontal text area corresponding to the initial text area;

[0008] The text in the image to be detected is determined according to the horizontal text area.

[0009] Optionally, determining that the initial text area is in a non-horizontal state includes:

[0010] Determine the minimum bounding rectangle corresponding to the initial text area;

[0011] Determine an area ratio between the area of ​​the initial text area and the rectangular area of ​​the minimum circumscribed rectangle;

[0012] When the area ratio is less than or equal to a preset ratio threshold, it is determined that the initial text area is in a non-horizontal state.

[0013] Optionally, before determining the minimum bounding rectangle corresponding to the initial text area, the method further includes:

[0014] For each pixel point in the initial text area, determine a moving direction corresponding to the pixel point according to the position of the pixel point, and determine a target position corresponding to the pixel point according to the moving direction and a preset moving distance;

[0015] Determining an extended text area corresponding to the initial text area according to a target position corresponding to each pixel point;

[0016] Determining the minimum bounding rectangle corresponding to the initial text area includes:

[0017] Determine a minimum bounding rectangle corresponding to the extended text area.

[0018] Optionally, obtaining an initial text area corresponding to the image to be detected includes:

[0019] The image to be detected is input into a pre-trained text region detection model to obtain the initial text region output by the text region detection model.

[0020] Optionally, the text region detection model includes a feature acquisition sub-model, a feature enhancement sub-model and a contour detection sub-model, the output end of the feature acquisition sub-model is coupled with the input end of the feature enhancement sub-model, and the output end of the feature enhancement sub-model is coupled with the input end of the contour detection sub-model; the step of inputting the image to be detected into a pre-trained text region detection model to obtain the initial text region output by the text region detection model includes:

[0021] Inputting the image to be detected into the feature acquisition sub-model to obtain a plurality of feature maps output by the feature acquisition sub-model, wherein different feature maps correspond to different sizes;

[0022] Inputting the plurality of feature maps into the feature enhancement sub-model, and performing enlargement enhancement processing and reduction enhancement processing on the plurality of feature maps through the feature enhancement sub-model, so as to obtain a plurality of target feature maps output by the feature enhancement sub-model;

[0023] According to the plurality of target feature maps, the initial text area is acquired through the contour detection sub-model.

[0024] Optionally, acquiring the initial text contour by using the contour detection sub-model according to the plurality of target feature maps comprises:

[0025] Performing splicing processing on a plurality of the target feature maps to obtain a target splicing feature map;

[0026] The target splicing feature map is input into the contour detection sub-model to obtain the initial text area output by the contour detection sub-model.

[0027] Optionally, the text region detection model is trained in the following manner:

[0028] Acquire multiple sample sets, the sample sets including sample images and true binary images corresponding to the sample images, the binary images being used to represent text areas in the sample images;

[0029] The target neural network model is trained by using a plurality of the sample sets to obtain the text region detection model.

[0030] Optionally, obtaining multiple sample sets includes:

[0031] Acquire a plurality of the sample images and a text boundary corresponding to each of the sample images;

[0032] For each of the sample images, a boundary distance is determined based on a preset adjustment coefficient and the area-to-perimeter ratio of the text boundary corresponding to the sample image, a true value threshold map corresponding to the sample image is determined based on the boundary distance, a true value probability map corresponding to the sample image is determined based on the text boundary, and the true value binary map corresponding to the sample image is determined based on the true value threshold map and the true value probability map.

[0033] Optionally, the training of a target neural network model by using a plurality of the sample sets to obtain the text region detection model comprises:

[0034] The model training step is executed in a loop until it is determined that the trained target neural network model meets the preset stop iteration condition according to the true binary image and the sample binary image, and the trained target neural network model is used as the text area detection model; the sample binary image is determined according to the sample threshold image and the sample probability image, and the sample threshold image and the sample probability image are images output after the sample image is input into the trained target neural network model;

[0035] The model training step includes:

[0036] Inputting a plurality of the sample images into the target neural network model to obtain the sample threshold map and the sample probability map corresponding to each of the sample images output by the target neural network model;

[0037] Determining the sample binary map according to the sample threshold map and the sample probability map;

[0038] When it is determined according to the true binary image and the sample binary image that the trained target neural network model does not satisfy the preset stop iteration condition, a target loss value is determined according to the true binary image and the sample binary image, and the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model, and the trained target neural network model is used as a new target neural network model.

[0039] In a second aspect, the present disclosure provides a text recognition device, the device comprising:

[0040] A region acquisition module is used to acquire an initial text region corresponding to the image to be detected;

[0041] A state determination module, configured to, when determining that the initial text region is in a non-horizontal state, perform deformation processing on the initial text region to obtain a horizontal text region corresponding to the initial text region;

[0042] The text recognition module is used to determine the text in the image to be detected according to the horizontal text area.

[0043] Optionally, the state determination module is further used to:

[0044] Determine the minimum bounding rectangle corresponding to the initial text area;

[0045] Determine an area ratio between the area of ​​the initial text area and the rectangular area of ​​the minimum circumscribed rectangle;

[0046] When the area ratio is less than or equal to a preset ratio threshold, it is determined that the initial text area is in a non-horizontal state.

[0047] Optionally, the device further comprises:

[0048] A position determination module, configured to determine, for each pixel point in the initial text area, a moving direction corresponding to the pixel point according to the position of the pixel point, and determine a target position corresponding to the pixel point according to the moving direction and a preset moving distance;

[0049] A region determination module, used for determining an extended text region corresponding to the initial text region according to a target position corresponding to each pixel point;

[0050] The state determination module is further used for:

[0051] Determine a minimum bounding rectangle corresponding to the extended text area.

[0052] Optionally, the region acquisition module is further used to:

[0053] The image to be detected is input into a pre-trained text region detection model to obtain the initial text region output by the text region detection model.

[0054] Optionally, the text region detection model includes a feature acquisition sub-model, a feature enhancement sub-model and a contour detection sub-model, the output end of the feature acquisition sub-model is coupled with the input end of the feature enhancement sub-model, and the output end of the feature enhancement sub-model is coupled with the input end of the contour detection sub-model; the region acquisition module is further used to:

[0055] Inputting the image to be detected into the feature acquisition sub-model to obtain a plurality of feature maps output by the feature acquisition sub-model, wherein different feature maps correspond to different sizes;

[0056] Inputting the plurality of feature maps into the feature enhancement sub-model, and performing enlargement enhancement processing and reduction enhancement processing on the plurality of feature maps through the feature enhancement sub-model, so as to obtain a plurality of target feature maps output by the feature enhancement sub-model;

[0057] According to the plurality of target feature maps, the initial text area is acquired through the contour detection sub-model.

[0058] Optionally, the region acquisition module is further used to:

[0059] Performing splicing processing on a plurality of the target feature maps to obtain a target splicing feature map;

[0060] The target splicing feature map is input into the contour detection sub-model to obtain the initial text area output by the contour detection sub-model.

[0061] Optionally, the region acquisition module is further used to:

[0062] Acquire multiple sample sets, the sample sets including sample images and true binary images corresponding to the sample images, the binary images being used to represent text areas in the sample images;

[0063] The target neural network model is trained by using a plurality of the sample sets to obtain the text region detection model.

[0064] Optionally, the region acquisition module is further used to:

[0065] Acquire a plurality of the sample images and a text boundary corresponding to each of the sample images;

[0066] For each of the sample images, a boundary distance is determined based on a preset adjustment coefficient and the area-to-perimeter ratio of the text boundary corresponding to the sample image, a true value threshold map corresponding to the sample image is determined based on the boundary distance, a true value probability map corresponding to the sample image is determined based on the text boundary, and the true value binary map corresponding to the sample image is determined based on the true value threshold map and the true value probability map.

[0067] Optionally, the region acquisition module is further used to:

[0068] The model training step is executed in a loop until it is determined that the trained target neural network model meets the preset stop iteration condition according to the true binary image and the sample binary image, and the trained target neural network model is used as the text area detection model; the sample binary image is determined according to the sample threshold image and the sample probability image, and the sample threshold image and the sample probability image are images output after the sample image is input into the trained target neural network model;

[0069] The model training step includes:

[0070] Inputting a plurality of the sample images into the target neural network model to obtain the sample threshold map and the sample probability map corresponding to each of the sample images output by the target neural network model;

[0071] Determining the sample binary map according to the sample threshold map and the sample probability map;

[0072] When it is determined according to the true binary image and the sample binary image that the trained target neural network model does not satisfy the preset stop iteration condition, a target loss value is determined according to the true binary image and the sample binary image, and the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model, and the trained target neural network model is used as a new target neural network model.

[0073] In a third aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0074] In a fourth aspect, the present disclosure provides a terminal, including:

[0075] a memory having a computer program stored thereon;

[0076] A processor is used to execute the computer program in the memory to implement the steps of the method described in the first aspect above.

[0077] Through the above technical scheme, the initial text area corresponding to the image to be detected is obtained; when it is determined that the initial text area is in a non-horizontal state, the initial text area is deformed to obtain a horizontal text area corresponding to the initial text area; according to the horizontal text area, the text in the image to be detected is determined. That is to say, in the present disclosure, when it is determined that the initial text area corresponding to the image to be detected is in a non-horizontal state, the initial text area is first deformed, and then text recognition is performed based on the horizontal text area after the deformation process. Since the shape of the horizontal text area is relatively regular, its contour will not fit the text line too closely, so that the text in the text line obtained by recognizing the horizontal text area will not be truncated, and the text line will be more complete, thereby improving the accuracy of text recognition.

[0078] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings:

[0080] Figure 1 is a flowchart of a text recognition method shown in an exemplary embodiment of the present disclosure;

[0081] Figure 2 is a schematic diagram of a text area shown in an exemplary embodiment of the present disclosure;

[0082] Figure 3 is a flowchart of another text recognition method shown in an exemplary embodiment of the present disclosure;

[0083] Figure 4 is a flowchart of a method for training a text region detection model shown in an exemplary embodiment of the present disclosure;

[0084] Figure 5 is a schematic diagram of an image shown in an exemplary embodiment of the present disclosure;

[0085] Figure 6 is a schematic diagram of a model structure shown in an exemplary embodiment of the present disclosure;

[0086] Figure 7 is a schematic diagram of an extended text area shown in an exemplary embodiment of the present disclosure;

[0087] Figure 8 is a schematic diagram of a circumscribed rectangle shown in an exemplary embodiment of the present disclosure;

[0088] Fig. 9is a block diagram of a text recognition device shown in an exemplary embodiment of the present disclosure;

[0089] Fig.10 is a block diagram of another text recognition device shown in an exemplary embodiment of the present disclosure;

[0090] Fig.11 It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0091] The specific implementation of the present disclosure is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the present disclosure, and is not used to limit the present disclosure.

[0092] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the device is located and with the authorization given by the owner of the corresponding device.

[0093] First, the application scenario of the present disclosure is described. Currently, commonly used text line detection methods include traditional methods (non-deep learning), target detection methods, and text detection methods. Traditional methods can be text positioning methods such as binarization, connected domain analysis, and projection analysis. This method can generally only be used in simple scenes, and the robustness to complex scenes is relatively poor, and it is difficult to have strong generalization. The target detection method can be YOLO series, MaskRCNN, RetinaNet, CenterNet, etc. The restriction of the rectangular frame in this method is not robust to the inclined text and curved text that are easy to produce in the photo scene. The text detection method can be PSENet, Craft, CTPN, EAST, FCENet, etc. CTPN, EAST, etc. are based on regression methods, which are difficult to describe curved text of arbitrary shapes; PSENet is based on segmentation, classifying each pixel of the input image into a binary map, and then clustering to obtain the final result. The processing is relatively complicated and the network process is long; FCENet uses the method of contour characterization model, which usually lacks the completeness of the description of the text contour, has slightly poor real-time performance, and the output form cannot be directly processed for text recognition.

[0094] In order to solve the above-mentioned problems, the present disclosure provides a text recognition method, device, storage medium and electronic device. When it is determined that the initial text area corresponding to the image to be detected is in a non-horizontal state, the initial text area is first deformed, and then text recognition is performed based on the horizontal text area after the deformation process. Since the shape of the horizontal text area is relatively regular, its outline will not be too close to the text line, so that the text in the text line recognized according to the horizontal text area will not be truncated, and the text line is more complete, thereby improving the accuracy of text recognition.

[0095] The present disclosure is described below in conjunction with specific embodiments.

[0096] Figure 1 is a flowchart of a text recognition method shown in an exemplary embodiment of the present disclosure. Figure 1 As shown, the method may include:

[0097] S101, obtaining an initial text region corresponding to an image to be detected.

[0098] The text area may be the outline of the text in the image to be detected. Figure 2 is a schematic diagram of a text area shown in an exemplary embodiment of the present disclosure, such as Figure 2 As shown, the curved area is the initial text area.

[0099] In this step, the image to be detected may be input into a pre-trained text region detection model to obtain the initial text region output by the text region detection model.

[0100] S102: When it is determined that the initial text region is in a non-horizontal state, a deformation process is performed on the initial text region to obtain a horizontal text region corresponding to the initial text region.

[0101] In this step, after obtaining the initial text area, it can be determined whether the initial text area is in a non-horizontal state, that is, whether the text in the image to be detected is curved text. In a possible implementation, the minimum bounding rectangle corresponding to the initial text area can be determined, and the area ratio between the area of ​​the initial text area and the rectangular area of ​​the minimum bounding rectangle can be determined. When the area ratio is less than or equal to a preset ratio threshold, it is determined that the initial text area is in a non-horizontal state.

[0102] Furthermore, when it is determined that the initial text area is in a non-horizontal state, the initial text area can be deformed by TPS (Thin Plate Spline). For example, N matching points can be first determined on the contour of the initial text area, and the N evenly distributed rectangular contour points are taken as the points after deformation. The width of this rectangle can be the same as the width of the minimum circumscribed rectangle of the initial text area, and the length of this rectangle can be 1.2 times the length of the minimum circumscribed rectangle of the initial text area. Afterwards, the curve text contour points corresponding to the initial text area can be directly transformed into rectangular text contour points by TPS, and the text in the curve contour is correspondingly transformed into the rectangular contour. At this time, the rectangular contour obtained is the horizontal text area.

[0103] S103: Determine the text in the image to be detected according to the horizontal text area.

[0104] In this step, after the horizontal text region is obtained, the text in the horizontal text region can be recognized by the method of the prior art to obtain the text in the image to be detected.

[0105] By adopting the above method, when it is determined that the initial text area corresponding to the image to be detected is in a non-horizontal state, the initial text area is first deformed, and then text recognition is performed based on the horizontal text area after the deformation process. Since the shape of the horizontal text area is relatively regular, its outline will not be too close to the text line, so that the text in the text line recognized according to the horizontal text area will not be truncated, and the text line is more complete, thereby improving the accuracy of text recognition.

[0106] Figure 3 is a flowchart of another text recognition method shown in an exemplary embodiment of the present disclosure, such as Figure 3 As shown, the method may include:

[0107] S301: Input the image to be detected into a pre-trained text region detection model to obtain the initial text region output by the text region detection model.

[0108] The text region detection model may include a feature acquisition submodel, a feature enhancement submodel and a contour detection submodel, wherein the output of the feature acquisition submodel is coupled to the input of the feature enhancement submodel, and the output of the feature enhancement submodel is coupled to the input of the contour detection submodel. The text region may be the contour of the text in the image to be detected.

[0109] Figure 4 is a flowchart of a method for training a text region detection model shown in an exemplary embodiment of the present disclosure. Figure 4 As shown, the method may include:

[0110] S1. Obtain multiple sample sets.

[0111] The sample set may include a sample image and a true binary image corresponding to the sample image, and the binary image may be used to represent a text area in the sample image.

[0112] In one possible implementation, multiple sample images and text boundaries corresponding to each sample image can be obtained. For each sample image, the boundary distance is determined based on a preset adjustment coefficient and the area-to-perimeter ratio of the text boundary corresponding to the sample image. A true value threshold map corresponding to the sample image is determined based on the boundary distance. A true value probability map corresponding to the sample image is determined based on the text boundary. And based on the true value threshold map and the true value probability map, the true value binary map corresponding to the sample image is determined.

[0113] For example, Figure 5 is a schematic diagram of an image shown in an exemplary embodiment of the present disclosure, such as Figure 5 As shown, it includes a first wireframe, a second wireframe and a third wireframe. The second wireframe is the text boundary. The text boundary can be manually annotated. After obtaining multiple sample images and the text boundary corresponding to each sample image, the preset adjustment coefficient can be obtained. The preset adjustment coefficient can be pre-set according to the thickness of the text line. For example, the preset adjustment coefficient can be 1.2. After that, the area-to-perimeter ratio of the text boundary can be determined, and the product of the area-to-perimeter ratio and the preset adjustment coefficient is used as the boundary distance. According to the boundary distance and the second wireframe, the first wireframe and the third wireframe are determined, and the area enclosed by the first wireframe and the third wireframe is used as the true value threshold map. Finally, the pixel value of the pixel point in the first wireframe can be set to 0.3. For each pixel point in the true value threshold map, the pixel value of the pixel point can be calculated by the following formula:

[0114] v=1-d / D (1)

[0115] Among them, v is the pixel value, d is the distance between the pixel point and the second wireframe, and D is the boundary distance.

[0116] Furthermore, after calculating the pixel value of each pixel in the threshold map, the true value probability map corresponding to the sample image can be represented by the pixel value. Figure 5 Taking the image shown as an example, the pixel value of the pixel point in the first wireframe is 0.3, the pixel value of the speed limit point in the second wireframe is 0, and the pixel value range from the second wireframe to the first wireframe is between 0-1 and increases. From the second wireframe to the third wireframe, the pixel value range is also between 0-1 and increases.

[0117] After obtaining the true value threshold map and the true value probability map corresponding to the sample image, the difference between the true value threshold map and the true value probability map can be used as the true value binary map.

[0118] S2. Train the target neural network model using multiple sample sets to obtain the text region detection model.

[0119] After obtaining multiple sample sets, the model training steps can be executed in a loop until the trained target neural network model is determined to meet the preset stop iteration condition based on the true binary image and the sample binary image, and the trained target neural network model is used as the text area detection model; the sample binary image is determined based on the sample threshold image and the sample probability image, and the sample threshold image and the sample probability image are the images output after the sample image is input into the trained target neural network model.

[0120] The target neural network model can be a network model optimized based on FCENet (Fourier Contour Embedding Net). For example, the ResNet18 lightweight network can be used as a skeleton network, and a feature enhancement model is added on this basis. The feature increment model can be FPEM (Feature Pyramid Enhancement Module). Figure 6 is a schematic diagram of a model structure shown in an exemplary embodiment of the present disclosure, such as Figure 6 As shown, FPEM can be a U-shaped structure, which includes two stages, one is the enhancement stage of feature map enlargement, and the other is the enhancement stage of feature map reduction. In the enhancement stage of feature map enlargement, the input is multiple feature maps ( Figure 6 The sizes of the feature maps are 1 / 4, 1 / 8, 1 / 16, and 1 / 32, respectively. Then, the 1 / 32 feature map is upsampled by a factor of 2 and added to the 1 / 16 feature map pixel-wise. 3*3 depthwise separable convolution, 1*1 convolution, BN (Batch Normalization) and Relu (Rectified Linear Unit) are performed to obtain the 1 / 16 feature map ( Figure 6 In the reduction stage, the 1 / 4 feature map is upsampled by a factor of 2, and then added to the 1 / 8 feature map pixel by pixel. 3*3 depthwise separable convolution, 1*1 convolution, BN and ReLU are performed to obtain the final 1 / 8 target feature map ( Figure 6 The last column), and so on, gradually calculate to 1 / 32 target feature map.

[0121] It should be noted that the output of FPEM can be used as the input of the next FPEM, so it can be a cascade structure. The cascade structure is conducive to the full fusion of image features and has a stronger feature extraction capability. The more cascades, the better the fusion and the larger the receptive field. In terms of computational cost, the architecture based on deep separable convolution consumes less time for calculation. After the cascaded FPEM is completed, feature fusion is performed to fuse each output of the cascaded FPEM. The usual fusion method uses channel splicing, but the speed will slow down as the number of channels increases. Therefore, this proposal adopts a pixel addition method. Through this calculation method, the final number of channels remains unchanged from before fusion, which can greatly reduce the amount of calculation.

[0122] Based on the above model structure, ResNet18 has a relatively small number of parameters and is faster during the inference process, but its receptive field is weaker, that is, its feature extraction capability is weaker. Based on this, after adding the feature enhancement model, the model's feature expression capability can be made stronger, thereby improving the accuracy of feature extraction on the basis of improving the operating efficiency of the text area detection model. In addition, the area surrounded by the first wireframe is the text center area of ​​the sample image, and the area surrounded by the second wireframe is the text area of ​​the sample image. During the training of the text area detection model, when the first wireframe and the second wireframe are relatively accurate, the locked text area and the text center area are also relatively accurate, thereby further improving the accuracy of text recognition.

[0123] The model training steps include:

[0124] S21. Input a plurality of the sample images into the target neural network model to obtain the sample threshold map and the sample probability map corresponding to each sample image output by the target neural network model.

[0125] S22. Determine the sample binary map according to the sample threshold map and the sample probability map.

[0126] After the sample threshold map and the sample probability map are obtained, the difference between the sample probability map and the sample threshold map can be used as the sample binary map.

[0127] S23. When it is determined according to the true binary image and the sample binary image that the trained target neural network model does not satisfy the preset stop iteration condition, a target loss value is determined according to the true binary image and the sample binary image, parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model, and the trained target neural network model is used as a new target neural network model.

[0128] In this step, after obtaining the image to be detected, the image to be detected can be input into the feature acquisition submodel to obtain multiple feature maps output by the feature acquisition submodel, and different feature maps have different sizes; the multiple feature maps are input into the feature enhancement submodel, and the multiple feature maps are enlarged and enhanced by the feature enhancement submodel to obtain multiple target feature maps output by the feature enhancement submodel; according to the multiple target feature maps, the initial text area is obtained by the contour detection submodel. Among them, the feature acquisition submodel can correspond to ResNet18, and the feature enhancement submodel can correspond to FPEM.

[0129] For example, after obtaining multiple feature maps output by the feature acquisition submodel, multiple target feature maps can be spliced ​​to obtain a target spliced ​​feature map, and the target spliced ​​feature map is input into the contour detection submodel to obtain the initial text area output by the contour detection submodel. For example, multiple target feature maps can be spliced ​​in a concat manner.

[0130] S302: Determine the minimum bounding rectangle corresponding to the initial text area.

[0131] In this step, after obtaining the initial text area corresponding to the image to be detected, for each pixel point in the initial text area, the moving direction corresponding to the pixel point is determined according to the position of the pixel point, and the target position corresponding to the pixel point is determined according to the moving direction and the preset moving distance; the extended text area corresponding to the initial text area is determined according to the target position corresponding to each pixel point; and the minimum bounding rectangle corresponding to the extended text area is determined. The preset moving distance can be predetermined based on experience, and for example, the preset moving distance can be 1 / 10 of the width of the minimum bounding rectangle of the initial text area.

[0132] For example, for each pixel point in the initial text area, two adjacent edges of the pixel point may be determined first, and the sum of the unit vectors of the pixel point moving away from the two adjacent edges is used as the moving direction corresponding to the pixel point. Figure 7 is a schematic diagram of an extended text area shown in an exemplary embodiment of the present disclosure, such as Figure 7 As shown in FIG. 1 , the inner hexagon is the initial text area, the direction indicated by the arrow is the moving direction, and r is the preset moving distance. After that, the target position corresponding to the pixel point can be determined according to the moving direction and the preset moving distance. After determining the target position corresponding to each pixel point in the initial text area, the extended text area can be obtained, as shown in FIG. Figure 7 As shown, the outer hexagon is the extended text area.

[0133] S303: Determine the area ratio between the area of ​​the initial text area and the area of ​​the minimum circumscribed rectangle.

[0134] In this step, if the initial text area is horizontal, the area ratio of the initial text area in the minimum circumscribed rectangle will be relatively large; if the initial text area is non-horizontal, the area ratio of the initial text area in the minimum circumscribed rectangle will be relatively small. Figure 8 is a schematic diagram of a circumscribed rectangle shown in an exemplary embodiment of the present disclosure, such as Figure 8 As shown, it can be clearly seen that the area of ​​the initial text area in the horizontal state is relatively large in the minimum bounding rectangle. Based on this, after determining the minimum bounding rectangle, the area ratio between the area of ​​the initial text area and the rectangular area of ​​the minimum bounding rectangle can be determined, and the area ratio can be used to further determine whether the initial text area is in the horizontal state.

[0135] S304: When the area ratio is less than or equal to a preset ratio threshold, determine that the initial text area is in a non-horizontal state.

[0136] The preset ratio threshold may be obtained through pre-testing. For example, the preset ratio threshold may be 80%.

[0137] S305: When it is determined that the initial text region is in a non-horizontal state, a deformation process is performed on the initial text region to obtain a horizontal text region corresponding to the initial text region.

[0138] S306: Determine the text in the image to be detected according to the horizontal text area.

[0139] In this step, if the initial text area is in a horizontal state, the text in the image to be detected can be determined directly based on the initial text area.

[0140] By adopting the above method, when it is determined that the initial text area corresponding to the image to be detected is in a non-horizontal state, the initial text area is first deformed, and then text recognition is performed based on the horizontal text area after deformation. Since the shape of the horizontal text area is relatively regular, its contour will not fit the text line too closely, so that the text in the text line identified according to the horizontal text area will not be truncated, and the text line will be more complete, thereby improving the accuracy of text recognition. In addition, after optimizing the model structure of the text area detection model, the model consumes less time, which can save computing resources and costs when deployed online, and after expanding the initial text area, the accuracy of handwritten text line recognition can be further improved.

[0141] Fig. 9is a block diagram of a text recognition device shown in an exemplary embodiment of the present disclosure. Fig. 9 As shown, the device may include:

[0142] The region acquisition module 901 is used to acquire the initial text region corresponding to the image to be detected;

[0143] A state determination module 902 is used to perform deformation processing on the initial text area to obtain a horizontal text area corresponding to the initial text area when it is determined that the initial text area is in a non-horizontal state;

[0144] The text recognition module 903 is used to determine the text in the image to be detected according to the horizontal text area.

[0145] Optionally, the state determination module 902 is further configured to:

[0146] Determine the minimum bounding rectangle corresponding to the initial text area;

[0147] Determine the area ratio between the area of ​​the initial text area and the rectangular area of ​​the minimum circumscribed rectangle;

[0148] When the area ratio is less than or equal to a preset ratio threshold, it is determined that the initial text area is in a non-horizontal state.

[0149] Optionally, Fig.10 is a block diagram of another text recognition device shown in an exemplary embodiment of the present disclosure, such as Fig.10 As shown, the device also includes:

[0150] A position determination module 904 is used to determine, for each pixel point in the initial text area, a moving direction corresponding to the pixel point according to the position of the pixel point, and determine a target position corresponding to the pixel point according to the moving direction and a preset moving distance;

[0151] The region determination module 905 is used to determine the extended text region corresponding to the initial text region according to the target position corresponding to each pixel point;

[0152] The state determination module 902 is further used for:

[0153] Determine the minimum bounding rectangle corresponding to the extended text area.

[0154] Optionally, the region acquisition module 901 is further used for:

[0155] The image to be detected is input into a pre-trained text region detection model to obtain the initial text region output by the text region detection model.

[0156] Optionally, the text region detection model includes a feature acquisition sub-model, a feature enhancement sub-model and a contour detection sub-model, the output end of the feature acquisition sub-model is coupled with the input end of the feature enhancement sub-model, and the output end of the feature enhancement sub-model is coupled with the input end of the contour detection sub-model; the region acquisition module 901 is further used to:

[0157] Inputting the image to be detected into the feature acquisition sub-model to obtain a plurality of feature maps output by the feature acquisition sub-model, wherein different feature maps correspond to different sizes;

[0158] Inputting the plurality of feature maps into the feature enhancement sub-model, and performing enlargement enhancement processing and reduction enhancement processing on the plurality of feature maps through the feature enhancement sub-model, so as to obtain a plurality of target feature maps output by the feature enhancement sub-model;

[0159] According to the multiple target feature maps, the initial text area is obtained through the contour detection sub-model.

[0160] Optionally, the region acquisition module 901 is further used for:

[0161] Performing splicing processing on a plurality of the target feature maps to obtain a target splicing feature map;

[0162] The target splicing feature map is input into the contour detection sub-model to obtain the initial text area output by the contour detection sub-model.

[0163] Optionally, the region acquisition module 901 is further used for:

[0164] Acquire multiple sample sets, the sample sets including sample images and true binary images corresponding to the sample images, the binary images being used to represent text areas in the sample images;

[0165] The target neural network model is trained by using multiple sample sets to obtain the text area detection model.

[0166] Optionally, the region acquisition module 901 is further used for:

[0167] Obtain a plurality of the sample images and a text boundary corresponding to each of the sample images;

[0168] For each sample image, the boundary distance is determined according to the preset adjustment coefficient and the area-to-perimeter ratio of the text boundary corresponding to the sample image, the true value threshold map corresponding to the sample image is determined according to the boundary distance, the true value probability map corresponding to the sample image is determined according to the text boundary, and the true value binary map corresponding to the sample image is determined according to the true value threshold map and the true value probability map.

[0169] Optionally, the region acquisition module 901 is further used for:

[0170] The model training step is executed cyclically until the trained target neural network model is determined to meet the preset stop iteration condition according to the true binary image and the sample binary image, and the trained target neural network model is used as the text area detection model; the sample binary image is determined according to the sample threshold image and the sample probability image, and the sample threshold image and the sample probability image are the images output after the sample image is input into the trained target neural network model;

[0171] The model training steps include:

[0172] Inputting a plurality of the sample images into the target neural network model to obtain the sample threshold map and the sample probability map corresponding to each of the sample images output by the target neural network model;

[0173] Determine the sample binary map according to the sample threshold map and the sample probability map;

[0174] When it is determined according to the true binary image and the sample binary image that the trained target neural network model does not satisfy the preset stop iteration condition, a target loss value is determined according to the true binary image and the sample binary image, and the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model, and the trained target neural network model is used as a new target neural network model.

[0175] Through the above-mentioned device, when it is determined that the initial text area corresponding to the image to be detected is in a non-horizontal state, the initial text area is first deformed, and then text recognition is performed based on the horizontal text area after the deformation process. Since the shape of the horizontal text area is relatively regular, its outline will not be too close to the text line, so that the text in the text line recognized according to the horizontal text area will not be truncated, and the text line is more complete, thereby improving the accuracy of text recognition.

[0176] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0177] Fig.11 FIG. 1 is a block diagram of an electronic device 1100 according to an exemplary embodiment of the present disclosure. Fig.11 As shown, the electronic device 1100 may include: a processor 1101 and a memory 1102. The electronic device 1100 may also include one or more of a multimedia component 1103, an input / output (I / O) interface 1104, and a communication component 1105.

[0178] The processor 1101 is used to control the overall operation of the electronic device 1100 to complete all or part of the steps in the above-mentioned text recognition method. The memory 1102 is used to store various types of data to support the operation of the electronic device 1100, and these data may include, for example, instructions for any application or method used to operate on the electronic device 1100, and application-related data, such as contact data, sent and received messages, pictures, audio, video, etc. The memory 1102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 1103 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 1102 or sent through the communication component 1105. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 1104 provides an interface between the processor 1101 and other interface modules, and the other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 1105 is used for wired or wireless communication between the electronic device 1100 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 1105 may include: Wi-Fi module, Bluetooth module, NFC module, etc.

[0179] In an exemplary embodiment, the electronic device 1100 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned text recognition method.

[0180] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by a processor, the steps of the above-mentioned text recognition method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 1102 including program instructions, and the above-mentioned program instructions can be executed by the processor 1101 of the electronic device 1100 to complete the above-mentioned text recognition method.

[0181] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned text recognition method when executed by the programmable device.

[0182] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, the technical solution of the present disclosure can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the present disclosure. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0183] In addition, various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A text recognition method, characterized in that: The method comprises: Get the initial text area corresponding to the image to be detected; In the case where it is determined that the initial text area is in a non-horizontal state, deforming the initial text area to obtain a horizontal text area corresponding to the initial text area; Determining text in the image to be detected according to the horizontal text area; The step of obtaining the initial text area corresponding to the image to be detected comprises: Inputting the image to be detected into a pre-trained text region detection model to obtain the initial text region output by the text region detection model; The text region detection model is trained in the following way: Acquire multiple sample sets, the sample sets including sample images and true binary images corresponding to the sample images, the binary images being used to represent text areas in the sample images; The target neural network model is trained by using a plurality of the sample sets to obtain the text region detection model; The training of the target neural network model by using the plurality of sample sets to obtain the text region detection model comprises: The model training step is executed in a loop until it is determined that the trained target neural network model meets the preset stop iteration condition according to the true binary image and the sample binary image, and the trained target neural network model is used as the text area detection model; the sample binary image is determined according to the sample threshold image and the sample probability image, and the sample threshold image and the sample probability image are images output after the sample image is input into the trained target neural network model; The model training step includes: Inputting a plurality of the sample images into the target neural network model to obtain the sample threshold map and the sample probability map corresponding to each of the sample images output by the target neural network model; Determining the sample binary map according to the sample threshold map and the sample probability map; When it is determined according to the true binary image and the sample binary image that the trained target neural network model does not satisfy the preset stop iteration condition, a target loss value is determined according to the true binary image and the sample binary image, and the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model, and the trained target neural network model is used as a new target neural network model.

2. The method according to claim 1, characterized in that: Determining that the initial text area is in a non-horizontal state includes: Determine the minimum bounding rectangle corresponding to the initial text area; Determine an area ratio between the area of ​​the initial text area and the rectangular area of ​​the minimum circumscribed rectangle; When the area ratio is less than or equal to a preset ratio threshold, it is determined that the initial text area is in a non-horizontal state.

3. The method according to claim 2, characterized in that Before determining the minimum bounding rectangle corresponding to the initial text area, the method further includes: For each pixel point in the initial text area, determine a moving direction corresponding to the pixel point according to the position of the pixel point, and determine a target position corresponding to the pixel point according to the moving direction and a preset moving distance; Determining an extended text area corresponding to the initial text area according to a target position corresponding to each pixel point; Determining the minimum bounding rectangle corresponding to the initial text area includes: Determine a minimum bounding rectangle corresponding to the extended text area.

4. The method according to claim 1, characterized in that: The text region detection model includes a feature acquisition sub-model, a feature enhancement sub-model and a contour detection sub-model, wherein the output end of the feature acquisition sub-model is coupled with the input end of the feature enhancement sub-model, and the output end of the feature enhancement sub-model is coupled with the input end of the contour detection sub-model; The step of inputting the image to be detected into a pre-trained text region detection model to obtain the initial text region output by the text region detection model comprises: Inputting the image to be detected into the feature acquisition sub-model to obtain a plurality of feature maps output by the feature acquisition sub-model, wherein different feature maps correspond to different sizes; Inputting the plurality of feature maps into the feature enhancement sub-model, and performing enlargement enhancement processing and reduction enhancement processing on the plurality of feature maps through the feature enhancement sub-model, so as to obtain a plurality of target feature maps output by the feature enhancement sub-model; According to the plurality of target feature maps, the initial text area is acquired through the contour detection sub-model.

5. The method according to claim 4, characterized in that The acquiring the initial text area according to the plurality of target feature maps by using the contour detection sub-model comprises: Performing splicing processing on a plurality of the target feature maps to obtain a target splicing feature map; The target splicing feature map is input into the contour detection sub-model to obtain the initial text area output by the contour detection sub-model.

6. The method according to claim 1, characterized in that The obtaining of multiple sample sets comprises: Acquire a plurality of the sample images and a text boundary corresponding to each of the sample images; For each of the sample images, a boundary distance is determined based on a preset adjustment coefficient and the area-to-perimeter ratio of the text boundary corresponding to the sample image, a true value threshold map corresponding to the sample image is determined based on the boundary distance, a true value probability map corresponding to the sample image is determined based on the text boundary, and the true value binary map corresponding to the sample image is determined based on the true value threshold map and the true value probability map.

7. A text recognition device, characterized in that: The device comprises: A region acquisition module is used to acquire an initial text region corresponding to the image to be detected; A state determination module, configured to, when determining that the initial text region is in a non-horizontal state, perform deformation processing on the initial text region to obtain a horizontal text region corresponding to the initial text region; A text recognition module, used for determining the text in the image to be detected according to the horizontal text area; The region acquisition module is further used for: Inputting the image to be detected into a pre-trained text region detection model to obtain the initial text region output by the text region detection model; The region acquisition module is further used for: Acquire multiple sample sets, the sample sets including sample images and true binary images corresponding to the sample images, the binary images being used to represent text areas in the sample images; The target neural network model is trained by using a plurality of the sample sets to obtain the text region detection model; The region acquisition module is further used for: The model training step is executed in a loop until it is determined that the trained target neural network model meets the preset stop iteration condition according to the true binary image and the sample binary image, and the trained target neural network model is used as the text area detection model; the sample binary image is determined according to the sample threshold image and the sample probability image, and the sample threshold image and the sample probability image are images output after the sample image is input into the trained target neural network model; The model training step includes: Inputting a plurality of the sample images into the target neural network model to obtain the sample threshold map and the sample probability map corresponding to each of the sample images output by the target neural network model; Determining the sample binary map according to the sample threshold map and the sample probability map; When it is determined according to the true binary image and the sample binary image that the trained target neural network model does not satisfy the preset stop iteration condition, a target loss value is determined according to the true binary image and the sample binary image, and the parameters of the target neural network model are updated according to the target loss value to obtain the trained target neural network model, and the trained target neural network model is used as a new target neural network model.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

9. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text Recognition Method and Terminal Device

    US20210326655A1