A street view text recognition method, system, device and medium

By using lightweight instance segmentation model and projection conversion technology in the street scene text recognition method, distorted interference in the image is removed, and combined with the scene text detection model and text recognition model, the problem of low recognition accuracy in the prior art is solved, achieving more efficient and accurate street scene text recognition.

CN115376118BActive Publication Date: 2025-05-30GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211024989.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-05-30
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

The existing street scene text recognition method mainly improves the text recognition part, and does not consider the impact of image problems such as blur, distortion, complex background and unclear light on the recognition results, resulting in more noise in the image and low accuracy of the recognition results.

Method used

The street scene image is segmented through a preset lightweight instance segmentation model to obtain the initial text area, and distortion factors such as distortion are removed through projection conversion. Then, the scene text detection model and the text recognition model are used to realize text area demarcation and text recognition respectively.

Benefits of technology

It improves the accuracy and efficiency of street scene text recognition, reduces noise in the image, and enhances the reliability of the recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376118B_ABST
    Figure CN115376118B_ABST
Patent Text Reader

Abstract

The present invention discloses a street view text recognition method, system, device and medium. When a street view image is received, the street view image is detected and recognized by a preset lightweight instance segmentation model, and the street view image is segmented. The initial text region segmented is subjected to projection transformation to obtain an intermediate text region. The intermediate text region is subjected to text region detection by a preset scene text detection model to determine the target text region where the scene text features are located. Then, the target characters in the target text region are recognized by a preset text recognition model to determine the image text corresponding to the street view image. By using the lightweight instance segmentation model to remove the non-text regions in the picture, and through projection transformation to remove interference factors such as distortion and aberration in the picture, and then combining the scene text detection model and the text recognition model for recognition, not only is the recognition efficiency fast, but also the recognition accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of character recognition, and in particular, to a street scene character recognition method, system, device and medium. Background Art

[0002] With the continuous development of artificial intelligence, more and more application scenarios have been explored. Among them, street scene character recognition is one of the current directions of artificial intelligence applications. Street scene character recognition is an optical character recognition (OCR) problem. OCR refers to the process of analyzing and recognizing an image file of text materials to obtain text and layout information, that is, recognizing the text in the image and returning it in the form of text.

[0003] The typical OCR problem-solving idea is image preprocessing - text region detection - text recognition. Among them, the technical bottlenecks affecting the recognition accuracy are text region detection and text recognition. At the same time, image problems such as blur, distortion, aberration, complex background, and unclear light are also factors affecting the recognition accuracy.

[0004] However, the existing street scene character recognition methods mainly improve the text recognition part, without considering the influence of image problems such as blur, distortion, aberration, complex background, and unclear light on the recognition result, resulting in more noise in the image and low accuracy of the recognition result. Summary of the Invention

[0005] The present invention provides a street scene character recognition method, system, device and medium, which solves the technical problem that the existing street scene character recognition methods mainly improve the text recognition part, without considering the influence of image problems such as blur, distortion, aberration, complex background, and unclear light on the recognition result, resulting in more noise in the image and low accuracy of the recognition result.

[0006] A street scene character recognition method provided by the present invention includes:

[0007] When receiving a street scene image, segment the street scene image through a preset lightweight instance segmentation model to obtain an initial text region;

[0008] Perform projection conversion on the initial text region to obtain an intermediate text region;

[0009] Detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located;

[0010] Recognize the target characters in the target text region through a preset text recognition model to determine the image text corresponding to the street scene image.

[0011] Optionally, the preset lightweight instance segmentation model includes multiple lightweight layers, a feature pyramid network layer, and a prediction class processing layer; the step of segmenting the street view image by the preset lightweight instance segmentation model when receiving the street view image to obtain the initial text region includes:

[0012] When receiving the street view image, extract the semantic features of the street view image at different scales through each of the lightweight layers respectively;

[0013] Perform multi-scale feature fusion on the semantic features through the feature pyramid network layer to obtain a semantic feature map;

[0014] Perform prediction on the semantic feature map through the prediction class processing layer to obtain prediction boxes corresponding to multiple prediction classes and class pixel probability maps within the prediction boxes;

[0015] Segment the corresponding class pixel probability maps according to the prediction classes by using the prediction boxes, and generate the initial text region in combination with the street view image.

[0016] Optionally, the prediction class processing layer includes a prototype feature segmentation layer and an instance class prediction layer; the step of performing prediction on the semantic feature map through the prediction class processing layer to obtain prediction boxes corresponding to multiple prediction classes and class pixel probability maps within the prediction boxes includes:

[0017] Segment the semantic feature map through the prototype feature segmentation layer to obtain multiple prototype feature maps;

[0018] Perform prediction on the semantic feature map through the instance class prediction layer to obtain multiple candidate boxes and multiple initial feature coefficients respectively corresponding to multiple prediction classes within the semantic feature map;

[0019] Remove duplicate candidate boxes within the multiple candidate boxes corresponding to the prediction classes according to the non-maximum suppression algorithm to obtain the prediction boxes corresponding to the prediction classes and multiple target feature coefficients;

[0020] Multiply all the prototype feature maps by the corresponding target feature coefficients respectively to obtain the class pixel probability maps within the prediction boxes.

[0021] Optionally, the step of segmenting the corresponding class pixel probability maps according to the prediction classes by using the prediction boxes and generating the initial text region in combination with the street view image includes:

[0022] Segment the corresponding class pixel probability maps according to the prediction boxes respectively to obtain multiple initial class pixel segmentation probability maps corresponding to the prediction classes;

[0023] Select the initial category pixel segmentation probability map according to a preset segmentation threshold to obtain the target category pixel segmentation probability map corresponding to the predicted category;

[0024] Multiply all the target category pixel segmentation probability maps by the street view image to generate the initial text region corresponding to the street view image.

[0025] Optionally, the step of performing a projective transformation on the initial text region to obtain an intermediate text region includes:

[0026] Perform a binarization operation on the initial text region to obtain a binarized region;

[0027] Calculate the minimum bounding rectangle corresponding to the white region within the binarized region to obtain the four vertex coordinates of the intermediate text region;

[0028] Calculate the projective transformation matrix corresponding to the vertex coordinates, and combine the preset specified coordinates to obtain the target vertex coordinates corresponding to each vertex coordinate;

[0029] Connect the target vertex coordinates in sequence to obtain the intermediate text region.

[0030] Optionally, the preset scene text detection model includes a feature extraction layer, a feature pyramid layer, and a trained inference layer; the step of detecting the intermediate text region through the preset scene text detection model to determine the target text region where the scene text features are located includes:

[0031] Extract multiple scene text features within the intermediate text region through the feature extraction layer;

[0032] Perform multi-scale feature fusion on the scene text features through the feature pyramid layer to obtain a scene feature map;

[0033] Infer the prediction probability map and threshold map corresponding to the scene feature map through the inference layer;

[0034] Calculate the approximate binary map corresponding to the feature map according to the pixel points corresponding to the prediction probability map and the threshold map, in combination with a preset approximate binary map formula;

[0035] Determine the target text region based on the approximate binary map.

[0036] Optionally, the preset text recognition model includes a convolutional network layer, a recurrent network layer, and a transcription layer; the step of recognizing the target characters within the target text region through the preset text recognition model to determine the image text corresponding to the street view image includes:

[0037] Extract multiple text feature maps within the target text region through the convolutional network layer, and convert the text feature maps into text feature sequences respectively;

[0038] Calculate the eigenvalue corresponding to the text feature sequence respectively through the recurrent network layer;

[0039] Perform exponential function conversion and scaling on all the eigenvalues to obtain a posterior probability matrix;

[0040] Through the transcription layer, calculate the character probability sequence corresponding to each column value in the posterior probability matrix by using the normalized exponential function;

[0041] Select the maximum value in the character probability sequence respectively, and use the character corresponding to the maximum value as the target character;

[0042] Use all the target characters as the image text corresponding to the street view image.

[0043] The present invention also provides a street view text recognition system, including:

[0044] An initial text region segmentation module, configured to segment the street view image through a preset lightweight instance segmentation model when receiving the street view image, to obtain an initial text region;

[0045] An intermediate text region obtaining module, configured to perform projection transformation on the initial text region to obtain an intermediate text region;

[0046] A target text obtaining module, configured to detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located;

[0047] An image text obtaining module, configured to recognize the target characters in the target text region through a preset text recognition model to determine the image text corresponding to the street view image.

[0048] The present invention also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor is caused to execute the steps of implementing the street view text recognition method as described in any one of the above.

[0049] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the street view text recognition method as described in any one of the above is implemented.

[0050] It can be seen from the above technical solutions that the present invention has the following advantages:

[0051] When receiving a street view image, the present invention detects and identifies the street view image through a preset lightweight instance segmentation model, and segments the street view image to obtain an initial text region. The initial text region obtained by segmentation is subjected to projection transformation to convert the text in the initial text region from the original size and position to a custom standard size and position, obtaining an intermediate text region. The intermediate text region is detected for text regions through a preset scene text detection model to determine the target text region where the scene text features are located. The target characters in the target text region are recognized through a preset text recognition model to determine the image text corresponding to the street view image, solving the technical problem that the existing street view text recognition methods mainly improve the text recognition part and do not consider the influence of image problems such as blur, distortion, aberration, complex background, and unclear light on the recognition result, resulting in more noise in the image and low accuracy of the recognition result. By using a lightweight instance segmentation model to remove non-text regions in the picture, and through projection transformation to remove interference factors such as distortion and aberration in the picture, and then combining a scene text detection model and a text recognition model to respectively implement text region delineation and text recognition operations, not only is the recognition efficiency fast, but also the recognition accuracy is high. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a flowchart of the steps of a street view text recognition method provided in Embodiment 1 of the present invention;

[0054] Figure 2 It is a flowchart of the steps of a street view text recognition method provided in Embodiment 2 of the present invention;

[0055] Figure 3 It is a structural block diagram of a street view text recognition system provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The embodiments of the present invention provide a street view text recognition method, system, device, and medium, which are used to solve the technical problem that the existing street view text recognition methods mainly improve the text recognition part and do not consider the influence of image problems such as blur, distortion, aberration, complex background, and unclear light on the recognition result, resulting in more noise in the image and low accuracy of the recognition result.

[0057] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the following described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0058] Please refer to Figure 1 , Figure 1 which is a step flowchart of a street view text recognition method provided in Embodiment 1 of the present invention.

[0059] A street view text recognition method provided by the present invention includes:

[0060] Step 101: When a street view image is received, segment the street view image through a preset lightweight instance segmentation model to obtain an initial text region.

[0061] The preset lightweight instance segmentation model for segmenting the street view image includes multiple lightweight layers, a feature pyramid network layer, and a prediction class processing layer, where the prediction class processing layer includes a prototype feature segmentation layer and an instance class prediction layer. The lightweight instance segmentation model uses a lightweight improved yolact deep neural network algorithm to detect and recognize the text region in the street view image, extract the text from the complex scene, and the final output result of the network only retains the pixel values of the original text region, and the remaining pixel values are all set to 0. Using the lightweight improved yolact deep neural network can completely extract the text region in the case of light difference or complex background. The lightweight instance segmentation model usually uses five lightweight layers, and the lightweight layer is also called the Shufflenet layer. Five lightweight Shufflenet layers are used to replace the convolutional module in yolact, thereby realizing pointwise grouped convolution. Grouped convolution can effectively reduce the capacity of the network and make the network lighter.

[0062] The initial text region refers to the image obtained by segmenting the street view image through the lightweight instance segmentation model and retaining the text region on the street view image.

[0063] In an embodiment of the present invention, when a street view image is received, the street view image is input into a lightweight instance segmentation model. First, semantic features of the street view image at different scales are respectively extracted through each lightweight layer. Then, multi-scale feature fusion is performed on the semantic features through a feature pyramid network layer to obtain a semantic feature map. Next, a prediction class processing layer performs prediction on the semantic feature map to obtain prediction boxes corresponding to multiple prediction classes and class pixel probability maps within the prediction boxes. Finally, the class pixel probability maps corresponding to the prediction boxes are segmented according to the prediction classes, and the initial text region is generated in combination with the street view image.

[0064] Step 102: Perform a projection transformation on the initial text region to obtain an intermediate text region.

[0065] The intermediate text region refers to determining the text position within the initial text region using a binarization algorithm and a contour detection algorithm, and then using a projection transformation technique to convert the text from its original size and position to a custom standard size and position. The region where the converted text is located is used as the image of the text region.

[0066] In an embodiment of the present invention, a binarization operation is performed on the initial text region to obtain a binarized region. The minimum bounding rectangle corresponding to the white region within the binarized region is calculated to obtain the four vertex coordinates of the intermediate text region. The projection transformation matrix corresponding to the vertex coordinates is calculated, and the target vertex coordinates corresponding to each vertex coordinate are obtained in combination with the preset specified coordinates. The target vertex coordinates are connected in sequence to obtain the intermediate text region.

[0067] Step 103: Detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located.

[0068] The preset scene text detection model is also called the DBNet deep neural network, and includes a feature extraction layer, a feature pyramid layer, and a trained inference layer. The feature extraction layer uses a 3*3 convolution kernel to extract multiple scene text features within the intermediate text region. The feature pyramid layer samples to the same size and simultaneously fuses the scene text features of different levels to finally obtain a scene feature map that is 1 / 4 of the original size. The trained inference layer uses the preset training samples to perform supervised training on the prediction probability map, the threshold map, and the approximate binary map. The prediction probability map and the approximate binary map use the same supervision signal, so that the trained inference layer can infer the corresponding prediction probability map and threshold map based on the scene feature map. Figure 1 / 4 of the scene feature map. The trained inference layer uses the preset training samples to perform supervised training on the prediction probability map, the threshold map, and the approximate binary map. The prediction probability map and the approximate binary map use the same supervision signal, so that the trained inference layer can infer the corresponding prediction probability map and threshold map based on the scene feature map.

[0069] The target text region refers to the image of the text region obtained after detecting the intermediate text region by inputting it into the scene text detection model.

[0070] In an embodiment of the present invention, the intermediate text region is input into a scene text detection model. First, multiple scene text features within the intermediate text region are extracted through a feature extraction layer. Then, multi-scale feature fusion is performed on the scene text features through a feature pyramid layer to obtain a scene feature map. Next, a prediction probability map and a threshold map corresponding to the scene feature map are inferred through an inference layer. According to the pixel points corresponding to the prediction probability map and the threshold map, and in combination with a preset approximate binary map formula, an approximate binary map corresponding to the feature map is calculated. Finally, based on the approximate binary map, the text region is framed to obtain the target text region.

[0071] Step 104: Identify the target characters within the target text region through a preset text recognition model to determine the image text corresponding to the street view image.

[0072] The preset text recognition model uses a CRNN deep neural network, including a convolutional network layer, a recurrent network layer, and a transcription layer. The convolutional network layer extracts useful information within the target text region, namely multiple text feature maps, through multiple convolution operations, and converts the text feature maps into text feature sequences respectively. The recurrent network layer uses a BLSTM network to calculate the feature values corresponding to the text feature sequences. The transcription layer uses a softmax function to calculate the character probability sequences corresponding to the values in each column of the posterior probability matrix, and selects the character corresponding to the position with the largest probability as the target character.

[0073] In an embodiment of the present invention, the target text region is input into the text recognition model. Multiple text feature maps within the target text region are extracted through the convolutional network layer, and the text feature maps are respectively converted into text feature sequences. Then, the feature values corresponding to the text feature sequences are calculated respectively through the recurrent network layer, and exponential function transformation and scaling are performed on all the feature values to obtain a posterior probability matrix. The transcription layer uses a softmax function to calculate the character probability sequences corresponding to the values in each column of the posterior probability matrix, respectively selects the maximum value within the character probability sequences, and takes the character corresponding to the maximum value as the target character. All the target characters are used as the image text corresponding to the street view image.

[0074] In the embodiment of the present invention, when a street view image is received, the street view image is detected and recognized by a preset lightweight instance segmentation model, and the street view image is segmented to obtain an initial text area. The initial text area obtained by segmentation is subjected to projection conversion to convert the text in the initial text area from the original size and position to a custom standard size and position, obtaining an intermediate text area. The intermediate text area is detected by a preset scene text detection model to determine the target text area where the scene text features are located. The target characters in the target text area are recognized by a preset text recognition model to determine the image text corresponding to the street view image, solving the technical problem that the existing street view text recognition methods mainly improve the text recognition part and do not consider the influence of image problems such as blur, distortion, aberration, complex background, and unclear light on the recognition result, resulting in a large amount of noise in the image and a low accuracy of the recognition result. By using the lightweight instance segmentation model to remove the non-text areas in the picture, and through projection conversion to remove interference factors such as distortion and aberration in the picture, and then combining the scene text detection model and the text recognition model to respectively implement text area delimitation and text recognition operations, not only the recognition efficiency is fast, but also the recognition accuracy is high.

[0075] Please refer to Figure 2 , Figure 2 which is a step flowchart of a street view text recognition method provided in the second embodiment of the present invention.

[0076] Step 201: When a street view image is received, semantic features of the street view image at different scales are respectively extracted through each lightweight layer.

[0077] In the embodiment of the present invention, when a street view image is received, the street view image is input into a trained lightweight instance segmentation model. The backbone network of the lightweight instance segmentation model is composed of five lightweight Shufflenet layers. The Shufflenet layers are used to respectively extract features of the street view image, and the output semantic feature sizes are 112×112, 56×56, 28×28, 14×14, and 7×7. Each lightweight layer respectively extracts shallow features such as low-level colors and high-level semantic features in the street view image. The output semantic features are transmitted to the feature pyramid network layer for the next step of processing.

[0078] Step 202: The semantic features are subjected to multi-scale feature fusion through the feature pyramid network layer to obtain a semantic feature map.

[0079] In the embodiment of the present invention, the semantic features at different scales are input into the feature pyramid network layer. The feature pyramid network layer extracts multi-scale feature representations of the semantic features, and then combines these feature representations to ensure that deep and large semantic feature maps can be extracted.

[0080] Step 203: Perform prediction on the semantic feature map through the prediction category processing layer to obtain prediction boxes corresponding to multiple prediction categories and class pixel probability maps within the prediction boxes.

[0081] Further, the prediction category processing layer includes a prototype feature segmentation layer and an instance category prediction layer. Step 203 may include the following sub-steps S11 - S14:

[0082] S11: Segment the semantic feature map through the prototype feature segmentation layer to obtain multiple prototype feature maps.

[0083] In an embodiment of the present invention, the semantic feature map is input into the prototype feature segmentation layer and the instance category prediction layer in parallel. After the semantic feature map is input into the prototype feature segmentation layer, i.e., the Protonet branch, 32 prototype feature maps will be generated. The prototype feature maps are also called prototype mask feature maps.

[0084] S12: Perform prediction on the semantic feature map through the instance category prediction layer to obtain multiple candidate boxes and multiple initial feature coefficients respectively corresponding to multiple prediction categories within the semantic feature map.

[0085] In an embodiment of the present invention, the instance category prediction layer is also called the Prediction Head branch. After the semantic feature map is input into the Prediction Head branch, multiple candidate boxes corresponding to the prediction category, the confidence level corresponding to each candidate box, the coordinates corresponding to each candidate box, and 32 initial feature coefficients corresponding one - to - one with the 32 prototype mask features output by the Protonet branch Figure 1 will be generated.

[0086] S13: Remove duplicate candidate boxes within the multiple candidate boxes corresponding to the prediction category according to the non - maximum suppression algorithm to obtain the prediction boxes corresponding to the prediction category and multiple target feature coefficients.

[0087] In an embodiment of the present invention, since there are a large number of overlapping regions in the multiple candidate boxes corresponding to the prediction category output by the Prediction Head branch, after the semantic feature map is output by the Prediction Head branch, the non - maximum suppression algorithm NMS is used to remove duplicates from the multiple candidate boxes. Only one candidate box is retained for each category, and this candidate box is used as the prediction box, the confidence level corresponding to the prediction box, the coordinates corresponding to the prediction box, and 32 target feature coefficients corresponding one - to - one with the 32 prototype mask features output by the Protonet branch Figure 1 will be generated.

[0088] S14: Multiply all the prototype feature maps by the corresponding target feature coefficients respectively to obtain the class pixel probability maps within the prediction boxes.

[0089] In an embodiment of the present invention, 32 prototype mask feature maps output by the Protonet branch are respectively multiplied by 32 corresponding target feature coefficients generated by the Prediction Head branch to obtain the class pixel probability maps to be segmented in each prediction box.

[0090] Step 204: Segment the corresponding class pixel probability maps according to the predicted classes, and generate initial text regions in combination with the street view images.

[0091] Further, step 204 performs the following sub-steps S21 - S23:

[0092] S21: Segment the corresponding class pixel probability maps according to the prediction boxes respectively to obtain multiple initial class pixel segmentation probability maps corresponding to the predicted classes.

[0093] In an embodiment of the present invention, according to the predicted classes, the prediction boxes of the predicted classes output after NMS in the Prediction Head branch are used to segment the corresponding class pixel probability maps to obtain multiple initial class pixel segmentation probability maps corresponding to the predicted classes.

[0094] S22: Select the initial class pixel segmentation probability maps according to a preset segmentation threshold to obtain the target class pixel segmentation probability maps corresponding to the predicted classes.

[0095] In an embodiment of the present invention, a segmentation threshold is set in advance based on the detection requirements. When multiple initial class pixel segmentation probability maps corresponding to the predicted classes are obtained, the target class pixel segmentation probability maps corresponding to the predicted classes are selected from the multiple initial class pixel segmentation probability maps according to the segmentation threshold.

[0096] S23: Multiply all the target class pixel segmentation probability maps by the street view images to generate the initial text regions corresponding to the street view images.

[0097] In an embodiment of the present invention, after determining the target pixel segmentation probability maps corresponding to each predicted class according to the segmentation threshold, since the target pixel segmentation probability map is a mask map containing the detection probability of the predicted class and the class pixel segmentation probability, multiplying the street view image by the mask map can generate the initial text region.

[0098] Step 205: Perform a projection transformation on the initial text region to obtain an intermediate text region.

[0099] Further, step 205 may include the following sub-steps S31 - S34:

[0100] S31: Perform a binarization operation on the initial text region to obtain a binarized region.

[0101] In the embodiment of the present invention, the Otsu method with an adaptive threshold is used to perform a binarization operation on the initial text region, and a binarized region is obtained, where the pixel values of the text region in the binarized region are 255 and those of the non-text region are 0.

[0102] S32. Calculate the minimum bounding rectangle corresponding to the white region in the binarized region to obtain the four vertex coordinates of the intermediate text region.

[0103] In the embodiment of the present invention, the white region refers to the text region in the binarized region, that is, the region with pixel value 255. The contour detection algorithm is used to calculate the minimum bounding rectangle corresponding to the white region in the binarized region, so as to determine the four vertex coordinates of the intermediate text region.

[0104] S33. Calculate the projection transformation matrix corresponding to the vertex coordinates, and combine the preset specified coordinates to obtain the target vertex coordinates corresponding to each vertex coordinate.

[0105] In the embodiment of the present invention, after determining the vertex coordinates (x, y) of the intermediate text region, the preset specified coordinates are (X, Y, Z). Based on the known vertex coordinates and the specified coordinates, the target vertex coordinates corresponding to each vertex coordinate are calculated in combination with the following transformation matrix formula.

[0106] The transformation matrix formula is as follows:

[0107]

[0108] Among them, (X, Y, Z) are the specified coordinates, (x, y, 1) are the vertex coordinates corresponding to the intermediate text region, and M is the projection transformation matrix to be obtained.

[0109] Let m 33 = 1, and then substitute the four vertices of the intermediate text region into the transformation matrix formula respectively to obtain 8 equations and solve for 8 unknowns. Finally, the projection transformation matrix is solved, and the specific formula is as follows:

[0110]

[0111]

[0112] Among them, (X’, Y’, Z’) are the corresponding specified coordinates in the two-dimensional coordinates.

[0113] S34. Connect the target vertex coordinates in sequence to obtain the intermediate text region.

[0114] In the embodiment of the present invention, after calculating the target vertex coordinates corresponding to each vertex coordinate through the transformation matrix formula, the target vertex coordinates are connected in sequence to obtain the intermediate text region.

[0115] Step 206: Detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located.

[0116] Further, the preset scene text detection model includes a feature extraction layer, a feature pyramid layer, and a trained inference layer. Step 206 may include the following sub-steps S41 - S45:

[0117] S41: Extract multiple scene text features within the intermediate text region through the feature extraction layer.

[0118] In an embodiment of the present invention, after inputting the intermediate text region into the scene text detection model, the feature extraction layer of the scene text detection model extracts the backbone of the intermediate text region, that is, extracts multiple scene text features within the intermediate text region through a 3*3 convolutional kernel.

[0119] S42: Perform multi-scale feature fusion on the scene text features through the feature pyramid layer to obtain a scene feature map.

[0120] In an embodiment of the present invention, input multiple scene text features into the feature pyramid layer. The feature pyramid layer upsamples to the same size and simultaneously fuses scene text features of different levels to finally obtain a scene feature map that is 1 / 4 of the street view image.

[0121] S43: Infer the corresponding prediction probability map and threshold map of the scene feature map through the inference layer.

[0122] In an embodiment of the present invention, the inference layer is pre-trained with a large number of prediction probability maps, threshold maps, and approximate binary maps for supervision, so that the inference layer can infer the corresponding prediction probability map and threshold map based on the scene feature map. When the scene feature map is input into the inference layer, the trained inference layer can quickly infer the corresponding prediction probability map and threshold map based on the scene feature map.

[0123] S44: Calculate the approximate binary map corresponding to the feature map according to the pixel points corresponding to the prediction probability map and the threshold map, in combination with a preset approximate binary map formula.

[0124] In an embodiment of the present invention, substitute the pixel points corresponding to the prediction probability map and the threshold map into the preset approximate binary map formula respectively to calculate the approximate binary map corresponding to the feature map.

[0125] The calculation formula for the approximate binary map is:

[0126]

[0127] where B i,j is the pixel point with coordinates (i, j) in the approximate binary map; P i,j is the pixel point with coordinates (i, j) in the prediction probability map, and Ti,j is the pixel at coordinates (i, j) in the threshold graph; k is the magnification factor, which is set to 50 according to experiments.

[0128] S45. Determine the target text region based on the approximate binary graph.

[0129] In the embodiment of the present invention, after obtaining the approximate binary graph corresponding to the feature graph, the target text region is determined based on the distribution of gray values on the approximate binary graph.

[0130] Step 207. Recognize the target characters in the target text region through a preset text recognition model, and determine the image text corresponding to the street view image.

[0131] Further, the preset text recognition model includes a convolutional network layer, a recurrent network layer, and a transcription layer. Step 207 may include the following sub-steps S51 - S56:

[0132] S51. Extract multiple text feature graphs in the target text region through the convolutional network layer, and convert the text feature graphs into text feature sequences respectively.

[0133] In the embodiment of the present invention, the target text region is input into the text recognition model. The convolutional network layer of the text recognition model extracts multiple text features in the target text region through multiple convolution operations to generate text feature graphs. For the convenience of subsequent use of the recurrent network layer, the above text feature graphs are respectively converted into corresponding text feature sequences.

[0134] S52. Calculate the eigenvalue corresponding to the text feature sequence through the recurrent network layer respectively.

[0135] In the embodiment of the present invention, the text feature sequence is input into the recurrent network layer, and the recurrent network layer calculates the eigenvalue corresponding to each text feature sequence respectively.

[0136] S53. Perform exponential function conversion and scaling on all eigenvalues to obtain a posterior probability matrix.

[0137] In the embodiment of the present invention, a softmax operation is performed on all eigenvalues, and all eigenvalues are respectively subjected to exponential function conversion and scaling, thereby generating a posterior probability matrix.

[0138] For example: for a k-dimensional vector x, softmax is used to convert the above result into a probability distribution p(x) of k categories. The specific calculation formula is:

[0139]

[0140] where x is a vector, x i and x j is one of its elements.

[0141] For the k-dimensional vector x, where x i ∈R, using the exponential function transformation, the value range of the elements can be transformed to (0, +∞). Then, by summing all the elements, the final result is scaled to [0, 1] to form a probability distribution, and finally a posterior probability matrix is formed.

[0142] S54. Calculate the word probability sequence corresponding to each column value in the posterior probability matrix by using the normalized exponential function through the transcription layer.

[0143] In the embodiment of the present invention, the posterior probability matrix is input into the transcription layer, and the transcription layer performs the normalized exponential function calculation on each column value of the posterior probability matrix to obtain the word probability sequence corresponding to each column value.

[0144] S55. Select the maximum value in the word probability sequence respectively, and use the character corresponding to the maximum value as the target character.

[0145] In the embodiment of the present invention, the maximum value is selected from all the position probability sequences respectively, and the character corresponding to the maximum value is used as the target character.

[0146] S56. Use all the target characters as the image text corresponding to the street view image.

[0147] In the embodiment of the present invention, the image text corresponding to the street view image is constructed by using all the target characters, and the image text is finally output and fed back to the user.

[0148] In an embodiment of the present invention, when a street view image is received, semantic features of the street view image at different scales are respectively extracted through each lightweight layer, and multi-scale feature fusion is performed on the semantic features through a feature pyramid network layer to obtain a semantic feature map. The semantic feature map is predicted through a prediction category processing layer to obtain prediction boxes corresponding to multiple prediction categories and a category pixel probability map within the prediction boxes. The category pixel probability map corresponding to the prediction box is segmented according to the prediction category, and an initial text region is generated in combination with the street view image. A binarization operation is performed on the initial text region to obtain a binarized region, the minimum bounding rectangle corresponding to the white region within the binarized region is calculated, the four vertex coordinates of the intermediate text region are obtained, the projection transformation matrix corresponding to the vertex coordinates is calculated, the target vertex coordinates corresponding to each vertex coordinate are obtained in combination with preset specified coordinates, and the target vertex coordinates are sequentially connected to obtain the intermediate text region. The intermediate text region is detected through a preset scene text detection model to determine the target text region where the scene text features are located, and the target characters within the target text region are recognized through a preset text recognition model to determine the image text corresponding to the street view image. Segmenting the street view image through a lightweight instance segmentation model can accurately segment the text region in the image, and using this method to remove interference information will make text recognition more accurate. The subsequent projection conversion can correct the distortion of the text in the image, facilitating subsequent text recognition. The last part uses two deep networks, DBNet and CRNN, to respectively implement the operations of delimiting the target text region and text recognition. Due to the addition of denoising processing and subsequent multi-model fusion, the overall recognition efficiency of the above method has been greatly improved compared with the current recognition algorithms in terms of accuracy.

[0149] Please refer to Figure 3 , Figure 3 The structural block diagram of a street view text recognition system provided in Embodiment 3 of the present invention.

[0150] An embodiment of the present invention provides a street view text recognition system, including:

[0151] An initial text region segmentation module 301, configured to segment a street view image through a preset lightweight instance segmentation model when a street view image is received, to obtain an initial text region.

[0152] An intermediate text region obtaining module 302, configured to perform projection conversion on the initial text region to obtain an intermediate text region.

[0153] A target text obtaining module 303, configured to detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located.

[0154] An image text obtaining module 304, configured to recognize the target characters within the target text region through a preset text recognition model to determine the image text corresponding to the street view image.

[0155] Optionally, the preset lightweight instance segmentation model includes multiple lightweight layers, a feature pyramid network layer, and a prediction class processing layer. The initial text region segmentation module 301 includes:

[0156] A semantic feature extraction module, configured to, when receiving a street view image, extract semantic features of the street view image at different scales through each lightweight layer.

[0157] A semantic feature map obtaining module, configured to perform multi-scale feature fusion on the semantic features through the feature pyramid network layer to obtain a semantic feature map.

[0158] A class pixel probability map obtaining module, configured to perform prediction on the semantic feature map through the prediction class processing layer to obtain prediction boxes corresponding to multiple prediction classes and class pixel probability maps within the prediction boxes.

[0159] An initial text region generation module, configured to segment the corresponding class pixel probability map according to the prediction class by using the prediction box, and generate an initial text region in combination with the street view image.

[0160] Optionally, the prediction class processing layer includes a prototype feature segmentation layer and an instance class prediction layer. The class pixel probability map obtaining module includes:

[0161] A prototype feature map obtaining module, configured to segment the semantic feature map through the prototype feature segmentation layer to obtain multiple prototype feature maps.

[0162] A candidate box and initial feature coefficient obtaining module, configured to perform prediction on the semantic feature map through the instance class prediction layer to obtain multiple candidate boxes and multiple initial feature coefficients respectively corresponding to multiple prediction classes within the semantic feature map.

[0163] A prediction box and target feature coefficient obtaining module, configured to remove duplicate candidate boxes within the multiple candidate boxes corresponding to the prediction class according to the non-maximum suppression algorithm to obtain a prediction box corresponding to the prediction class and multiple target feature coefficients.

[0164] A class pixel probability map obtaining sub-module, configured to multiply all the prototype feature maps by the corresponding target feature coefficients respectively to obtain the class pixel probability map within the prediction box

[0165] Optionally, the initial text region generation module includes:

[0166] An initial class pixel segmentation probability map obtaining module, configured to segment the corresponding class pixel probability map according to the prediction box respectively to obtain multiple initial class pixel segmentation probability maps corresponding to the prediction class.

[0167] The target category pixel segmentation probability map obtaining module is used to select the initial category pixel segmentation probability map according to a preset segmentation threshold, and obtain the target category pixel segmentation probability map corresponding to the predicted category.

[0168] The initial text region generation sub-module is used to multiply all the target category pixel segmentation probability maps by the street view image to generate the initial text region corresponding to the street view image.

[0169] Optionally, the intermediate text region obtaining module 302 includes:

[0170] The binary region obtaining module is used to perform a binarization operation on the initial text region to obtain a binary region.

[0171] The vertex coordinate obtaining module is used to calculate the minimum circumscribed rectangle corresponding to the white region in the binary region, and obtain the four vertex coordinates of the intermediate text region.

[0172] The target vertex coordinate obtaining module is used to calculate the projective transformation matrix corresponding to the vertex coordinates, and combine the preset specified coordinates to obtain the target vertex coordinates corresponding to each vertex coordinate.

[0173] The intermediate text region obtaining sub-module is used to sequentially connect the target vertex coordinates to obtain the intermediate text region.

[0174] Optionally, the preset scene text detection model includes a feature extraction layer, a feature pyramid layer, and a trained inference layer. The target text obtaining module 303 includes:

[0175] The scene text feature extraction module is used to extract multiple scene text features in the intermediate text region through the feature extraction layer.

[0176] The scene feature map obtaining module is used to perform multi-scale feature fusion on the scene text features through the feature pyramid layer to obtain a scene feature map.

[0177] The prediction probability map and threshold map inference module is used to infer the prediction probability map and threshold map corresponding to the scene feature map through the inference layer.

[0178] The approximate binary map calculation module is used to calculate the approximate binary map corresponding to the feature map according to the pixel points corresponding to the prediction probability map and the threshold map, in combination with a preset approximate binary map formula.

[0179] The target text obtaining sub-module is used to determine the target text region based on the approximate binary map.

[0180] Optionally, the preset text recognition model includes a convolutional network layer, a recurrent network layer, and a transcription layer. The image text obtaining module 304 includes:

[0181] A text feature sequence obtaining module, configured to extract multiple text feature maps within a target text region through a convolutional network layer, and convert the text feature maps into text feature sequences respectively.

[0182] An eigenvalue calculation module, configured to calculate eigenvalues corresponding to the text feature sequences respectively through a recurrent network layer.

[0183] A posterior probability matrix obtaining module, configured to perform exponential function conversion and scaling on all eigenvalues to obtain a posterior probability matrix.

[0184] A character probability sequence calculation module, configured to calculate a character probability sequence corresponding to each column value in the posterior probability matrix by using a normalized exponential function through a transcription layer.

[0185] A target character obtaining module, configured to select the maximum value in the character probability sequence respectively, and use the character corresponding to the maximum value as the target character.

[0186] An image text obtaining sub-module, configured to use all target characters as the image text corresponding to the street view image.

[0187] An embodiment of the present invention further provides an electronic device, which includes: a memory and a processor, and a computer program is stored in the memory; when the computer program is executed by the processor, the processor is caused to execute the street view text recognition method according to any one of the above embodiments.

[0188] The memory may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. The memory has a storage space for program codes for executing any method steps in the above methods. For example, the storage space for program codes may include respective program codes for implementing various steps in the above methods. These program codes may be read out from or written into one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. The program codes may be compressed in an appropriate form. When these codes are run by a computing processing device, the computing processing device is caused to execute each step in the street view text recognition method described above.

[0189] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the street view text recognition method according to any one of the above embodiments is implemented.

[0190] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0191] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0192] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0194] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0195] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A street view text recognition method, characterized in that, it includes: When receiving a street view image, segment the street view image through a preset lightweight instance segmentation model to obtain an initial text region; Perform projection transformation on the initial text region to obtain an intermediate text region; Detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located; Identify the target characters in the target text region through a preset text recognition model to determine the image text corresponding to the street view image; The preset lightweight instance segmentation model includes multiple lightweight layers, a feature pyramid network layer, and a prediction class processing layer; the step of segmenting the street view image through the preset lightweight instance segmentation model to obtain an initial text region when receiving a street view image includes: When receiving a street view image, extract the semantic features of the street view image at different scales through each of the lightweight layers; Perform multi-scale feature fusion on the semantic features through the feature pyramid network layer to obtain a semantic feature map; Perform prediction on the semantic feature map through the prediction class processing layer to obtain prediction boxes corresponding to multiple prediction classes and class pixel probability maps within the prediction boxes; Segment the corresponding class pixel probability maps according to the prediction classes using the prediction boxes, and generate an initial text region in combination with the street view image; The prediction class processing layer includes a prototype feature segmentation layer and an instance class prediction layer; the step of performing prediction on the semantic feature map through the prediction class processing layer to obtain prediction boxes corresponding to multiple prediction classes and class pixel probability maps within the prediction boxes includes: Segment the semantic feature map through the prototype feature segmentation layer to obtain multiple prototype feature maps; Perform prediction on the semantic feature map through the instance class prediction layer to obtain multiple candidate boxes and multiple initial feature coefficients respectively corresponding to multiple prediction classes within the semantic feature map; Remove duplicate candidate boxes within the multiple candidate boxes corresponding to the prediction classes according to the non-maximum suppression algorithm to obtain the prediction boxes corresponding to the prediction classes and multiple target feature coefficients; Multiply all the prototype feature maps by the corresponding target feature coefficients respectively to obtain the class pixel probability maps within the prediction boxes.

2. The street view text recognition method according to claim 1, characterized in that, the step of segmenting the corresponding class pixel probability maps according to the prediction classes using the prediction boxes and generating an initial text region in combination with the street view image includes: Segment the corresponding class pixel probability maps according to the prediction boxes respectively to obtain multiple initial class pixel segmentation probability maps corresponding to the prediction classes; Select the initial class pixel segmentation probability maps according to a preset segmentation threshold to obtain target class pixel segmentation probability maps corresponding to the prediction classes; Multiply all the target class pixel segmentation probability maps by the street view image to generate the initial text region corresponding to the street view image.

3. The street view text recognition method according to claim 1, characterized in that, The step of performing projection transformation on the initial text region to obtain an intermediate text region includes: Performing a binarization operation on the initial text region to obtain a binarized region; Calculating the minimum bounding rectangle corresponding to the white region within the binarized region to obtain the four vertex coordinates of the intermediate text region; Calculating the projection transformation matrix corresponding to the vertex coordinates, and combining the preset specified coordinates to obtain the target vertex coordinates corresponding to each vertex coordinate; Sequentially connecting the target vertex coordinates to obtain the intermediate text region.

4. The street view text recognition method according to claim 1, characterized in that the preset scene text detection model includes a feature extraction layer, a feature pyramid layer, and a trained inference layer; the step of detecting the intermediate text region through the preset scene text detection model to determine the target text region where the scene text features are located includes: Extracting a plurality of scene text features within the intermediate text region through the feature extraction layer; Performing multi-scale feature fusion on the scene text features through the feature pyramid layer to obtain a scene feature map; Inferring the prediction probability map and the threshold map corresponding to the scene feature map through the inference layer; Calculating the approximate binary map corresponding to the feature map according to the pixel points corresponding to the prediction probability map and the threshold map, in combination with the preset approximate binary map formula; Determining the target text region based on the approximate binary map.

5. The street view text recognition method according to claim 1, characterized in that the preset text recognition model includes a convolutional network layer, a recurrent network layer, and a transcription layer; the step of recognizing the target characters within the target text region through the preset text recognition model to determine the image text corresponding to the street view image includes: Extracting a plurality of text feature maps within the target text region through the convolutional network layer, and respectively converting the text feature maps into text feature sequences; Calculating the feature values corresponding to the text feature sequences respectively through the recurrent network layer; Performing exponential function transformation and scaling on all the feature values to obtain a posterior probability matrix; Calculating the text probability sequence corresponding to each column value within the posterior probability matrix through the transcription layer using the normalized exponential function; Respectively selecting the maximum value within the text probability sequence, and taking the character corresponding to the maximum value as the target character; Taking all the target characters as the image text corresponding to the street view image.

6. A street view text recognition system applied to the street view text recognition method according to claim 1, characterized in that it includes: An initial text region segmentation module, configured to segment the street view image through a preset lightweight instance segmentation model when receiving a street view image to obtain an initial text region; An intermediate text region obtaining module, configured to perform projection transformation on the initial text region to obtain an intermediate text region; A target text obtaining module, configured to detect the intermediate text region through a preset scene text detection model to determine the target text region where the scene text features are located; An image text obtaining module, configured to recognize target characters within the target text region through a preset text recognition model, and determine image text corresponding to the street view image.

7. An electronic device, characterized in that it includes a memory and a processor, a computer program is stored in the memory, and when the computer program is executed by the processor, the processor is caused to execute the steps of the street view text recognition method according to any one of claims 1-5.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed, it implements the street view text recognition method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Natural scene text detection method and system based on attention mechanism feature fusion and enhancement

    CN114255456A

  • License plate recognition method and apparatus, storage medium and terminal

    WO2022111355A1