Text recognition method, device, computer equipment and storage medium

By extracting feature of text images and classifying text instances, and splitting text sets, the problem of inaccurate text recognition in traditional technology is solved, and high-accurate text recognition is achieved.

CN113822116BActive Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110620895.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-03
Publication Date
2025-06-06
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

Traditional text recognition technology is difficult to accurately identify the situation where text is typing, resulting in inaccurate recognition.

Method used

By obtaining text images, performing feature extraction, classifying text instances of each pixel in the text image, determining the correspondence between each pixel point and text instance category, splitting the text image, obtaining the instance text image corresponding to the text instance category, and performing text recognition.

Benefits of technology

The accuracy of text recognition is improved, and by splitting the text into pieces, high-quality data to be identified, and accurate text recognition results are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822116B_ABST
    Figure CN113822116B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of machine learning, and provides a text recognition method, device, computer equipment and storage medium. The method comprises: obtaining a text image; performing feature extraction on the text image to obtain feature information of the text image; classifying each pixel in the text image into text instances according to the feature information, and determining the correspondence between each pixel and the text instance category, where the text instance category is an independent text entry category; splitting the text image according to the correspondence between each pixel and the text instance category to obtain an instance text image corresponding to the text instance category; performing text recognition on the instance text image to obtain a text recognition result. The present method can obtain accurate text recognition results and improve the accuracy of text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine learning, and in particular to a text recognition method, apparatus, computer device and storage medium. Background Art

[0002] With the development of computer technology, optical character recognition technology has emerged. Optical character recognition can be applied to text recognition scenarios with duplicate text. The existence of duplicate text means that there are foreground text and background text that interfere with each other in the image to be recognized, such as Figure 1 As shown in the automatic homework grading scenario, the existence of duplicate text means that there are handwritten answers and printed questions that interfere with each other, such as Figure 2 As shown in the figure, in the intelligent bill recognition scenario, the existence of duplicate text refers to the existence of user description information and bill background templates that interfere with each other (in Figure 2 They are marked with different blocks in the figure).

[0003] In traditional technology, text recognition is mainly performed on images to be recognized by constructing a large amount of corresponding training data and training a text recognition model to learn text objects (foreground text or background text).

[0004] However, the traditional technology can only learn one text object (foreground text or background text) and cannot further identify the existing duplicate text, which leads to the problem of inaccurate text recognition. Summary of the invention

[0005] Based on this, it is necessary to provide a text recognition method, apparatus, computer device and storage medium that can improve the accuracy of text recognition in response to the above technical problems.

[0006] A text recognition method, the method comprising:

[0007] Get text image;

[0008] Perform feature extraction on the text image to obtain feature information of the text image;

[0009] Classify each pixel in the text image into a text instance according to the feature information, and determine the corresponding relationship between each pixel and the text instance category, where the text instance category is an independent text entry category;

[0010] According to the correspondence between each pixel and the text instance category, the text image is split to obtain the instance text image corresponding to the text instance category;

[0011] Perform text recognition on the instance text image to obtain the text recognition result.

[0012] A text recognition device, comprising:

[0013] An acquisition module, used for acquiring text images;

[0014] A feature extraction module is used to extract features from text images to obtain feature information of the text images;

[0015] A classification module is used to classify each pixel in the text image into a text instance according to the feature information, and determine the corresponding relationship between each pixel and the text instance category, where the text instance category is an independent text entry category;

[0016] A splitting module is used to split the text image according to the correspondence between each pixel point and the text instance category to obtain an instance text image corresponding to the text instance category;

[0017] The recognition module is used to perform text recognition on the instance text image to obtain the text recognition result.

[0018] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0019] Get text image;

[0020] Perform feature extraction on the text image to obtain feature information of the text image;

[0021] Classify each pixel in the text image into a text instance according to the feature information, and determine the corresponding relationship between each pixel and the text instance category, where the text instance category is an independent text entry category;

[0022] According to the correspondence between each pixel and the text instance category, the text image is split to obtain the instance text image corresponding to the text instance category;

[0023] Perform text recognition on the instance text image to obtain the text recognition result.

[0024] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0025] Get text image;

[0026] Perform feature extraction on the text image to obtain feature information of the text image;

[0027] Classify each pixel in the text image into a text instance according to the feature information, and determine the corresponding relationship between each pixel and the text instance category, where the text instance category is an independent text entry category;

[0028] According to the correspondence between each pixel and the text instance category, the text image is split to obtain the instance text image corresponding to the text instance category;

[0029] Perform text recognition on the instance text image to obtain the text recognition result.

[0030] The above-mentioned text recognition method, device, computer equipment and storage medium obtain feature information of the text image by acquiring a text image, extracting features from the text image, classifying each pixel in the text image into text instances according to the feature information, determining the correspondence between each pixel and the text instance category, and splitting the text image according to the correspondence between each pixel and the text instance category to obtain instance text images corresponding to the text instance category. By performing text instance classification on each pixel in the text image, the templated text in the text image can be split to obtain non-templated instance text images, thereby providing high-quality data to be recognized for text recognition. Furthermore, by performing text recognition on the instance text image, accurate text recognition results can be obtained, thereby improving the accuracy of text recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of a printed text in one embodiment;

[0032] Figure 2 A schematic diagram of another embodiment in which a template text exists;

[0033] Figure 3 A flowchart of a text recognition method in one embodiment;

[0034] Figure 4 is a schematic diagram of a text image in one embodiment;

[0035] Figure 5 A schematic diagram of sampling using a trained downsampling network and a trained upsampling network in one embodiment;

[0036] Figure 6 A schematic diagram of a process of obtaining a corresponding relationship by using a pixel segmentation model in one embodiment;

[0037] Figure 7 A flowchart of a text recognition method in one embodiment;

[0038] Figure 8 is a schematic diagram of an example text image in one embodiment;

[0039] Fig. 9 is a structural block diagram of a text recognition device in one embodiment;

[0040] Fig.10 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0041] The solution provided in the embodiment of the present application relates to the field of machine learning technology. Machine Learning (ML) is a multi-disciplinary cross-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching.

[0042] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0043] In one embodiment, Figure 3 As shown, a text recognition method is provided. This embodiment uses the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server can be implemented as an independent server or a server cluster composed of multiple servers, or it can be a node in a blockchain. In this embodiment, the method includes the following steps:

[0044] Step 302, obtaining a text image.

[0045] The text image refers to a printed text image.

[0046] Specifically, when text recognition is required, the terminal will obtain the image to be processed, perform target detection on the image to be processed, divide the image to be processed into a template text area, an independent text area and a background area, and extract the text image from the image to be processed according to the template text area. The image to be processed refers to a complete image including the template text area, the independent text area and the background area, and the independent text area refers to a non-templated text area.

[0047] Furthermore, the terminal can perform target detection on the processed image through the trained target detection model. The trained target detection model can be obtained by training the sample training data. The sample training data here refers to the sample image containing the text area. The background area, the printed text area and the independent text area are marked in the sample image at the same time. When training, we can define the background area as category 0, the printed text area as category 1, and the independent text area as category 2. For example, Mask R-CNN (Mask Regions with CNN (Convolutional Neural Networks)) can be used to perform 3-category detection, detect the candidate box of each type of area and obtain the corresponding category. Mask R-CNN can use Resnet-50 (Deep residual network) as the backbone network. After extracting features, RPN (Region Proposal Network) is first used to determine the candidate interest area with the extracted features as input, and then each candidate interest area found by RPN is classified and located. For the real target corresponding to the candidate ROI, its label is sorted in descending order according to the coordinates of the upper left corner of the text (the upper leftmost label = 1, with priority given to left being greater than right and top being greater than bottom) and given a specific label.

[0048] Step 304: extract features from the text image to obtain feature information of the text image.

[0049] The feature information refers to information used to characterize the features of pixels in a text image. For example, the feature information may specifically refer to information obtained by combining pixel coordinate information of pixels and image information of the text image. For example, the feature information may specifically refer to a feature map obtained by extracting features from a text image in combination with pixel coordinate information of pixels.

[0050] Specifically, the terminal will first obtain the image channel data of the text image and the pixel coordinate information of each pixel point, splice the pixel coordinate information in the image feature in the form of an additional channel as additional position information, update the image channel data, obtain the image data to be sampled, and then obtain the feature information of the text image by sampling the image data to be sampled. Furthermore, when sampling the image data to be sampled, in order to obtain richer feature information, the sampling can be performed by downsampling first and then upsampling. Furthermore, when sampling and upsampling, a multi-level sampling network is used for sampling. When upsampling, the input of each level of the upsampling network is the upsampled feature map output by the upsampling network of the previous level and the downsampled feature map output by the downsampling network of the same level. In this way, the purpose of fusion learning can be achieved.

[0051] Step 306 , classify each pixel in the text image into a text instance according to the feature information, and determine the correspondence between each pixel and a text instance category, where the text instance category is an independent text entry category.

[0052] Among them, a text instance refers to an independent text item in a text image that is not copied, and an independent text item refers to text that does not have text overlapping areas and is classified by content. Figure 4 As shown, the text image includes three text instances with duplicate printing (represented by different letters (X, Y, Z), that is, texts composed of the same letters without text overlapping areas are independent text entries). The correspondence between each pixel point and the text instance category is used to characterize the attribution relationship between the pixel point and the text instance category, that is, whether the pixel point belongs to the text instance category. For example, the correspondence between each pixel point and the text instance category can be that the pixel point belongs to the text instance category, or it can be that the pixel point does not belong to the text instance category.

[0053] Specifically, after obtaining the feature information, the terminal will determine the number of text instances corresponding to the text image according to the feature information, and each text instance corresponds to a text instance category. After determining the text instance category, the terminal will classify each pixel in the text image according to the text instance category, and determine the corresponding relationship between each pixel and the text instance category, that is, whether each pixel belongs to the text instance category. Among them, the text instance category is mainly determined according to the number of text instances corresponding to the text image, that is, the terminal does not know the text instance category corresponding to the text image in advance. After obtaining the feature information, the terminal will first determine the number of text instances corresponding to the text image according to the feature information, and then determine the text instance category according to the number of text instances. Each text instance in the text image corresponds to a text instance category. Further, when classifying each pixel in the text image and determining the corresponding relationship between each pixel and the text instance category, the terminal can use the feature information to determine the category probability that each pixel in the text image belongs to the text instance category, and use the category probability to determine the corresponding relationship.

[0054] Step 108 , splitting the text image according to the correspondence between each pixel point and the text instance category to obtain an instance text image corresponding to the text instance category.

[0055] The example text image refers to an image containing only non-printed text.

[0056] Specifically, since the correspondence between each pixel point and the text instance category is used to characterize the relationship between the pixel point and the text instance category, the terminal can divide each pixel point in the text image according to the correspondence between each pixel point and the text instance category, and divide out a set of pixel points belonging to the text instance category. Through the set of pixel points belonging to the text instance category, the instance text image corresponding to the text instance category can be obtained.

[0057] Step 110: Perform text recognition on the example text image to obtain a text recognition result.

[0058] Specifically, after obtaining the instance text image, the terminal will use the instance text image as the data to be recognized for text recognition, and obtain a text recognition result by performing text recognition on the instance text image. The text recognition result includes all independent texts corresponding to the template text in the text image.

[0059] The above-mentioned text recognition method obtains a text image, performs feature extraction on the text image, obtains feature information of the text image, classifies each pixel in the text image into text instances according to the feature information, determines the correspondence between each pixel and the text instance category, and splits the text image according to the correspondence between each pixel and the text instance category to obtain instance text images corresponding to the text instance category. By performing text instance classification on each pixel in the text image, the templated text in the text image can be split to obtain non-templated instance text images, thereby providing high-quality data to be recognized for text recognition. Furthermore, by performing text recognition on the instance text image, accurate text recognition results can be obtained, thereby improving the accuracy of text recognition.

[0060] In one embodiment, feature extraction is performed on a text image to obtain feature information of the text image including:

[0061] Obtain image channel data of the text image and pixel coordinate information of each pixel point;

[0062] According to the pixel coordinate information, the image channel data is updated to obtain the image data to be sampled;

[0063] The image data to be sampled is sampled to obtain the feature information of the text image.

[0064] Among them, the channel is used to indicate how much data can be stored at each point, and the image channel data refers to the stored data corresponding to each pixel in the text image. For example, when the text image is an RGB image, the image channel data can specifically refer to the grayscale value corresponding to the R channel, G channel and B channel stored at each pixel. Furthermore, the text image can also be a four-channel image, which includes an A (alpha) channel in addition to the R channel, G channel and B channel to indicate transparency. Pixel coordinate information refers to the coordinate information of the pixel in the preset image coordinate system corresponding to the text image. For example, the pixel coordinate information can specifically refer to two-dimensional coordinate information, namely, X-axis coordinate information and Y-axis coordinate information. For example, the preset image coordinate system corresponding to the text image can specifically use the image center point of the text image as the coordinate origin.

[0065] Specifically, after acquiring the text image, the terminal will extract the image channel data from the text image, determine the pixel coordinate information of each pixel, and splice the pixel coordinate information into the image channel data in the form of an additional channel to update the image channel data, obtain the image data to be sampled, and obtain the feature information of the text image by sampling the image data to be sampled. Among them, when determining the pixel coordinate information of each pixel, the terminal will determine the coordinate origin of the preset image coordinate system corresponding to the text image according to the preset image coordinate system origin determination rule, and then determine the pixel coordinates of each pixel in the text image according to the coordinate origin. Among them, the preset image coordinate system origin determination rule is used to determine the coordinate origin in the text image. For example, the image coordinate system origin determination rule can be specifically to take the first pixel in the upper left corner as the coordinate origin, the first pixel in the upper right corner as the coordinate origin, the center point of the image as the coordinate origin, etc., which is not specifically limited in this embodiment.

[0066] Specifically, when sampling the image data to be sampled, in order to obtain richer feature information, the terminal will use the method of downsampling first and then upsampling. Furthermore, when sampling is performed both in downsampling and upsampling, a multi-level sampling network is used for sampling. When upsampling, the input of each level of the upsampling network is the upsampling feature map output by the upsampling network of the previous level and the downsampling feature map output by the downsampling network of the same level. In this way, the purpose of fusion learning can be achieved.

[0067] In this embodiment, by acquiring the image channel data of the text image and the pixel coordinate information of each pixel point, splicing the pixel coordinate information into the image channel data, updating the image channel data, and obtaining the image data to be sampled, the spatial position relationship of each pixel point can be fully utilized during feature extraction to obtain richer feature information.

[0068] In one embodiment, sampling the image data to be sampled to obtain feature information of the text image includes:

[0069] Down-sample the image data to be sampled to obtain a multi-scale down-sampled feature map;

[0070] Upsampling is performed based on the multi-scale downsampled feature map to obtain the feature information of the text image.

[0071] Among them, downsampling of the image data to be sampled can be achieved through a trained downsampling network. A trained downsampling network refers to a pre-trained network for downsampling. For example, the trained downsampling network may specifically include multiple levels of downsampling networks, each of which may specifically be composed of a leaky linear correction unit (leakyrelu), a convolution unit, and a batch normalization unit (batch normalization), wherein the convolution unit may specifically be a convolution unit with a convolution kernel size of 4 and a step size of 2. The multi-scale downsampling feature map refers to the downsampled image feature data output by each level of the downsampling network in the trained downsampling network. The trained downsampling network can be obtained by training with preset sample sampling image data, and the sample sampling image data refers to image data with the same type and quantity of image channel data as the image data to be sampled. The present embodiment does not limit the specific training method here, as long as accurate downsampling can be achieved.

[0072] Among them, upsampling according to the multi-scale downsampling feature map can be achieved through a trained upsampling network. The trained upsampling network refers to a pre-trained network for upsampling. For example, the trained upsampling network may specifically include multiple levels of upsampling networks, each of which may specifically be composed of a linear correction unit (relu), a deconvolution unit, a batch normalization unit (batch normalization) and a merging unit (concat), wherein the convolution kernel size and step size of the deconvolution unit are the same as those of the convolution unit in each level of the downsampling network, and the merging unit is used to merge the upsampling feature map to be fused output by the batch normalization unit and the downsampling feature map output by the downsampling network of the same level in the trained downsampling network, wherein the downsampling feature map output by the downsampling network of the same level is a part of the input of the upsampling network of this level, and the other part of the input of the upsampling network of this level is the upsampling feature map output by the upsampling network of the previous level. Figure, by processing the upsampled feature map output by the upsampling network of the previous level through a linear correction unit, a deconvolution unit, and a batch normalization unit, the upsampled feature map to be fused can be obtained, and by performing feature fusion on the upsampled feature map to be fused and the downsampled feature map output by the downsampling network of the same level, the upsampled feature map output by the upsampling network of this level can be obtained. The upsampled feature map output by the upsampling network of this level is the input of the upsampling network of the next level. When the upsampling network of the current level is the last level, the obtained upsampled feature map is the feature information of the text image. The trained upsampling network can be obtained by training with preset sample sampling image data. The sample sampling image data refers to image data whose type and quantity of image channel data are the same as those of the image data to be sampled. The present embodiment does not limit the specific training method here, as long as accurate upsampling can be achieved.

[0073] Specifically, when it is necessary to sample the image data to be sampled, the terminal will obtain the trained downsampling network and the trained upsampling network, and first use the downsampling networks at each level in the trained downsampling network to sequentially downsample the image data to be sampled, and obtain the downsampled feature map corresponding to each level of the downsampling network, that is, the multi-scale downsampled feature map. Among them, when downsampling is performed sequentially, the input of the first level of the downsampling network is the image data to be sampled, and the input of each subsequent level of the downsampling network is the downsampled feature map output by the downsampling network at the previous level.

[0074] Specifically, after completing the downsampling, the terminal will input the downsampled feature map output by the last level of the downsampling network as input data into the trained upsampling network, and start upsampling using the upsampling networks at each level in the trained upsampling network. When upsampling, the input of the first level of the upsampling network is the downsampled feature map output by the last level of the downsampling network, and then the input of each level of the upsampling network is the upsampled feature map output by the upsampling network at the previous level and the downsampled feature map output by the downsampling network at the same level.

[0075] Specifically, each level of the upsampling network includes a linear correction unit, a deconvolution unit, a batch normalization unit and a merging unit. The linear correction unit, the deconvolution unit and the batch normalization unit are used to process the upsampling feature map output by the upsampling network of the previous level to obtain the upsampling feature map to be fused. The merging unit is used to merge the upsampling feature map to be fused output by the batch normalization unit and the downsampling feature map output by the downsampling network of the same level to obtain the upsampling feature map corresponding to the upsampling network of this level.

[0076] For example, Figure 5 As shown, the trained downsampling network and the trained upsampling network are connected in sequence, where the first 8 layers (D1-D8) are downsampling networks at each level, and each downsampling network at each level consists of a leaky linear correction unit, a convolution unit, and a batch normalization unit. When the convolution unit is a convolution unit with a convolution kernel size of 4 and a step size of 2, the trained downsampling network can convert the image data to be sampled corresponding to a text image of size 512*512 into a 1*1 feature map. The last 8 layers (S8-S1) are upsampling networks at each level, and each upsampling network at each level consists of a linear correction unit, a deconvolution unit, a batch normalization unit, and a merging unit, where the convolution kernel size and step size of the deconvolution unit are the same as those of the convolution unit in each downsampling network at each level. Furthermore, before D1, there are also two coordinate information convolution layers (C1, C2), which are used to update the pixel coordinate information of each pixel point to the image channel data, as shown in FIG. Figure 5As shown, the input of the coordinate information convolution layer is the text image. For example, assuming that the input data size of the coordinate information convolution layer is N*C*H*W, where N is the batch size, C is the number of image channels, H is the image height, and W is the image width, after passing through the coordinate information convolution layer, the output data size is N*(C+2)*H*W, where the output data is the image data to be sampled.

[0077] In this embodiment, by first downsampling the image data to be sampled to obtain a multi-scale downsampling feature map, and then upsampling is performed according to the multi-scale downsampling feature map to obtain feature information of the text image, so that richer feature information can be obtained.

[0078] In one embodiment, text instance classification is performed on each pixel in the text image according to the feature information, and determining the correspondence between each pixel and the text instance category includes:

[0079] Classify each pixel in the text image into text instances according to the feature information, and determine the category probability of each pixel belonging to each text instance category;

[0080] Compare the category probability with the preset probability threshold to determine the correspondence between each pixel and each text instance category.

[0081] The category probability refers to the probability that each pixel belongs to each text instance category. The preset probability threshold refers to a pre-set probability value used to determine whether a pixel belongs to a text instance category. When the category probability is greater than the preset probability threshold, the pixel is considered to belong to the text instance category and is part of the text instance.

[0082] Specifically, the terminal will classify each pixel in the text image as a text instance based on the obtained feature information, determine the category probability of each pixel belonging to each text instance category, compare the category probability with the preset probability threshold, and when the category probability is greater than the preset probability threshold, the pixel is considered to belong to the text instance category, and when the category probability is not greater than the preset probability threshold, the pixel is considered not to belong to the text instance category, thereby determining the correspondence between all pixels in the text image and all text instance categories. Among them, the text instance category corresponds to the number of text instances. The terminal can determine the number of text instances in the text image based on the feature information, and then determine the text instance category based on the number of text instances, and each text instance corresponds to a text instance category.

[0083] Furthermore, the above process of determining the correspondence between each pixel and each text instance category can be obtained through a trained classification network. The trained classification network takes feature information as input, and by convolving the feature information, the category probability of each pixel in the text image belonging to each text instance category can be obtained. By comparing the category probability and the preset probability threshold, the correspondence between each pixel and each text instance category can be determined. Among them, the trained classification network can be obtained by training with preset sample classification data. This embodiment does not limit the specific training method here, as long as accurate text instance classification of pixels can be achieved.

[0084] In this embodiment, by classifying each pixel in the text image as a text instance according to the feature information, determining the category probability of each pixel belonging to each text instance category, and comparing the category probability with the preset probability threshold, the correspondence between each pixel and each text instance category can be determined.

[0085] In one embodiment, text recognition is performed on the example text image to obtain a text recognition result including:

[0086] Get the trained text recognition model;

[0087] The trained text recognition model is used to perform text recognition on the instance text image to obtain the text recognition result.

[0088] The trained text recognition model refers to a pre-trained model for text recognition, which can be obtained by training sample text recognition data, and the sample text recognition data refers to a sample text image pre-labeled with a non-printed text area, and the sample text image does not include a printed text area. The trained text recognition model in this embodiment can be based on various common text recognition networks, and this embodiment is not specifically limited here.

[0089] Specifically, after obtaining the instance text image, the terminal will obtain a trained text recognition model, input the instance text image into the trained text recognition model, perform text recognition on the instance text image through the trained text recognition model, and obtain a text recognition result, which includes all independent texts corresponding to the template text in the text image.

[0090] In this embodiment, by acquiring a trained text recognition model and performing text recognition on an example text image using the trained text recognition model, accurate text recognition results can be obtained.

[0091] In one embodiment, the correspondence between each pixel point and the text instance category in the above embodiment is obtained by a pixel segmentation model;

[0092] The construction process of the pixel segmentation model includes:

[0093] Obtaining a template sample image, a training label corresponding to the template sample image, and a model to be trained, wherein the model to be trained includes a feature extraction network and a text instance classification network;

[0094] Extract features of the printed sample image through a feature extraction network to obtain sample feature information of the printed sample image;

[0095] The text instance classification network is used to classify each sample pixel in the printed sample image according to the sample feature information, and the sample category probability of each sample pixel belonging to the sample text instance in the training label is predicted;

[0096] According to the sample category probability and training labels, the model loss function is obtained;

[0097] According to the model loss function, the training model is adjusted to obtain a pixel segmentation model.

[0098] Among them, the template sample image refers to a sample image including a template text area, and the training label corresponding to the template sample image is used to mark the correspondence between each sample pixel in the template sample image and each sample text instance in the template text area. For example, the training label can be a matrix corresponding to the size of the template sample image and the number of channels is the number of sample text instances in the template sample image, and the mask corresponding to each sample text instance is stored. The mask is a string of binary codes that performs a bitwise AND operation on the target field to shield the current input bit. In this embodiment, the mask is used to determine the correspondence between the sample pixel and the sample text instance. Specifically, when the sample pixel belongs to the sample text instance, the corresponding mask is 1, and when the sample pixel does not belong to the sample text instance, the corresponding mask is 0.

[0099] The feature extraction network includes a coordinate information convolution layer, a downsampling network, and an upsampling network, wherein the coordinate information convolution layer is used to update the pixel coordinate information of each sample pixel point to the image channel data of the printed sample image, and obtain the image data to be sampled corresponding to the printed sample image, the downsampling network is used to downsample the image data to be sampled corresponding to the printed sample image, and obtain a multi-scale downsampling feature map, and the upsampling network is used to upsample according to the multi-scale downsampling feature map, and obtain the sample feature information of the printed sample image. The data processing process of the downsampling network on the image data to be sampled corresponding to the printed sample image is the same as the process of the trained downsampling network on the image data to be sampled, which is not repeated here in this embodiment, and the data processing process of the upsampling network on the multi-scale downsampling feature map corresponding to the printed sample image is the same as the process of the trained upsampling network on the image data to be sampled, which is not repeated here in this embodiment. The text instance classification network refers to a network used to perform convolution on sample feature information to predict the sample category probability of each sample pixel belonging to the sample text instance in the training label. The process of processing the sample feature information can refer to the process of processing the feature information by the above-mentioned trained classification network.

[0100] Specifically, when text recognition is required, the terminal will obtain a sample text set, use the sample text in the sample text set to construct a sample image carrying a template and a training label corresponding to the template sample image, and obtain a model to be trained, perform feature extraction on the template sample image through the feature extraction network in the model to be trained, obtain sample feature information of the template sample image, input the sample feature information into the text instance classification network in the model to be trained, and perform text instance classification on each sample pixel in the template sample image according to the sample feature information through the text instance classification network, predict the sample category probability that each sample pixel belongs to the sample text instance in the training label, calculate the model loss function according to the sample category probability and the training label, and adjust the network parameters of the feature extraction network and the text instance classification network in the model to be trained according to the model loss function until the model loss function meets the preset model adjustment requirements, and obtain the pixel segmentation model. Among them, the preset model adjustment requirements can be specifically that the model loss function is less than the preset loss function threshold, the model loss function converges, etc., and this embodiment is not specifically limited here.

[0101] For example, the process of using the pixel segmentation model to obtain the correspondence between each pixel and the text instance category can be as follows: Figure 6As shown, the terminal obtains a text image and a pixel segmentation model, obtains image channel data of the text image and pixel coordinate information of each pixel point, uses the coordinate information convolution layer in the feature extraction network in the pixel segmentation model to splice the pixel coordinate information to the image channel data to update the image channel data, obtains the image data to be sampled, uses the downsampling network in the feature extraction network to downsample the image data to obtain a multi-scale downsampled feature map, uses the upsampling network in the feature extraction network to upsample according to the multi-scale downsampled feature map, obtains the feature information of the text image, uses the text instance classification network feature information in the pixel segmentation model to classify each pixel point in the text image into a text instance, and determines the category probability (i.e., P) of each pixel point belonging to each text instance category. 0 , P 1…… P n ), compare the category probability with the preset probability threshold (i.e., threshold), determine the correspondence between each pixel and each text instance category, and return the pixel to the text instance.

[0102] In this embodiment, by acquiring the template sample images and the training labels corresponding to the template sample images, the template sample images and the training labels corresponding to the template sample images can be used to train the model to be trained, thereby obtaining a pixel segmentation model.

[0103] In one embodiment, obtaining the template sample image and the training label corresponding to the template sample image includes:

[0104] Get a sample text set;

[0105] Rendering sample texts in the sample text set into colored text line images;

[0106] Get the text mask area corresponding to the color text line image;

[0107] Paste the text mask area into the preset background canvas according to the preset overlap degree to obtain a printed sample image, and determine the number of sample text instances according to the number of text mask areas in the printed sample image;

[0108] The training labels are obtained according to the number of sample text instances and the correspondence between the text mask area in the printed sample image and the sample text instances.

[0109] The sample text refers to independent non-printed text, that is, an independent text entry. The text mask area refers to the area where the sample text is located in the color text line image. The preset overlap refers to pasting the text mask area with a preset overlap, so that there is overlap between the text mask areas from different color text line images. The preset overlap can be set as needed. The preset background canvas refers to a preset blank canvas. The text mask area corresponds to the sample text instance, and a text mask area is a sample text instance.

[0110] Specifically, the terminal obtains a sample text set, renders the sample text in the sample text set into multiple independent color text line images, converts the color text line image into a corresponding grayscale image, uses a preset grayscale value threshold to screen the pixels in the corresponding grayscale image, selects the pixels whose grayscale values ​​are less than the preset grayscale values, obtains a text mask area corresponding to the color text line image, pastes the text mask area into a preset background canvas according to a preset overlap, obtains a printed sample image, and determines the number of sample text instances according to the number of text mask areas in the printed sample image, determines the correspondence between sample pixel points and sample text instances in the printed sample image according to the correspondence between the text mask areas and the sample text instances in the printed sample image, creates a matrix corresponding to the size of the printed sample image and with the number of channels equal to the number of sample text instances, and stores the correspondence between sample pixel points and sample text instances in the printed sample image, wherein the correspondence with the size of the printed sample image refers to the correspondence with the number of pixels in the printed sample image.

[0111] It should be noted that the present embodiment does not limit the method of rendering the sample text to obtain a color text line image. For example, the sample text can be rendered using the python pygame toolkit. In the color text line image, the image background can be pure black, and the sample text is a random color. Preferably, according to the coverage of the printed text that often appears in real application scenarios, the preset overlap range in the present embodiment is 25%-75%. When the text mask area is pasted into the preset background canvas according to the preset overlap, it is necessary to ensure that the coverage between the subsequent sample text and the previously existing sample text meets the preset overlap requirement.

[0112] In this embodiment, by obtaining a sample text set, rendering the sample text in the sample text set as a color text line image, obtaining a text mask area corresponding to the color text line image, pasting the text mask area to a preset background canvas according to a preset overlap, obtaining a template sample image, and obtaining training labels based on the number of sample text instances and the correspondence between the text mask area and the sample text instances in the template sample image, thereby realizing the acquisition of the template sample image and the training labels corresponding to the template sample image.

[0113] In one embodiment, Figure 7 As shown, a flowchart is used to illustrate the text recognition method of the present application, and the text recognition method specifically includes the following steps:

[0114] The first is text area detection. When text recognition is required, the terminal will obtain the image to be processed, perform target detection on the image to be processed, divide the image to be processed into a printed text area, an independent text area, and a background area, and extract the text image from the image to be processed based on the printed text area.

[0115] The second is pixel segmentation. After obtaining the text image, the terminal will obtain the image channel data of the text image and the pixel coordinate information of each pixel point, update the image channel data according to the pixel coordinate information, obtain the image data to be sampled, downsample the image data to be sampled, obtain the multi-scale downsampled feature map, upsample according to the multi-scale downsampled feature map, obtain the feature information of the text image, classify each pixel point in the text image into text instances according to the feature information, determine the category probability of each pixel point belonging to each text instance category, compare the category probability with the preset probability threshold, determine the correspondence between each pixel point and each text instance category, and the text instance category is an independent text entry category.

[0116] The third is to return the printed pixels to each text instance. The terminal will split the text image according to the correspondence between each pixel and each text instance category to obtain the instance text image corresponding to the text instance category.

[0117] Fourth, OCR (Optical Character Recognition) recognition. The terminal will obtain a trained text recognition model, perform text recognition on the instance text image through the trained text recognition model, and obtain the text recognition result.

[0118] The effect of the text recognition method of the present application is described below.

[0119] like Figure 4 As shown in the figure, the text image includes three text instances with duplicate printing (represented by different letters, the same letters form a text instance, the colors of different text instances are different, and the color of the background area is different from the color of the text instance (not explicitly shown in the figure)). After determining the corresponding relationship between each pixel and each text instance category through duplicate pixel segmentation, the terminal can return the duplicate pixels to each text instance to obtain the instance text image corresponding to the text instance category, as shown in FIG. Figure 8As shown, it should be noted that the background area and the text instance have different colors (not explicitly shown in the figure). After obtaining the instance text image, the text recognition result can be obtained by performing text recognition on the instance text image. Among them, after determining the correspondence between each pixel point and each text instance category, the terminal can also visualize the correspondence between each pixel point and each text instance category, that is, different colors are used to represent different text instances.

[0120] It should be understood that, although the steps in each flow chart involved in the above-described embodiment are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a portion of the steps in each flow chart involved in the above-described embodiment may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.

[0121] In one embodiment, Fig. 9 As shown, a text recognition device is provided, which can be a part of a computer device using a software module or a hardware module, or a combination of the two. The device specifically includes: an acquisition module 902, a feature extraction module 904, a classification module 906, a splitting module 908 and a recognition module 910, wherein:

[0122] An acquisition module 902 is used to acquire a text image;

[0123] The feature extraction module 904 is used to extract features from the text image to obtain feature information of the text image;

[0124] The classification module 906 is used to classify each pixel in the text image into a text instance according to the feature information, and determine the correspondence between each pixel and a text instance category, where the text instance category is an independent text entry category;

[0125] A splitting module 908 is used to split the text image according to the correspondence between each pixel point and each text instance category to obtain an instance text image corresponding to the text instance category;

[0126] The recognition module 910 is used to perform text recognition on the example text image to obtain a text recognition result.

[0127] The above-mentioned text recognition device obtains a text image, performs feature extraction on the text image, obtains feature information of the text image, classifies each pixel in the text image into text instances according to the feature information, determines the correspondence between each pixel and the text instance category, and splits the text image according to the correspondence to obtain an instance text image corresponding to the text instance category. It can classify the text instances of each pixel in the text image and split the templated text in the text image to obtain a non-templated instance text image, thereby providing high-quality data to be recognized for text recognition. Furthermore, it can obtain accurate text recognition results by performing text recognition on the instance text image, thereby improving the accuracy of text recognition.

[0128] In one embodiment, the feature extraction module is also used to obtain image channel data of the text image and pixel coordinate information of each pixel point, update the image channel data according to the pixel coordinate information, obtain the image data to be sampled, sample the image data to be sampled, and obtain feature information of the text image.

[0129] In one embodiment, the feature extraction module is further used to downsample the image data to be sampled to obtain a multi-scale downsampled feature map, and upsample according to the multi-scale downsampled feature map to obtain feature information of the text image.

[0130] In one embodiment, the classification module is also used to classify each pixel in the text image into a text instance based on the feature information, determine the category probability of each pixel belonging to each text instance category, compare the category probability with a preset probability threshold, and determine the correspondence between each pixel and each text instance category.

[0131] In one embodiment, the recognition module is further used to obtain a trained text recognition model, and perform text recognition on the instance text image through the trained text recognition model to obtain a text recognition result.

[0132] In one embodiment, the correspondence between each pixel point and the text instance category in the above embodiment is obtained through a pixel segmentation model. The device also includes a model construction module. The model construction module is used to obtain a template sample image, a training label corresponding to the template sample image, and a model to be trained. The model to be trained includes a feature extraction network and a text instance classification network. The feature extraction network is used to extract features of the template sample image to obtain sample feature information of the template sample image. The text instance classification network is used to classify each sample pixel point in the template sample image according to the sample feature information, and the sample category probability of each sample pixel point belonging to the sample text instance in the training label is predicted. According to the sample category probability and the training label, a model loss function is obtained. According to the model loss function, the model to be trained is adjusted to obtain a pixel segmentation model.

[0133] In one embodiment, the model building module is also used to obtain a sample text set, render the sample text in the sample text set as a color text line image, obtain a text mask area corresponding to the color text line image, paste the text mask area into a preset background canvas according to a preset overlap degree, obtain a printed sample image, and determine the number of sample text instances based on the number of text mask areas in the printed sample image, and obtain training labels based on the number of sample text instances and the correspondence between the text mask areas and the sample text instances in the printed sample image.

[0134] For the specific definition of the text recognition device, please refer to the definition of the text recognition method above, which will not be repeated here. Each module in the above text recognition device can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0135] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.10 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a text recognition method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball or a touch pad set on the computer device housing, or an external keyboard, touch pad or mouse, etc.

[0136] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0137] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0138] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0139] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0141] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0142] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A text recognition method, It is characterized in that The method comprises: Get text image; Extracting features of the text image using a pixel segmentation model to obtain feature information of the text image; By using the pixel segmentation model, each pixel in the text image is classified into text instance according to the feature information, and a corresponding relationship between each pixel and a text instance category is determined, wherein the text instance category is an independent text entry category; According to the correspondence between each pixel point and the text instance category, the text image is split to obtain an instance text image corresponding to the text instance category; Performing text recognition on the example text image to obtain a text recognition result; The pixel segmentation model construction process includes: Obtain a sample text set, render the sample text in the sample text set into a color text line image, obtain a text mask area corresponding to the color text line image, paste the text mask area into a preset background canvas according to a preset overlap degree, obtain a printed sample image, and determine the number of sample text instances according to the number of text mask areas in the printed sample image, and obtain training labels according to the number of sample text instances and the correspondence between the text mask areas in the printed sample image and the sample text instances; A model to be trained is obtained, and the model to be trained is trained using the printed sample images and the training labels corresponding to the printed sample images to obtain a pixel segmentation model.

2. The method according to claim 1, It is characterized in that The extracting features of the text image to obtain feature information of the text image includes: Obtaining image channel data of the text image and pixel coordinate information of each pixel point; According to the pixel coordinate information, the image channel data is updated to obtain the image data to be sampled; The image data to be sampled is sampled to obtain feature information of the text image.

3. The method according to claim 2, It is characterized in that The sampling of the image data to be sampled to obtain the feature information of the text image includes: Downsampling the image data to be sampled to obtain a multi-scale downsampling feature map; Upsampling is performed according to the multi-scale down-sampled feature map to obtain feature information of the text image.

4. The method according to claim 1, It is characterized in that The step of classifying each pixel in the text image as a text instance according to the feature information by using a pixel segmentation model, and determining a corresponding relationship between each pixel and a text instance category comprises: Using a pixel segmentation model, classifying each pixel in the text image into text instances according to the feature information, and determining the category probability that each pixel belongs to each text instance category; The category probability is compared with a preset probability threshold to determine the corresponding relationship between each pixel point and each text instance category.

5. The method according to claim 1, It is characterized in that The performing text recognition on the example text image to obtain a text recognition result includes: Get the trained text recognition model; The trained text recognition model is used to perform text recognition on the example text image to obtain a text recognition result.

6. The method according to any one of claims 1 to 5, It is characterized in that The model to be trained includes a feature extraction network and a text instance classification network; the model to be trained is trained by the printed sample image and the training label corresponding to the printed sample image to obtain a pixel segmentation model including: Extracting features of the printed sample image through the feature extraction network to obtain sample feature information of the printed sample image; Using the text instance classification network, classifying each sample pixel in the printed sample image according to the sample feature information, and predicting the sample category probability that each sample pixel belongs to the sample text instance in the training label; Obtaining a model loss function according to the sample category probability and the training label; According to the model loss function, the model to be trained is adjusted to obtain a pixel segmentation model.

7. A text recognition device, It is characterized in that The device comprises: A model building module is used to obtain a sample text set, render the sample text in the sample text set as a color text line image, obtain a text mask area corresponding to the color text line image, paste the text mask area into a preset background canvas according to a preset overlap, obtain a printed sample image, and determine the number of sample text instances according to the number of text mask areas in the printed sample image, obtain training labels according to the number of sample text instances and the correspondence between the text mask areas in the printed sample image and the sample text instances, obtain a model to be trained, train the model to be trained using the printed sample image and the training labels corresponding to the printed sample image, and obtain a pixel segmentation model; An acquisition module, used for acquiring text images; A feature extraction module, used to extract features from the text image through a pixel segmentation model to obtain feature information of the text image; A classification module, used to classify each pixel in the text image into a text instance according to the feature information through the pixel segmentation model, and determine the corresponding relationship between each pixel and a text instance category, wherein the text instance category is an independent text entry category; A splitting module, used for splitting the text image according to the correspondence between each pixel point and the text instance category to obtain an instance text image corresponding to the text instance category; The recognition module is used to perform text recognition on the example text image to obtain a text recognition result.

8. The device according to claim 7, It is characterized in that The feature extraction module is also used to obtain the image channel data of the text image and the pixel coordinate information of each pixel point, update the image channel data according to the pixel coordinate information, obtain the image data to be sampled, sample the image data to be sampled, and obtain the feature information of the text image.

9. The device according to claim 8, It is characterized in that The feature extraction module is also used to downsample the image data to be sampled to obtain a multi-scale downsampled feature map, and upsample according to the multi-scale downsampled feature map to obtain feature information of the text image.

10. The device according to claim 7, It is characterized in that The classification module is also used to classify each pixel in the text image into text instances according to the feature information through a pixel segmentation model, determine the category probability that each pixel belongs to each text instance category, compare the category probability with a preset probability threshold, and determine the correspondence between each pixel and each text instance category.

11. The device according to claim 7, It is characterized in that The recognition module is also used to obtain a trained text recognition model, and perform text recognition on the instance text image through the trained text recognition model to obtain a text recognition result.

12. The device according to any one of claims 7 to 11, It is characterized in that The model to be trained includes a feature extraction network and a text instance classification network; the model construction module is also used to perform feature extraction on the template sample image through the feature extraction network to obtain sample feature information of the template sample image, and perform text instance classification on each sample pixel in the template sample image according to the sample feature information through the text instance classification network, and predict the sample category probability that each sample pixel belongs to the sample text instance in the training label, and obtain a model loss function based on the sample category probability and the training label, and adjust the model to be trained based on the model loss function to obtain a pixel segmentation model.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

14. A computer-readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Character recognition method and device for image, electronic equipment and readable storage medium

    CN111476067A

  • Model training method, image processing method and device, computer system and medium

    CN111723815A