Text recognition method, device, computer equipment and storage medium
By using a shared feature extraction scheme and specific dimensional transformation processing, the problems of large model size and high training difficulty in horizontal and vertical text recognition were solved, thus improving the accuracy of text recognition.
Patent Information
- Application Number
- CN202010834389.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-08-19
AI Technical Summary
Existing methods for recognizing horizontal and vertical text have large model sizes and are difficult to train. The number of vertical text samples is insufficient, which affects the accuracy of text recognition.
A shared feature extraction scheme is adopted to perform specific dimensional transformation processing on horizontal and vertical text line images, converting the feature map into a two-dimensional matrix, and analyzing it according to the writing direction to improve recognition accuracy.
It simplifies the model size and training difficulty, and improves the accuracy of horizontal and vertical text recognition.
Smart Images

Figure CN114078250B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a text recognition method, device, computer equipment and storage medium. Background Art
[0002] Currently, based on the writing direction of text lines in an image, text recognition for an image includes recognition of vertical text and horizontal text.
[0003] In view of the different writing directions, existing horizontal and vertical text recognition methods can be mainly divided into two categories. The first category is based on character detection, and the second category is based on the overall recognition of text lines.
[0004] For methods based on overall text line recognition, vertical text can be rotated 90 degrees counterclockwise and mixed with horizontal text to train the model. After training, the model can be used to recognize both vertical and horizontal text. However, in this solution, the model needs to learn the features of each character in both upright and sideways modes, so the model scale and training difficulty will be larger. In addition, the number of vertical text samples is generally lower than that of horizontal text, which is not conducive to ensuring the model's text recognition accuracy. Summary of the Invention
[0005] Embodiments of the present invention provide a text recognition method, apparatus, computer equipment, and storage medium. Horizontal and vertical texts can share a common feature extraction scheme, which is beneficial for simplifying the model size and training difficulty, avoiding the problem of uneven amounts of horizontal and vertical texts, and improving text recognition accuracy.
[0006] An embodiment of the present invention provides a text recognition method, which includes:
[0007] Get the text line image to be recognized;
[0008] Acquire the writing direction of the to-be-recognized text line image, wherein the types of the writing direction include horizontal writing direction and vertical writing direction;
[0009] Extracting a feature map of the to-be-recognized text line image based on a trained text recognition model;
[0010] Performing a specific dimensional transformation process on the feature map by the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0011] The two-dimensional matrix is analyzed by the text recognition model to determine the text in the text line image to be recognized.
[0012] On the other hand, the present application also provides a text recognition device, comprising:
[0013] A first acquiring unit, configured to acquire an image of a text line to be recognized;
[0014] A second acquiring unit is configured to acquire a writing direction of the to-be-recognized text line image, wherein the types of the writing direction include a horizontal writing direction and a vertical writing direction;
[0015] A feature acquisition unit, configured to extract a feature map of the text line image to be recognized based on a trained text recognition model;
[0016] a conversion unit, configured to perform a specific dimensional transformation process on the feature map using the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; and if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0017] An analyzing unit is configured to analyze the two-dimensional matrix using the text recognition model to determine the text in the to-be-recognized text line image.
[0018] On the other hand, the present application further provides a storage medium having a computer program stored thereon, wherein the computer program implements the steps of the text recognition method described above when executed by a processor.
[0019] On the other hand, the present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the text recognition method described above when executing the program.
[0020] The present embodiment discloses a text recognition method, apparatus, computer equipment and storage medium, which can obtain a text line image to be recognized; obtain the writing direction of the text line image to be recognized, wherein the types of writing directions include horizontal writing direction and vertical writing direction; extract a feature map of the text line image to be recognized based on a trained text recognition model; in the present embodiment, both horizontal text lines and vertical text lines use a feature extraction scheme to extract feature maps, so the feature extraction capability of the model is relatively strong and the scale of the model can also be controlled; after the feature map is extracted, the feature map is subjected to a specific dimension transformation processing by the text recognition model to convert the feature map into a two-dimensional matrix, and the specific dimension processing corresponds to the writing direction, wherein, if the writing direction is a horizontal writing direction, the dimension of the two-dimensional matrix includes the horizontal dimension of the first image and the channel dimension of the first image; if the writing direction is a vertical writing direction, the dimension of the two-dimensional matrix includes the vertical dimension of the second image and the channel dimension of the second image; the two-dimensional matrix is analyzed by the text recognition model to determine the text in the text line image to be recognized, and the model in the present embodiment has a strong feature extraction capability for horizontal text lines and vertical text lines, which is conducive to improving the accuracy of text recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 1 is a flowchart of a text recognition method provided by an embodiment of the present invention;
[0023] Figure 2a 1 is a flow chart of a method for training a text recognition model provided by an embodiment of the present invention;
[0024] Figure 2b is a schematic diagram of a horizontal text line image and a vertical text line image provided by an embodiment of the present invention;
[0025] Figure 2c Schematic diagram of a text recognition model without an RNN layer provided in an embodiment of the present invention for recognizing horizontal and vertical text line images;
[0026] Figure 2d Schematic diagram of the principle of the text recognition model training method provided by an embodiment of the present invention;
[0027] Figure 3 is a structural diagram of a text recognition device provided by an embodiment of the present invention;
[0028] Figure 4It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0031] In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is given to enable any person skilled in the art to make and use the invention. In the following description, details are listed for the purpose of explanation. It should be understood that one of ordinary skill in the art will recognize that the invention can be practiced without these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0032] This embodiment discloses a text recognition method, apparatus, storage medium, and computer device, which are described in detail below.
[0033] First, an embodiment of the present invention provides a text recognition method, which includes: obtaining a text line image to be recognized; obtaining the writing direction of the text line image to be recognized, wherein the types of the writing directions include horizontal writing directions and vertical writing directions; extracting a feature map of the text line image to be recognized based on a trained text recognition model; performing a specific dimensional transformation processing on the feature map through the text recognition model to convert the feature map into a two-dimensional matrix, and the specific dimensional transformation processing corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension, and if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension; analyzing the two-dimensional matrix through the text recognition model to determine the text in the text line image to be recognized.
[0034] This embodiment provides a text recognition method applicable to a computer device. The computer device may be a terminal or other device, such as a mobile phone, tablet computer, laptop computer, desktop computer, etc. The computer device may also be a server or other device, and the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited thereto.
[0035] For example, the text recognition method can be integrated into the terminal or server.
[0036] like Figure 1 FIG2 is a flow chart of an embodiment of a text recognition method in an embodiment of the present invention. The text recognition method is applied to a text recognition device. The text recognition device can be integrated into a computer device, for example, integrated into a terminal or a server. This embodiment has no limitation on this. Figure 1 , the text recognition method may include:
[0037] 101. Obtain an image of a text line to be recognized;
[0038] For images that require text recognition, they may include one or more lines of text (written vertically or horizontally). If the image to be recognized contains a line of text, then the image to be recognized is a text line image to be recognized. If the image to be recognized contains multiple lines of text, it is necessary to obtain a text line image through processing.
[0039] In an optional embodiment, the step of “obtaining an image of a text line to be recognized” may include:
[0040] Obtain the image to be recognized;
[0041] If the text in the image to be recognized is not distributed in rows, preprocess the image to be recognized so that the text in the image to be recognized is distributed in rows;
[0042] If the image to be identified distributed by lines includes multiple lines of text, the preprocessed image to be identified is segmented into lines to obtain at least one text line image;
[0043] Select a text line image to be recognized from the text line images.
[0044] When the text in the image to be recognized is not distributed in rows, the preprocessing may include: selecting the text area in the image to be recognized using a polygonal frame, and reshaping the text in the polygonal frame into linear row text.
[0045] For an image to be recognized where text is distributed in lines, the image to be recognized can be segmented into lines to obtain at least one text line image. For example, a quadrilateral is first used to frame the text line of the image to be recognized, and then the text line is segmented into lines to obtain a text line image.
[0046] 102. Acquire a writing direction of the text line image to be recognized, where the types of writing directions include horizontal writing direction and vertical writing direction;
[0047] In this embodiment, the writing direction of the text line image to be identified can be a horizontal writing direction or a vertical writing direction. If it is impossible to determine whether the writing direction of the text line image to be identified is the horizontal writing direction or the vertical writing direction, it can be assumed that the writing direction of the text line image to be identified includes both the horizontal writing direction and the vertical writing direction. Subsequently, text recognition is performed on the text line image to be identified according to both the horizontal writing direction and the vertical writing direction, and then the correct text recognition result is selected from the recognition results.
[0048] In this embodiment, in order for the text recognition model to determine the writing direction of the text line image to be recognized, the writing direction of the text line image to be recognized can be marked after step 102, for example, a label can be set for the text line image to be recognized, and the label records the writing direction identifier of the text line image to be recognized, wherein the writing direction identifier can be expressed in any form, for example, the writing direction identifier is defined as a parameter V, and the parameter V has two values, False and True, V=False, indicating a horizontal writing direction, and V=Ture, indicating a vertical writing direction.
[0049] Optionally, the step of “obtaining the writing direction of the text line image to be recognized” may include:
[0050] Determining whether the actual length relationship between the horizontal length and the vertical length of the text line image to be recognized matches the preset length relationship corresponding to the horizontal writing direction or the vertical writing direction;
[0051] Based on the judgment result, the writing direction of the text line image to be recognized is determined.
[0052] The length relationship may be a ratio of the horizontal length to the vertical length of the text line image.
[0053] For example, the horizontal length (width) and vertical length (height) of the text line image Iu to be recognized can be Wu and Hu respectively, the preset length relationship corresponding to the horizontal writing direction is O1, and the preset length relationship corresponding to the vertical writing direction is O2.
[0054] in,
[0055] If O1 is satisfied, the text line image to be recognized is text written horizontally (i.e., horizontal text); if O2 is satisfied, the text line image to be recognized is text written vertically (i.e., vertical text). Of course, it is understandable that the ratio "2" in the above preset correspondence can be set to other values as needed, such as 1.5, 1.8, 3, or any other value.
[0056] In this embodiment, before the text line image to be recognized is recognized by the text recognition model, the text line image to be recognized can also be resized so that the height of the text line image to be recognized in the horizontal writing direction in the input text recognition model is equal to the width of the text line image to be recognized in the vertical writing direction. The specific values of the height of the text line image to be recognized in the horizontal writing direction and the width of the text line image to be recognized in the vertical writing direction can be determined according to the input requirements of the text recognition model. This embodiment has no limitation on this. The resizing process can make the sizes of the text line images in the horizontal and vertical writing directions meet the requirements of the text recognition model.
[0057] For example, define the following operations:
[0058] O1: Set v = False, scale I proportionally u Make the image height 32
[0059] O2: Set v = True, scale I proportionally u Make the image width 32.
[0060] Of course, it is understandable that the above value "32" can also be determined according to actual conditions, such as the size of the text in the text line image to be recognized.
[0061] 103. Extracting a feature map of the text line image to be recognized based on the trained text recognition model;
[0062] To facilitate understanding, the training process of the text recognition model is first illustrated with an example.
[0063] Before step 101, you can first use Figure 2a The model training method shown is used to train the text recognition model. Figure 2a , the training methods of the text recognition model include:
[0064] 201. Obtain training samples, where the training samples include horizontal text samples and vertical text samples, and sample labels of the training samples include the writing direction of the text in the samples and character identifiers corresponding to the characters in the text;
[0065] In this embodiment, step 201 includes two stages: data preparation and data processing.
[0066] In the data preparation stage, the original data is first obtained. The original data includes horizontal text line images and vertical text line images, that is, images with a line of text written horizontally or vertically. Among them, the text area occupies the vast majority of the image, such as Figure 2b shown.
[0067] During the data processing phase, the raw data can be processed, resulting in a set of augmented images and corresponding annotated data. The data can be annotated by using a string to record the text content in the image and a numerical value to indicate the direction of the text, for example, 0 for horizontal text and 1 for vertical text.
[0068] Specifically, the data processing stage can include three parts: label mapping, input normalization and data augmentation.
[0069] Optional processing of raw data includes:
[0070] Marking mapping step: based on a preset mapping relationship between characters and character identifiers, obtaining character identifiers corresponding to characters in the text in the image of the original data;
[0071] Input normalization step: scale the image in the original data so that the height of the horizontal text line image is the same as the width of the vertical text line image;
[0072] Data augmentation step: performing data augmentation on the scaled image to obtain an augmented image;
[0073] The augmented images are used as training samples, and sample labels are added to the training samples. The sample labels include the character identifiers corresponding to the images and the writing direction.
[0074] After the scaling process, the height of the horizontal text line image and the width of the vertical text line image can be unified to the same number of pixels, for example, 32 pixels.
[0075] In this embodiment, a correspondence between characters and character identifiers may be set, for example, a correspondence between Chinese characters and character identifiers may be set, wherein the character identifiers may be represented by numerical values, for example, by positive integer values of 1, 2, 3, ..., n.
[0076] Through label mapping, the character string corresponding to the text in the image can be converted into a numerical sequence that can be input into the text recognition model, that is, a word table Vob with a length of T is established, and each character in the string is mapped to a serial number in the word table with a value range of [0, T-1]. Characters that are not in the word table are called out-of-table characters or unknown characters, and can be uniformly mapped to T.
[0077] When performing input normalization, an interpolation method may be used to proportionally scale the image so that the height of the horizontal text line image and the width of the vertical text line image are unified to 32 pixels.
[0078] Data augmentation is a method of expanding data. By augmenting data, the generalization ability of the model can be improved, and the accuracy of the model prediction can be improved to a certain extent. The data augmentation methods used in this embodiment include but are not limited to: perspective transformation, Gaussian blur, noise addition, and HSV (Hue, Saturation, Value) channel color transformation. Furthermore, by randomly selecting and combining data augmentation methods, ten times the amount of data can be obtained.
[0079] 202. Obtain a text recognition model to be trained;
[0080] In this embodiment, a text recognition model can be established based on training samples. The text recognition model may include: a feature extraction layer, a data transformation layer, a matrix processing layer, a classification layer, and a loss calculation layer.
[0081] Among them, the matrix processing layer can be implemented based on RNN (Recurrent Neural Network), and the loss calculation layer can be a CTC (Connectionist Temporal Classification) layer.
[0082] 203. Extract feature maps from training samples based on the text recognition model;
[0083] This embodiment can extract feature maps of training samples based on the feature extraction layer.
[0084] The input of the text recognition model is a normalized image I (assuming its size is H*W*C, where H and W are the height and width respectively, and C is the number of image channels, usually 1 for grayscale images and 3 for color images), an image annotation Label (a character identifier, i.e., a sequence of character numbers in the character list of the text string in the image), and a Boolean parameter v indicating whether the text on it is horizontal or vertical:
[0085]
[0086] The feature extraction layer in this embodiment can be a convolutional neural network, which can be composed of a convolution layer, a pooling layer, a batch normalization layer, etc. through adjacent layer connections or skip connections.
[0087] In one example, the feature extraction layer can downsample the input training sample by 8 times (or other multiples) in both height and width directions to obtain a feature map F. That is, the input of the feature extraction layer is the training sample image I, and the output is a three-dimensional feature map F, whose shape is [H f ,W f ,C f ]1(H f =H / 8,W f =W / 8, Cf is the number of feature channels). Among them, Hf can be understood as the vertical dimension of the original image of the feature map, Wf can be understood as the horizontal dimension of the original image of the feature map, and Cf can be understood as the channel dimension of the original image of the feature map.
[0088] From the above, when v=False, H f =4; when v=True, W f =4.
[0089] In this embodiment, notations such as [w1, w2, w3] represent the shape of a matrix. For example, the matrix [w1, w2, w3] has a dimension of 3, with the lengths of each dimension being w1, w2, and w3, respectively. Here, w1, w2, and w3 are algebraic notations representing integer values. Similar notations can be used analogously.
[0090] The above feature extraction layer is shared by horizontal text samples and vertical text samples.
[0091] 204. Performing a specific dimensional transformation process on the feature map using the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the training sample, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; and if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0092] The specific dimension transformation processing of this embodiment is implemented by the data transformation layer.
[0093] The specific dimensional transformation processing of this embodiment is to perform dimensional processing on the three-dimensional feature map, and convert the three-dimensional feature map into a two-dimensional matrix. The specific dimensional transformation processing corresponds to the writing direction of the training sample, which means that if the writing direction is different, the dimensions of the merged images during the specific dimensional transformation processing are different, and the dimensions of the merged two-dimensional matrix are different. For the specific dimensional transformation processing scheme, please refer to the following description.
[0094] Optionally, the step of "performing a specific dimension transformation on the feature map using the text recognition model to convert the feature map into a two-dimensional matrix" may include:
[0095] If the writing direction is a horizontal writing direction, merging the original image longitudinal dimension and the original image channel dimension of the feature map through the text recognition model to obtain a first image channel dimension, and using the original image transverse dimension as the first image transverse dimension to obtain a first two-dimensional matrix including the first image transverse dimension and the first image channel dimension;
[0096] If the writing direction is vertical writing, the original image horizontal dimension and the original image channel dimension of the feature map are merged through the text recognition model to obtain a second image channel dimension, and the original image vertical dimension is used as the second image vertical dimension to obtain a second two-dimensional matrix including the second image vertical dimension and the second image channel dimension.
[0097] Among them, the order of the dimensions of the feature map is the vertical dimension of the original image, the horizontal dimension of the original image, and the channel dimension of the original image. If you want to merge the vertical dimension of the original image and the channel dimension of the original image, you can first swap the vertical dimension of the original image and the horizontal dimension of the original image of the feature map, and then merge the vertical dimension of the original image and the channel dimension of the original image.
[0098] Optionally, the data processing of the data transformation layer includes dimension swapping and dimension merging. The input of the data transformation layer is the feature map F and parameter v output by the feature extraction layer, and the output is a two-dimensional matrix M. The data processing of the data transformation layer does not change the total number of values in the feature map (the total number of values is H f *W f *C f ), only the shape of the matrix and the order of the internal values will be changed.
[0099] Among them, dimension swapping changes the permutation order but does not change the total number of dimensions. For example, swapping the two dimensions of a 2*3 matrix [[1,2,3],[4,5,6]] results in a 3*2 matrix [[1,4],[2,5],[3,6]].
[0100] Dimension merging does not change the permutation order and merges adjacent dimensions into one dimension. For example, merging the two dimensions of a 3*2 matrix [[1,4],[2,5],[3,6]] results in [1,4,2,5,3,6]. From the perspective of shape, its transformation rules are as follows:
[0101]
[0102] 205. Analyze based on the two-dimensional matrix through the text recognition model to determine the predicted character identifier corresponding to the text in the training sample;
[0103] In this example, the structure of the text recognition model may include a matrix processing layer or may not include a matrix processing layer. This embodiment places no restrictions on this.
[0104] For the solution where the text recognition model includes a matrix processing layer, the step of "analyzing based on the two-dimensional matrix through the text recognition model to determine the predicted character identifier corresponding to the text in the training sample" may include:
[0105] If the writing direction is the horizontal writing direction, expand the first two-dimensional matrix by the horizontal dimension of the first picture through the text recognition model to obtain the first column vector; if the writing direction is the vertical writing direction, expand the second two-dimensional matrix by the vertical dimension of the second picture through the text recognition model to obtain the first column vector;
[0106] Use the first column vector as the time series vector and learn the time series features of the first column vector through the matrix processing layer of the text recognition model to obtain the second column vector;
[0107] Classify the second column vector (through the classification layer) and determine the predicted character identifier corresponding to the text line in the training sample based on the classification result.
[0108] The predicted character identifier in this embodiment refers to the character identifier predicted by the text recognition model for each character in the training sample. It can be understood that in the classification result, for a character such as "文" in the training sample, the text recognition model may predict multiple character identifiers, and the classification result also includes the prediction probability of each character identifier.
[0109] For solutions where the text recognition model does not include a matrix processing layer, the step of "using the text recognition model to analyze based on the two-dimensional matrix to determine the predicted character identifiers corresponding to the text in the training sample" may include:
[0110] If the writing direction is a horizontal writing direction, the first two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector;
[0111] Classify the first column vector (via the classification layer) and determine the predicted character identifier for the text line in the training example based on the classification result.
[0112] Among them, the matrix processing layer can be an RNN network layer, and the input is a two-dimensional matrix M (output by the data transformation layer), which is expanded into a vector sequence (first column vector) of dimension Ct according to its first dimension (for images in the horizontal writing direction, the first dimension of the two-dimensional matrix is the horizontal dimension of the first image, and for images in the vertical writing direction, the first dimension of the two-dimensional matrix is the vertical dimension of the second image). This sequence is regarded as the time series input in the RNN layer. At each moment, the RNN network outputs a vector of dimension Cr, and finally obtains a new column vector R (second column vector) with a column length of W. f (when the text line is horizontal) or H f (when text lines are vertical).
[0113] refer to Figure 2c The classification layer of this embodiment can be a fully connected layer (refer to Figure 2c The FC layer in the algorithm has an input dimension of Cr (using the matrix processing layer) or Ct (not using the matrix processing layer), and an output dimension of T+2, corresponding to the classification cases of characters in the word list (number T, corresponding to category numbers in the range [0, T-1]), unknown characters (corresponding to category number T), and blanks or non-characters (denoted as CTC_Blank, corresponding to category number T+1). Each vector in R (if the RNN layer is not used, the column vector obtained by expanding the matrix M along its first dimension is used as R) is passed through the classification layer to obtain a column of classification vectors P with dimension T+2.
[0114] 206. Calculate the total loss of the text recognition model based on the predicted character identifiers of the training samples and the character identifiers in the sample labels;
[0115] In this embodiment, step 206 can be implemented by a loss calculation layer, which can be a CTC (Connectionist Temporal Classification) layer, which calculates the CTC loss. Its input is a column of classification vectors P obtained by the classification layer and the sample labels of the training samples, i.e., the labels in this embodiment. The output is a numerical value CTC_loss (Connectionist Temporal Classification loss), which represents the CTC loss. In this embodiment, CTC_loss can be used to derive each model variable and optimize the model parameters.
[0116] Taking into account the number of horizontal text samples and vertical text samples, in this embodiment, the losses corresponding to the horizontal and vertical text samples can be calculated respectively by two processing modules.
[0117] Optionally, in this embodiment, the step of “calculating the total loss of the text recognition model based on the predicted character identifiers of the training samples and the character identifiers in the sample labels” may include:
[0118] Calculating, by a first processing module, a first loss corresponding to the horizontal text sample based on the character identifier and the predicted character identifier of the horizontal text sample in the training sample;
[0119] Calculating, by a second processing module, a second loss corresponding to the vertical text sample based on the character identifiers and the predicted character identifiers of the vertical text sample in the training sample, wherein the number of horizontal text samples involved in calculating the first loss is the same as the number of vertical text samples involved in calculating the second loss;
[0120] The first loss and the second loss are summed to get the total loss of the text recognition model.
[0121] The first processing module and the second processing module are implemented by different GPUs (Graphics Processing Units).
[0122] For example, reference Figure 2d The model training principle diagram is shown in the figure. GPU1 and GPU2 are used to calculate horizontal text samples and vertical text samples respectively. After the text recognition model is input to obtain the prediction results and the CTC_loss of the character identifier in the label, the two CTC_loss are combined. Figure 2d The CTC_loss1 and CTC_loss2 in the above example are added together as the total loss to update the text recognition model. The number of horizontal text samples used to calculate the first loss is the same as the number of vertical text samples used to calculate the second loss.
[0123] Based on this approach, the number of samples from the horizontal and vertical rows of each loss is the same, and different computational graphs are executed on the two GPUs, saving the time and storage space consumed by conditional judgment.
[0124] 207. Adjust the parameters of the text recognition model based on the total loss until the text recognition model training is completed.
[0125] In this embodiment, the condition for completing text recognition model training may be that the number of training times is not less than a preset training number threshold, or that the total loss is less than a preset loss threshold, or that the first loss is less than a first loss threshold and the second loss is less than a second loss threshold. This embodiment is not limited to this.
[0126] 104. Performing a specific dimensional transformation process on the feature map using the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is horizontal, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; if the writing direction is vertical, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0127] Optionally, the dimensions of the feature map of this embodiment include the vertical dimension of the original image, the horizontal dimension of the original image, and the channel dimension of the original image. The step of "performing a specific dimension transformation process on the feature map using the text recognition model to convert the feature map into a two-dimensional matrix" may include:
[0128] If the writing direction is a horizontal writing direction, the original image vertical dimension and the original image channel dimension of the feature map are merged through the text recognition model to obtain a first image channel dimension, and the original image horizontal dimension is used as the first image horizontal dimension to obtain a first two-dimensional matrix including the first image horizontal dimension and the first image channel dimension.
[0129] If the writing direction is a vertical writing direction, the original image horizontal dimension and the original image channel dimension of the feature map are merged through the text recognition model to obtain a second image channel dimension, and the original image vertical dimension is used as the second image vertical dimension to obtain a second two-dimensional matrix including the second image vertical dimension and the second image channel dimension.
[0130] The specific process of the dimensionality transformation processing in this embodiment can refer to the dimensionality transformation processing described in the above training process, and will not be repeated here.
[0131] 105. Analyze the two-dimensional matrix through the text recognition model to determine the text in the text line image to be recognized.
[0132] In this embodiment, there are at least two analysis schemes for the two-dimensional matrix.
[0133] Optionally, in one example, the step of “analyzing the two-dimensional matrix using a text recognition model to determine the text in the text line image to be recognized” may include:
[0134] If the writing direction is a horizontal writing direction, the first two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector;
[0135] Using the first column vector as a time series vector, learning the time series features of the first column vector through the text recognition model to obtain a second column vector;
[0136] The second column vector is classified, and the text corresponding to the text line in the to-be-recognized text line image is determined based on the classification result.
[0137] The step of “classifying the second column vector and determining the text corresponding to the text line in the text line image to be recognized based on the classification result” may include:
[0138] Classify the second column vector using the text recognition model to obtain the character identifier corresponding to the text in the text line image to be recognized;
[0139] Based on the preset mapping relationship between the character identifier and the character, the text corresponding to the text line image to be recognized is obtained.
[0140] Optionally, in another example, the step of “analyzing the text in the to-be-recognized text line image using a text recognition model based on a two-dimensional matrix” may include:
[0141] If the writing direction is a horizontal writing direction, the first two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector;
[0142] The first column vector is classified, and the text corresponding to the text line in the text line image to be recognized is determined based on the classification result.
[0143] The step of “classifying the first column vector and determining the text corresponding to the text line in the text line image to be recognized based on the classification result” may include:
[0144] Classify the first column vector using the text recognition model to obtain the character identifier corresponding to the text in the text line image to be recognized;
[0145] Based on the preset mapping relationship between the character identifier and the character, the text corresponding to the text line image to be recognized is obtained.
[0146] Among them, the process of the text recognition model obtaining the character identifier corresponding to the text of the text line image to be recognized based on the two-dimensional matrix can refer to the above-mentioned model training process and will not be repeated here.
[0147] In this embodiment, when actually recognizing text, based on the preset length relationship corresponding to the above-mentioned horizontal writing direction and the vertical writing direction, the writing direction of the text line image to be recognized may not be determined. At this time, it can be assumed that the writing direction of the text line image to be recognized includes both the horizontal writing direction and the vertical writing direction, and two classification results of the text line image to be recognized are obtained based on the text recognition model, and then the correct text is selected.
[0148] Optionally, in one example, the step of “determining the writing direction of the text line image to be recognized based on the judgment result” may include:
[0149] If the actual length relationship matches the preset length relationship corresponding to the horizontal writing direction, determining that the writing direction of the to-be-recognized text line image is the horizontal writing direction;
[0150] If the actual length relationship matches the preset length relationship corresponding to the vertical writing direction, determining that the writing direction of the to-be-recognized text line image is the vertical writing direction;
[0151] If the actual length relationship does not match the preset length relationship corresponding to the horizontal writing direction and the vertical writing direction, the writing direction of the to-be-recognized text line image is set to include both the horizontal writing direction and the vertical writing direction.
[0152] Correspondingly, after the step of "analyzing the two-dimensional matrix using a text recognition model to determine the text in the text line image to be recognized", the following steps may also be included:
[0153] If the writing direction of the text line image to be recognized includes both a horizontal writing direction and a vertical writing direction, the correct text of the text line image to be recognized is determined from the text recognized by the text recognition model in the scenario where the horizontal writing direction is used as the writing direction of the text line image to be recognized, and the text recognized by the text recognition model in the scenario where the vertical writing direction is used as the writing direction of the text line image to be recognized.
[0154] Among them, in order to select the correct text, after obtaining the character identifier corresponding to the text of the text line image to be recognized based on the two-dimensional matrix, the confidence levels of the character identifiers in the two recognition results can be obtained, and the text corresponding to the character identifier with the larger confidence level can be determined as the correct text.
[0155] Alternatively, after obtaining the character identifier corresponding to the text of the text line image to be recognized based on the two-dimensional matrix, the character identifiers can be mapped to a line of text according to their positions, and the correct text can be determined based on the semantics of the text.
[0156] By adopting this embodiment, there is no need to rotate the vertical text line image, and the feature extraction layer trained in the horizontal direction can also be applied in the vertical direction. In this way, only a small number of vertical text line images are required for training to achieve good results. This reduces the requirements for training samples and is also conducive to controlling the scale of the model. In the related technology, for each character model classifier, it needs to learn its upright and upside-down two modes, which increases the difficulty of model training and reduces the classification accuracy of the model. And for some characters such as "一", it is easy to be confused with the upright mode of the character "1" after being upside down, and the error probability is extremely high. The method proposed in this embodiment avoids such drawbacks and is conducive to improving the recognition accuracy.
[0157] To better implement the above method, correspondingly, the embodiment of the present invention also provides another text recognition device. Refer to Figure 3 , the text recognition device includes:
[0158] A first acquisition unit 301, configured to acquire a text line image to be recognized;
[0159] A second acquisition unit 302, configured to acquire the writing direction of the text line image to be recognized, where the types of the writing direction include a horizontal writing direction and a vertical writing direction;
[0160] A feature acquisition unit 303, configured to extract a feature map of the text line image to be recognized based on a trained text recognition model;
[0161] A conversion unit 304, configured to perform a specific dimension transformation process on the feature map through the text recognition model, and convert the feature map into a two-dimensional matrix. The specific dimension transformation process corresponds to the writing direction of the text line image to be recognized. Wherein, if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first picture horizontal dimension and a first picture channel dimension; if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second picture vertical dimension and a second picture channel dimension;
[0162] An analysis unit 305, configured to analyze the two-dimensional matrix through a text recognition model to determine the text in the text line image to be recognized.
[0163] In some examples, a conversion unit is configured to:
[0164] If the writing direction is a horizontal writing direction, combining the original image longitudinal dimension and the original image channel dimension of the feature map into a first image channel dimension through the text recognition model, and using the original image transverse dimension as the first image transverse dimension to obtain a first two-dimensional matrix including the first image transverse dimension and the first image channel dimension;
[0165] If the writing direction is a vertical writing direction, the original image horizontal dimension and the original image channel dimension of the feature map are merged into a second image channel dimension through the text recognition model, and the original image vertical dimension is used as the second image vertical dimension to obtain a second two-dimensional matrix including the second image vertical dimension and the second image channel dimension.
[0166] In some examples, the analysis unit is configured to:
[0167] If the writing direction is a horizontal writing direction, the first two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector;
[0168] The first column vector is used as the time series vector, and the time series features of the first column vector are learned through the text recognition model to obtain the second column vector;
[0169] The second column vector is classified, and the text corresponding to the text line in the text line image to be recognized is determined based on the classification result.
[0170] In some examples, the analysis unit is configured to:
[0171] If the writing direction is a horizontal writing direction, the two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector;
[0172] The first column vector is classified, and the text corresponding to the text line in the text line image to be recognized is determined based on the classification result.
[0173] In some examples, the second acquisition unit is configured to:
[0174] Determining whether the actual length relationship between the horizontal length and the vertical length of the text line image to be recognized matches the preset length relationship corresponding to the horizontal writing direction or the vertical writing direction;
[0175] Based on the judgment result, the writing direction of the text line image to be recognized is determined.
[0176] In some examples, the second acquisition unit is configured to:
[0177] If the actual length relationship matches the preset length relationship corresponding to the horizontal writing direction, determining that the writing direction of the to-be-recognized text line image is the horizontal writing direction;
[0178] If the actual length relationship matches the preset length relationship corresponding to the vertical writing direction, determining that the writing direction of the to-be-recognized text line image is the vertical writing direction;
[0179] If the actual length relationship does not match the preset length relationship corresponding to the horizontal writing direction and the vertical writing direction, setting the writing direction of the to-be-recognized text line image to include both the horizontal writing direction and the vertical writing direction;
[0180] The device of this embodiment also includes: a selection unit, which is used to analyze the two-dimensional matrix through a text recognition model to determine the text in the text line image to be recognized. If the writing direction of the text line image to be recognized includes both a horizontal writing direction and a vertical writing direction, determine the correct text of the text line image to be recognized from the text recognized by the text recognition model in the scenario where the horizontal writing direction is used as the writing direction of the text line image to be recognized, and the text recognized by the text recognition model in the scenario where the vertical writing direction is used as the writing direction of the text line image to be recognized.
[0181] In some examples, the text recognition apparatus further includes a training unit configured to:
[0182] Get the text recognition model to be trained;
[0183] Extracting a feature map from the training sample based on the text recognition model;
[0184] Performing a specific dimensional transformation process on the feature map by the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the training sample, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; and if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0185] Determining predicted character identifiers corresponding to text in the training sample by analyzing the text recognition model based on the two-dimensional matrix;
[0186] Calculate the total loss of the text recognition model based on the predicted character identifications and the character identifications in the sample labels of the training samples;
[0187] Adjust the parameters of the text recognition model based on the total loss until the text recognition model is trained.
[0188] In some examples, the training unit is configured to:
[0189] Through the first processing module, calculate the first loss corresponding to the horizontal text samples based on the character identifications and the predicted character identifications in the horizontal text samples of the training samples;
[0190] Through the second processing module, calculate the second loss corresponding to the vertical text samples based on the character identifications and the predicted character identifications in the vertical text samples of the training samples;
[0191] Sum the first loss and the second loss to obtain the total loss of the text recognition model.
[0192] Optionally, the number of horizontal text samples participating in the calculation of the first loss is the same as the number of vertical text samples participating in the calculation of the second loss.
[0193] By adopting this embodiment, there is no need to rotate the vertical text line images, and the feature extraction layer trained horizontally can also be applied vertically. In this way, only a small number of vertical text line images are needed for training to achieve good results. This reduces the requirements for training samples and is also beneficial to controlling the scale of the model. In the related art, for each character model classifier, it needs to learn its upright and inverted two modes, which increases the difficulty of model training and reduces the classification accuracy of the model; and for some characters such as "一", it is easy to be confused with the upright mode of the character "1" after being inverted, and the error probability is extremely high. The method proposed in this embodiment avoids such drawbacks and is beneficial to improving the recognition accuracy.
[0194] The embodiment of the present invention also provides a computer device, which integrates any text recognition device provided by the embodiment of the present invention. The computer device includes:
[0195] One or more processors;
[0196] A memory; and
[0197] One or more applications, wherein one or more applications are stored in the memory and are configured to be executed by the processor to perform the steps in the text recognition method in any one of the embodiments of the above text recognition method embodiments.
[0198] The embodiment of the present invention also provides a computer device, which integrates any text recognition device provided by the embodiment of the present invention. As Figure 4, which shows a schematic diagram of the structure of a computer device involved in an embodiment of the present invention, specifically:
[0199] The computer device may include one or more processing core processors 401, one or more computer readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will understand that Figure 4 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0200] Processor 401 is the control center of the computer device. It connects the various components of the entire computer device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 402 and accessing data stored in memory 402, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, processor 401 may include one or more processing cores; preferably, processor 401 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.
[0201] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and text recognition by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0202] The computer device also includes a power supply 403 for supplying power to various components. Preferably, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0203] The computer device may further include an input unit 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0204] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 401 in the computer device loads the executable files corresponding to one or more application processes into the memory 402 according to the following instructions, and the processor 401 runs the application stored in the memory 402 to implement various functions.
[0205] For example, if the computer device is the above-mentioned text recognition device, when the application program in the computer device is executed, the following steps are implemented:
[0206] Get the text line image to be recognized;
[0207] Acquire the writing direction of the to-be-recognized text line image, wherein the types of the writing direction include horizontal writing direction and vertical writing direction;
[0208] Extracting a feature map of the to-be-recognized text line image based on a trained text recognition model;
[0209] Performing a specific dimensional transformation process on the feature map by the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0210] The two-dimensional matrix is analyzed by the text recognition model to determine the text in the text line image to be recognized.
[0211] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0212] To this end, an embodiment of the present invention provides a computer-readable storage medium, which may include a read-only memory (ROM), a random access memory (RAM), a disk, or an optical disk. A computer program is stored on the computer-readable storage medium, and the computer program is loaded by a processor to execute the steps of any text recognition method provided in an embodiment of the present invention. For example, the computer program loaded by the processor may execute the following steps:
[0213] Get the text line image to be recognized;
[0214] Acquire the writing direction of the to-be-recognized text line image, wherein the types of the writing direction include horizontal writing direction and vertical writing direction;
[0215] Extracting a feature map of the to-be-recognized text line image based on a trained text recognition model;
[0216] Performing a specific dimensional transformation process on the feature map by the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension;
[0217] The two-dimensional matrix is analyzed by the text recognition model to determine the text in the text line image to be recognized.
[0218] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above and will not be repeated here.
[0219] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to implement as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments and will not be repeated here.
[0220] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0221] The above is a detailed introduction to a text recognition method, device, computer equipment and storage medium provided in an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A text recognition method, characterized in that: include: Get the text line image to be recognized; Acquire the writing direction of the to-be-recognized text line image, wherein the types of the writing direction include horizontal writing direction and vertical writing direction; Extracting a feature map of the to-be-recognized text line image based on a trained text recognition model; Performing a specific dimensional transformation process on the feature map by the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension; Analyzing the two-dimensional matrix using the text recognition model to determine the text in the text line image to be recognized; The step of performing a specific dimension transformation on the feature map by using the text recognition model to convert the feature map into a two-dimensional matrix includes: If the writing direction is a horizontal writing direction, merging the original image longitudinal dimension and the original image channel dimension of the feature map through the text recognition model to obtain the first image channel dimension, and using the original image transverse dimension as the first image transverse dimension to obtain a first two-dimensional matrix including the first image transverse dimension and the first image channel dimension; If the writing direction is a vertical writing direction, the original image horizontal dimension and the original image channel dimension of the feature map are merged through the text recognition model to obtain the second image channel dimension, and the original image vertical dimension is used as the second image vertical dimension to obtain a second two-dimensional matrix including the second image vertical dimension and the second image channel dimension.
2. The text recognition method according to claim 1, characterized in that The analyzing the two-dimensional matrix by the text recognition model to determine the text in the to-be-recognized text line image includes: If the writing direction is a horizontal writing direction, the first two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector; Using the first column vector as a time series vector, learning the time series features of the first column vector through the text recognition model to obtain a second column vector; The second column vector is classified, and the text corresponding to the text line in the to-be-recognized text line image is determined based on the classification result.
3. The text recognition method according to claim 1, characterized in that The determining the text in the to-be-recognized text line image by analyzing the text recognition model based on the two-dimensional matrix includes: If the writing direction is a horizontal writing direction, the first two-dimensional matrix is expanded according to the horizontal dimension of the first image using the text recognition model to obtain a first column vector; if the writing direction is a vertical writing direction, the second two-dimensional matrix is expanded according to the vertical dimension of the second image using the text recognition model to obtain a first column vector; The first column vector is classified, and the text corresponding to the text line in the to-be-recognized text line image is determined based on the classification result.
4. The text recognition method according to claim 1, characterized in that The obtaining of the writing direction of the to-be-recognized text line image includes: Determining whether an actual length relationship between a horizontal length and a vertical length of the to-be-recognized text line image matches a preset length relationship corresponding to a horizontal writing direction or a vertical writing direction; Based on the judgment result, the writing direction of the to-be-recognized text line image is determined.
5. The text recognition method according to claim 4, characterized in that: The determining the writing direction of the to-be-recognized text line image based on the judgment result includes: If the actual length relationship matches the preset length relationship corresponding to the horizontal writing direction, determining that the writing direction of the to-be-recognized text line image is the horizontal writing direction; If the actual length relationship matches the preset length relationship corresponding to the vertical writing direction, determining that the writing direction of the to-be-recognized text line image is the vertical writing direction; If the actual length relationship does not match the preset length relationship corresponding to the horizontal writing direction and the vertical writing direction, setting the writing direction of the to-be-recognized text line image to include both the horizontal writing direction and the vertical writing direction; After analyzing the two-dimensional matrix by the text recognition model to determine the text in the text line image to be recognized, the method further includes: If the writing direction of the text line image to be recognized includes both the horizontal writing direction and the vertical writing direction, the correct text of the text line image to be recognized is determined from the text recognized by the text recognition model in the scenario where the horizontal writing direction is used as the writing direction of the text line image to be recognized, and the text recognized by the text recognition model in the scenario where the vertical writing direction is used as the writing direction of the text line image to be recognized.
6. The text recognition method according to any one of claims 1 to 5, characterized in that: Also includes: Acquire training samples, wherein the training samples include horizontal text samples and vertical text samples, and the sample labels of the training samples include the writing direction of the text in the samples and the character identifiers corresponding to the characters in the text; Get the text recognition model to be trained; Extracting a feature map from the training sample based on the text recognition model; Performing a specific dimensional transformation process on the feature map by the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the training sample, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; and if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension; Determining predicted character identifiers corresponding to text in the training sample by analyzing the text recognition model based on the two-dimensional matrix; Calculating a total loss of the text recognition model based on the predicted character identifiers of the training samples and the character identifiers in the sample labels; Parameters of the text recognition model are adjusted based on the total loss until the text recognition model training is completed.
7. A text recognition device, characterized in that: include: A first acquiring unit, configured to acquire an image of a text line to be recognized; A second acquiring unit is configured to acquire a writing direction of the to-be-recognized text line image, wherein the types of the writing direction include a horizontal writing direction and a vertical writing direction; A feature acquisition unit extracts a feature map of the text line image to be recognized using a trained text recognition model; a conversion unit, configured to perform a specific dimensional transformation process on the feature map using the text recognition model to convert the feature map into a two-dimensional matrix, wherein the specific dimensional transformation process corresponds to the writing direction of the text line image to be recognized, wherein if the writing direction is a horizontal writing direction, the dimensions of the two-dimensional matrix include a first image horizontal dimension and a first image channel dimension; and if the writing direction is a vertical writing direction, the dimensions of the two-dimensional matrix include a second image vertical dimension and a second image channel dimension; an analyzing unit, configured to analyze the two-dimensional matrix using the text recognition model to determine the text in the to-be-recognized text line image; Among them, the conversion unit is specifically used to, if the writing direction is a horizontal writing direction, merge the original image vertical dimension and the original image channel dimension of the feature map through the text recognition model to obtain the first image channel dimension, and use the original image horizontal dimension as the first image horizontal dimension to obtain a first two-dimensional matrix including the first image horizontal dimension and the first image channel dimension; if the writing direction is a vertical writing direction, merge the original image horizontal dimension and the original image channel dimension of the feature map through the text recognition model to obtain the second image channel dimension, and use the original image vertical dimension as the second image vertical dimension to obtain a second two-dimensional matrix including the second image vertical dimension and the second image channel dimension.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Character recognition method, device and equipment and readable storage medium
CN111126410A
Text recognition method and device, storage medium and electronic equipment
CN111400497A