A feature classification model acquisition system, method, device and electronic device

By establishing a feature classification model between the platform server and the user terminal to obtain a system, the problem of pixel-level annotation is solved in the prior art training feature classification model, and the effect of reducing text recognition costs and improving efficiency is achieved.

CN114387463BActive Publication Date: 2025-05-13ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011108960.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-16
Publication Date
2025-05-13
Estimated Expiration
2040-10-16

AI Technical Summary

Technical Problem

Existing text recognition methods based on intensive predictions require character pixel-level annotation of each character in the sample text image when training feature classification models, resulting in high text recognition costs.

Method used

The system is obtained by establishing a feature classification model between the platform server and the user terminal, and the sample text image sent by the user terminal and the corresponding character category annotation information are used to generate and train the feature classification model and provide it to the user terminal for use, avoiding pixel-level annotation of each character.

Benefits of technology

It reduces the cost of text recognition, improves the efficiency of text recognition, and reduces the need for character pixel-level annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387463B_ABST
    Figure CN114387463B_ABST
Patent Text Reader

Abstract

The present application provides a method for obtaining a target feature classification model, including: obtaining a sample text image, and obtaining character category annotation information corresponding to the characters in the sample text image, wherein the character category annotation information is annotation information for the character categories corresponding to the characters in the sample text image; obtaining a first feature classification model based on the character category annotation information, the sample text image feature map, and the character feature information; training the first feature classification model based on the character category annotation information and the sample text image feature map, and obtaining a second feature classification model for obtaining the character categories corresponding to the characters in the specified text image based on the specified text image feature map corresponding to the specified text image. The feature classification model obtaining method of the present application does not require character pixel-level annotation for the characters in the sample text image, thereby reducing the cost of text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly to a feature classification model acquisition system and acquisition method. The present application also relates to a target feature classification model acquisition device, electronic device and storage medium. The present application also relates to a text recognition method, device, electronic device and storage medium. Background Art

[0002] In daily life and work, we often need to process a large amount of text data. In order to improve the processing efficiency of text data, text recognition methods have become an important research topic in computer vision technology.

[0003] In the prior art, a text recognition method based on dense prediction is a mainstream text recognition method. The existing text recognition method based on dense prediction requires a feature classification model that is pre-trained for character classification before text recognition. However, the existing text recognition method based on dense prediction requires character pixel-level annotation of each character in the sample text image when training the feature classification model, which greatly increases the cost of text recognition. Summary of the invention

[0004] The present application provides a system and method for obtaining a feature classification model, an apparatus, an electronic device, and a storage medium to reduce the cost of text recognition.

[0005] The present application provides a feature classification model acquisition system, including: a platform server and a user terminal;

[0006] The platform server is used to obtain the sample text image sent by the user terminal, and obtain the character category annotation information corresponding to the characters in the sample text image, wherein the character category annotation information is the annotation information for the character category corresponding to the characters in the sample text image; for the sample text image, obtain the sample text image feature map corresponding to the sample text image, and obtain the character feature information corresponding to the characters in the sample text image in the sample text image feature map; obtain a first feature classification model for the character category corresponding to the characters in the text image based on the character category annotation information, the sample text image feature map and the character feature information; perform model training on the first feature classification model based on the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the characters in the specified text image based on the specified text image feature map corresponding to the specified text image; and provide the second feature classification model to the user terminal;

[0007] The user terminal is used to send the sample text image to the platform server; and obtain the second feature classification model sent by the platform server.

[0008] The present application provides a method for obtaining a feature classification model, comprising:

[0009] Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image;

[0010] For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map;

[0011] According to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map;

[0012] The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0013] Optionally, obtaining a sample text image and obtaining character category labeling information corresponding to characters in the sample text image includes:

[0014] Obtain an initial sample text image;

[0015] Performing image quality enhancement processing on the initial sample image to obtain the sample text image;

[0016] Character categories are annotated for the characters in the sample text image to obtain the character category annotated information.

[0017] Optionally, the method includes: performing image noise reduction processing on the sample text image to obtain the sample text image.

[0018] Optionally, it includes: inputting the sample text image into a pre-trained target feature extractor, performing image feature extraction on the sample text image, and obtaining a feature map of the sample text image.

[0019] Optionally, obtaining character feature information corresponding to the characters in the sample text image in the sample text image feature map includes: inputting the sample text image and the sample text image feature map into a target attention module to obtain the character feature information, the target attention module being used to obtain character feature information corresponding to the characters in the text image in the text image feature map based on the text image and the text image feature map corresponding to the text image.

[0020] Optionally, also include:

[0021] Obtain a target sample text image;

[0022] For the target sample text image, obtaining a target sample text image feature map corresponding to the target sample text image and character feature information corresponding to characters in the target sample text image in the target sample text image feature map;

[0023] The target attention module is obtained according to the target sample text image, the target sample text image feature map, and character feature information corresponding to characters in the target sample text image in the target sample text image feature map.

[0024] Optionally, obtaining a sample text image and obtaining character category annotation information corresponding to characters in the sample text image includes: obtaining a first sample text image and obtaining first character category annotation information corresponding to characters in the first sample text image;

[0025] The method of obtaining, for the sample text image, a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to the characters in the sample text image in the sample text image feature map, includes: obtaining, for the first sample text image, a first sample text image feature map corresponding to the first sample text image, and obtaining first character feature information corresponding to the characters in the first sample text image in the first sample text image feature map.

[0026] Optionally, the step of obtaining, according to the character category annotation information, the sample text image feature map, and the character feature information, a first feature classification model for obtaining a character category corresponding to a character in the text image based on the text image feature map corresponding to a text image and character feature information corresponding to characters in the text image in the text image feature map, comprises:

[0027] Inputting the first sample text image feature map and the first character feature information into a feature classification model to be trained to obtain a first character category corresponding to the character in the first sample text image;

[0028] According to the first character category and the first character category labeling information, obtaining a first target character category loss corresponding to the first character category and the first character category labeling information;

[0029] If the first target character category loss is less than the first specified character category loss, the feature classification model to be trained is used as the first feature classification model.

[0030] Optionally, also include:

[0031] If the first target character category loss is not less than the first designated character category loss, obtaining a second sample text image, and obtaining second character category labeling information corresponding to the characters in the second sample text image;

[0032] For the second sample text image, obtaining a second sample text image feature map corresponding to the second sample text image, and obtaining second character feature information corresponding to characters in the second sample text image in the second sample text image feature map;

[0033] Inputting the second sample text image feature map and the second character feature information into the feature classification model to be trained to obtain a second character category corresponding to the character in the second sample text image;

[0034] According to the second character category and the second character category labeling information, the first target character category loss corresponding to the second character category and the second character category labeling information is obtained, and so on, until the first feature classification model is obtained.

[0035] Optionally, the performing model training on the first feature classification model according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining a character category corresponding to a character in the specified text image according to the specified text image feature map corresponding to the specified text image, comprises:

[0036] Inputting the first character category annotation information and the first sample text image feature map corresponding to the first sample text image into the first feature classification model to obtain the first character category corresponding to the characters in the first sample text image;

[0037] According to the first character category and the first character category labeling information, obtaining a second target character category loss corresponding to the first character category and the first character category labeling information;

[0038] If the second target character category loss is less than the second specified character category loss, the feature classification model to be trained is used as the second feature classification model.

[0039] Optionally, the method further includes: if the second target character category loss is not less than the second specified character category loss, inputting the second character category labeling information and the second sample text image feature map corresponding to the second sample text image into the first feature classification model to obtain the second character category corresponding to the character in the second sample text image;

[0040] According to the second character category and the second character category labeling information, obtaining a second target character category loss corresponding to the second character category and the second character category labeling information;

[0041] According to the second character category and the second character category labeling information, a second target character category loss corresponding to the second character category and the second character category labeling information is obtained, and so on, until the second feature classification model is obtained.

[0042] On the other hand, the present application also provides a feature classification model obtaining device, including:

[0043] A character category annotation information obtaining unit, used to obtain a sample text image, and obtain character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image;

[0044] A character feature information obtaining unit, used for obtaining, for the sample text image, a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map;

[0045] A first feature classification model obtaining unit, for obtaining, based on the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to a text image and character feature information corresponding to a character in the text image in the text image feature map;

[0046] The second feature classification model acquisition unit is used to perform model training on the first feature classification model according to the character category annotation information and the sample text image feature map, and obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0047] In another aspect, the present application further provides a processor; and

[0048] The memory is used to store a program of a method for obtaining a feature classification model. After the device is powered on and the program of the method for obtaining a feature classification model is run by the processor, the following steps are performed:

[0049] Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image;

[0050] For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map;

[0051] According to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map;

[0052] The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0053] On the other hand, the present application further provides a storage medium storing a program of a method for obtaining a feature classification model, wherein the program is executed by a processor to perform the following steps:

[0054] Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image;

[0055] For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map;

[0056] According to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map;

[0057] The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0058] In another aspect, the present application further provides a text recognition method, comprising:

[0059] Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image;

[0060] Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map;

[0061] The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

[0062] Optionally, performing text recognition processing on the target text image using the target feature classification model to obtain a text recognition result of the target text image includes:

[0063] Inputting the target text image feature map into the target feature classification model to obtain a two-dimensional dense prediction result map corresponding to the target text image feature map, wherein the two-dimensional dense prediction result map is an image used to represent image features corresponding to different pixels in the target text image feature map;

[0064] The text recognition result is obtained according to the two-dimensional dense prediction result graph.

[0065] Optionally, obtaining the text recognition result according to the two-dimensional dense prediction result graph includes:

[0066] According to the pixel-level character category prediction results in the two-dimensional dense prediction result image, adjacent characters of the same category are neighborhood-merged to obtain the character category of the characters in the target text image, and the character coordinates corresponding to the characters in the target text image are obtained; according to the image features corresponding to the different pixels, the different pixels are neighborhood-merged to obtain the target character category of the characters in the target text image;

[0067] Clustering the characters in the target text image according to the ordinates of the character coordinates to obtain ordinate clusters corresponding to the characters in the target text image and obtaining a row sequence corresponding to the target character category in the text recognition result for the target character category;

[0068] According to the ordinate cluster, sort the characters in the same ordinate cluster from left to right according to the abscissa to obtain a single-line text recognition result corresponding to the text recognition result; sort the target character categories by line according to the line sequence to obtain a line text recognition result corresponding to the text recognition result;

[0069] The single-line text recognition results are connected according to the order of the different vertical coordinate clusters from top to bottom to obtain the text recognition result. The text recognition result is obtained according to the line text recognition results.

[0070] In another aspect, the present application further provides a text recognition device, comprising:

[0071] A target text image feature map obtaining unit, used to obtain a target text image to be recognized, and obtain a target text image feature map corresponding to the target text image;

[0072] a target feature classification model obtaining unit, used for obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image according to a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained according to character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image according to a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map;

[0073] The text recognition result obtaining unit is used to perform text recognition processing on the target text image using the target feature classification model to obtain the text recognition result of the target text image.

[0074] In another aspect, the present application further provides an electronic device, comprising:

[0075] Processor; and

[0076] The memory is used to store a program of the text recognition method. After the device is powered on and the program of the text recognition method is run by the processor, the following steps are performed:

[0077] Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image;

[0078] Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map;

[0079] The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

[0080] In another aspect, the present application further provides a storage medium storing a program of a text recognition method, wherein the program is executed by a processor to perform the following steps:

[0081] Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image;

[0082] Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map;

[0083] The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

[0084] Compared with the prior art, this application has the following advantages:

[0085] The feature classification model acquisition system and method provided in the present application first obtain a sample text image, and obtain character category annotation information corresponding to the characters in the sample text image, where the character category annotation information is annotation information for the character categories corresponding to the characters in the sample text image; secondly, for the sample text image, obtain a sample text image feature map corresponding to the sample text image, and obtain character feature information corresponding to the characters in the sample text image in the sample text image feature map; thirdly, based on the character category annotation information, the sample text image feature map and the character feature information, obtain a first feature classification model for obtaining the character categories corresponding to the characters in the text image based on the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map; finally, model training is performed on the first feature classification model based on the character category annotation information and the sample text image feature map, to obtain a second feature classification model for obtaining the character categories corresponding to the characters in the specified text image based on the specified text image feature map corresponding to the specified text image. The feature classification model acquisition method provided in the present application does not need to perform pixel-level annotation on characters in a sample text image when obtaining a second feature classification model for obtaining a character category corresponding to a specified text image based on a specified text image feature map corresponding to a specified text image, thereby reducing the cost of text recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 This is a first schematic diagram of an application scenario of the feature classification model acquisition method provided in the present application.

[0087] Figure 2 This is a second schematic diagram of an application scenario of the feature classification model acquisition method provided in the present application.

[0088] Figure 3 This is a flowchart of a method for obtaining a feature classification model provided in the first embodiment of the present application.

[0089] Figure 4 This is a schematic diagram of a feature classification model acquisition device provided in the second embodiment of the present application.

[0090] Figure 5 A schematic diagram of an electronic device provided in an embodiment of the present application.

[0091] Figure 6 This is a flowchart of a method for obtaining a feature classification model provided in the fifth embodiment of the present application.

[0092] Figure 7This is a schematic diagram of a text recognition device provided in the sixth embodiment of the present application.

[0093] Figure 8 This is a schematic diagram of a feature classification model acquisition system provided in the ninth embodiment of the present application. DETAILED DESCRIPTION

[0094] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0095] In order to more clearly demonstrate the method for obtaining a feature classification model provided by the embodiment of the present application, the application scenario of the method for obtaining a feature classification model provided by the embodiment of the present application is first introduced. In practical applications, text recognition methods based on dense prediction often need to obtain recognition results for characters in text images through a pre-trained feature classification model. In order to solve the problem that each character in the sample text image needs to be labeled at the pixel level when training the feature classification model, the embodiment of the present application provides a method for obtaining a feature classification model, thereby reducing the cost of text recognition methods based on dense prediction.

[0096] The method for obtaining a feature classification model provided in this application can be executed by a server or a client with a related text recognition application installed, such as a smartphone, tablet computer, or PC (Personal Computer) with a related text recognition application installed. The so-called server is generally a server or a server cluster in terms of specific implementation.

[0097] The following specifically takes the execution subject as the server as an example to explain in detail the application scenario of the method for obtaining the feature classification model provided by this application. Figure 1 As shown, it is a first schematic diagram of an application scenario of the feature classification model acquisition method provided in this application.

[0098] After obtaining the sample text image, the server 101 firstly performs character category annotation for the characters in the sample text image to obtain character category annotation information. Secondly, the sample text image is input into the pre-trained target feature extractor 101-1, and image feature extraction is performed on the sample text image to obtain the sample text image feature map corresponding to the sample text image. Thirdly, the sample text image and the sample text image feature map are input into the target attention module 101-2 to obtain character feature information. The target attention module 101-2 is used to obtain the character feature information corresponding to the characters in the text image in the text image feature map according to the text image and the text image feature map corresponding to the text image. Then, according to the character category annotation information, the sample text image feature map and the character feature information, the first feature classification model module 101-3 is used to obtain the first feature classification model corresponding to the character category of the characters in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map. Finally, the first feature classification model is trained according to the character category annotation information and the sample text image feature map through the second feature classification model module 101-4 to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0099] The method for obtaining the feature classification model provided in this application can also be applied to the scenario where the client and the server interact. Figure 2 As shown, it is a second schematic diagram of the application scenario of the feature classification model acquisition method provided in this application.

[0100] First, the server 201 obtains the sample text image provided by the client 202 and the character category annotation information corresponding to the characters in the sample text image. Secondly, the server 201 obtains the sample text image feature map corresponding to the sample text image, and obtains the character feature information corresponding to the characters in the sample text image in the sample text image feature map. Thirdly, the server 201 obtains a first feature classification model for obtaining the character category corresponding to the characters in the text image based on the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map based on the character category annotation information, the sample text image feature map, and the character feature information. Finally, the server 201 performs model training on the first feature classification model based on the character category annotation information and the sample text image feature map, and obtains a second feature classification model for obtaining the character category corresponding to the characters in the specified text image based on the specified text image feature map corresponding to the specified text image.

[0101] The embodiments of the present application do not specifically limit the application scenarios of the feature classification model acquisition method provided in the embodiments of the present application, such as: the feature classification model acquisition method provided in the present application can also be applied to the scenario of obtaining the feature classification model through the client, which will not be described one by one here. The embodiments corresponding to the application scenarios of the above-mentioned feature classification model acquisition method are provided to facilitate the understanding of the feature classification model acquisition method provided in the present application, and are not used to limit the feature classification model acquisition method provided in the present application.

[0102] First embodiment

[0103] In the first embodiment of the present application, a method for obtaining a feature classification model is provided as follows: Figure 3 Provide explanation.

[0104] Figure 3 This is a flowchart of a method for obtaining a feature classification model provided in the first embodiment of the present application. Figure 3 The method for obtaining the feature classification model shown includes: steps S301 to S304.

[0105] The execution subject of the feature classification model acquisition method provided in the first embodiment of the present application can be a server, or a client with a related text recognition application installed, such as a smartphone, tablet computer, and PC with a related text recognition application installed. The so-called server is generally a server or a server cluster in terms of specific implementation.

[0106] In step S301, a sample text image is obtained, and character category annotation information corresponding to characters in the sample text image is obtained.

[0107] In the first embodiment of the present application, the so-called text image is an image including one or more text characters, wherein the characters include letters, numbers, and characters with complex structures based on radicals, such as Chinese, Korean, and Japanese. The so-called sample text image is an image including text characters that is obtained in advance for performing a feature classification model.

[0108] In the first embodiment of the present application, the process of obtaining a sample text image is: first obtain an initial sample text image, and then perform image quality enhancement processing on the initial sample image to obtain a sample text image. The so-called initial sample text image is a text image acquired by an image acquisition device, or a text image stored in an image database for image storage. The specific implementation method of performing image quality enhancement processing on the initial sample image to obtain a sample text image is at least: performing image noise reduction processing on the sample text image to obtain a sample text image.

[0109] In the first embodiment of the present application, the specific implementation method for obtaining the character category annotation information corresponding to the characters in the sample text image is: character category annotation is performed on the characters in the sample text image to obtain the character category annotation information. The so-called character category annotation information is the annotation information of the character category corresponding to the characters in the sample text image, that is, the annotation information is obtained after the character category annotation is performed on the characters in the sample text image. The so-called character category is the category of the character used to identify the character. Specifically, if a text image includes two characters "4S shop", then the character categories of these three characters are "4", "S" and "shop" respectively.

[0110] In step S302, for the sample text image, a sample text image feature map corresponding to the sample text image is obtained, and character feature information corresponding to characters in the sample text image in the sample text image feature map is obtained.

[0111] The so-called feature map is an image generated after feature extraction of an image, and is composed of the image features of the image. For a sample text image, obtaining a sample text image feature map corresponding to the sample text image means inputting the sample text image into a pre-trained target feature extractor, performing image feature extraction on the sample text image, and obtaining the sample text image feature map. In a specific implementation, the sample text image can be input into a pre-trained target feature extractor based on CNN (Convolutional Neural Networks) to perform image feature extraction on the sample text image. CNN is essentially a multi-layer perceptron that uses local connections and shared weights to extract image features including color, texture, shape, and topological structure of the image.

[0112] Since a text image often includes more than just text characters, the features in the text image feature map often include more than just character features. In addition, a text image often includes more than just one text character. The so-called character feature information includes the feature information of the character and the position information of the feature of the character in the text image corresponding to the character in the feature map. The so-called character feature information includes the stroke feature information of the character and the outline feature information of the character.

[0113] In the first embodiment of the present application, the specific implementation method of obtaining the character feature information corresponding to the characters in the sample text image in the sample text image feature map is: inputting the sample text image and the sample text image feature map into the target attention module to obtain the character feature information. The so-called target attention module is used to obtain the character feature information corresponding to the characters in the text image in the text image feature map based on the text image and the text image feature map corresponding to the text image.

[0114] In addition, before obtaining character feature information through the target attention module, first, the target sample text image is obtained; then, for the target sample text image, the target sample text image feature map corresponding to the target sample text image and the character feature information corresponding to the characters in the target sample text image in the target sample text image feature map are obtained; finally, based on the target sample text image, the target sample text image feature map and the character feature information corresponding to the characters in the target sample text image in the target sample text image feature map, the target attention module is obtained.

[0115] In step S303, based on the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining the character category corresponding to the characters in the text image based on the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map.

[0116] In the first embodiment of the present application, the specific implementation method of obtaining the first feature classification model is as follows: first, the first sample text image feature map and the first character feature information are input into the feature classification model to be trained to obtain the first character category corresponding to the character in the first sample text image. Then, based on the first character category and the first character category annotation information, the first target character category loss corresponding to the first character category and the first character category annotation information is obtained. Finally, if the first target character category loss is less than the first specified character category loss, the feature classification model to be trained is used as the first feature classification model.

[0117] Before inputting the first sample text image feature map and the first character feature information into the feature classification model to be trained, it is necessary to first obtain the first sample text image and obtain the first character category annotation information corresponding to the characters in the first sample text image, and then, for the first sample text image, obtain the first sample text image feature map corresponding to the first sample text image and obtain the first character feature information corresponding to the characters in the first sample text image in the first sample text image feature map.

[0118] In addition, if the first target character category loss is not less than the first specified character category loss, the following steps need to be performed in sequence: First: obtain a second sample text image, and obtain the second character category annotation information corresponding to the characters in the second sample text image. Second: for the second sample text image, obtain a second sample text image feature map corresponding to the second sample text image, and obtain the second character feature information corresponding to the characters in the second sample text image in the second sample text image feature map. Third: input the second sample text image feature map and the second character feature information into the feature classification model to be trained, and obtain the second character category corresponding to the characters in the second sample text image. Fourth: based on the second character category and the second character category annotation information, obtain the first target character category loss corresponding to the second character category and the second character category annotation information, and so on, until the first feature classification model is obtained.

[0119] It should be noted that the formula for obtaining the target character category loss is: first target character category loss = -log(y t s t ), where s is the character category annotated in the sample text image, that is, s t ={s1, s2, s3, ...s l},s t is the character category annotated in the sample text image, specifically, the character category annotated in the sample text image and corresponding to the character category obtained in step t by inputting the sample text image feature map and character feature information into the feature classification model to be trained; t The character category is obtained by inputting the sample text image feature map and character feature information into the feature classification model to be trained, that is, y t ={y1, y2, y3, ...y l}, specifically, the sample text image feature map and character feature information are input into the feature classification model to be trained, and the character category is obtained in the tth step. This formula is used to supervise the serialized output of the feature classification model to be trained using cross entropy loss.

[0120] In step S304, the first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0121] In the first embodiment of the present application, in the first embodiment of the present application, the specific implementation method of obtaining the second feature classification model is: first, the first character category annotation information and the first sample text image feature map corresponding to the first sample text image are input into the first feature classification model to obtain the first character category corresponding to the character in the first sample text image. Then, based on the first character category and the first character category annotation information, the second target character category loss corresponding to the first character category and the first character category annotation information is obtained. Finally, if the second target character category loss is less than the second specified character category loss, the feature classification model to be trained is used as the second feature classification model.

[0122] In addition, if the second target character category loss is not less than the second designated character category loss, the following steps need to be performed in sequence: First: input the second character category annotation information and the second sample text image feature map corresponding to the second sample text image into the first feature classification model to obtain the second character category corresponding to the character in the second sample text image. Second: based on the second character category and the second character category annotation information, obtain the second target character category loss corresponding to the second character category and the second character category annotation information. Third: based on the second character category and the second character category annotation information, obtain the second target character category loss corresponding to the second character category and the second character category annotation information, and so on, until the second feature classification model is obtained.

[0123] It should be noted that the formula for obtaining the target character category loss is: Among them, z is the character category confidence corresponding to the characters in the sample text image obtained by inputting the character category annotation information and the sample text image feature map corresponding to the sample text image into the first feature classification model. The so-called character category confidence is the confidence of the character category corresponding to the characters in the sample text image. That is, z is a confidence tensor whose value range is between (0, 1), and H, W, and C are the width, height, and character category of the sample text image feature map, respectively. c is the hollowing parameter h of the C-dimensional vector, s is the character category annotated in the sample text image, s={s1,s1,...,s t}. This formula is used to suppress the overflow confidence of categories outside the character category annotation.

[0124] In the first embodiment of the present application, a method for obtaining a feature classification model is provided. First, a sample text image is obtained, and character category annotation information corresponding to the characters in the sample text image is obtained, where the character category annotation information is annotation information for the character categories corresponding to the characters in the sample text image; secondly, for the sample text image, a sample text image feature map corresponding to the sample text image is obtained, and character feature information corresponding to the characters in the sample text image in the sample text image feature map is obtained; thirdly, based on the character category annotation information, the sample text image feature map, and the character feature information, a first feature classification model is obtained for obtaining the character categories corresponding to the characters in the text image based on the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map; finally, model training is performed on the first feature classification model based on the character category annotation information and the sample text image feature map, to obtain a second feature classification model for obtaining the character categories corresponding to the characters in the specified text image based on the specified text image feature map corresponding to the specified text image. The feature classification model obtaining method provided in the first embodiment of the present application does not need to perform pixel-level annotation on characters in the sample text image when obtaining the second feature classification model for obtaining the character category corresponding to the characters in the specified text image according to the specified text image feature map corresponding to the specified text image, thereby reducing the cost of text recognition.

[0125] Second embodiment

[0126] Corresponding to the application scenario of the feature classification model acquisition method provided by the present application and the feature classification model acquisition method provided by the first embodiment, the second embodiment of the present application also provides a feature classification model acquisition device. Since the device embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple, and the relevant parts can be referred to the application scenario and the partial description of the first embodiment. The device embodiment described below is only exemplary.

[0127] Please refer to Figure 4 , which is a schematic diagram of a feature classification model obtaining device provided in the second embodiment of the present application.

[0128] The feature classification model obtaining device provided in the second embodiment of the present application includes:

[0129] The character category annotation information obtaining unit 401 is used to obtain a sample text image and obtain character category annotation information corresponding to the characters in the sample text image, wherein the character category annotation information is annotation information of the character category corresponding to the characters in the sample text image;

[0130] The character feature information obtaining unit 402 is used to obtain, for the sample text image, a sample text image feature map corresponding to the sample text image, and obtain character feature information corresponding to the characters in the sample text image in the sample text image feature map;

[0131] A first feature classification model obtaining unit 403 is used to obtain, according to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model for obtaining a character category corresponding to a character in the text image based on the text image feature map corresponding to a text image and character feature information corresponding to a character in the text image in the text image feature map;

[0132] The second feature classification model obtaining unit 404 is used to perform model training on the first feature classification model according to the character category annotation information and the sample text image feature map, and obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0133] Optionally, the character category annotation information obtaining unit 401 is specifically used to obtain an initial sample text image; perform image quality enhancement processing on the initial sample image to obtain the sample text image; perform character category annotation on characters in the sample text image to obtain the character category annotation information.

[0134] Optionally, performing image quality enhancement processing on the initial sample image to obtain the sample text image includes: performing image noise reduction processing on the sample text image to obtain the sample text image.

[0135] Optionally, the character feature information obtaining unit 402 is specifically configured to input the sample text image into a pre-trained target feature extractor, perform image feature extraction on the sample text image, and obtain the sample text image feature map.

[0136] Optionally, the character feature information obtaining unit 402 specifically inputs the sample text image and the sample text image feature map into a target attention module to obtain the character feature information. The target attention module is used to obtain the character feature information corresponding to the characters in the text image in the text image feature map based on the text image and the text image feature map corresponding to the text image.

[0137] Optionally, also include:

[0138] Obtain a target sample text image;

[0139] For the target sample text image, obtaining a target sample text image feature map corresponding to the target sample text image and character feature information corresponding to characters in the target sample text image in the target sample text image feature map;

[0140] The target attention module is obtained according to the target sample text image, the target sample text image feature map, and character feature information corresponding to characters in the target sample text image in the target sample text image feature map.

[0141] Optionally, obtaining a sample text image and obtaining character category annotation information corresponding to characters in the sample text image includes: obtaining a first sample text image and obtaining first character category annotation information corresponding to characters in the first sample text image;

[0142] The method of obtaining, for the sample text image, a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to the characters in the sample text image in the sample text image feature map, includes: obtaining, for the first sample text image, a first sample text image feature map corresponding to the first sample text image, and obtaining first character feature information corresponding to the characters in the first sample text image in the first sample text image feature map.

[0143] Optionally, the first feature classification model obtaining unit 403 is specifically used to input the first sample text image feature map and the first character feature information into the feature classification model to be trained to obtain a first character category corresponding to the character in the first sample text image; obtain a first target character category loss corresponding to the first character category and the first character category annotation information based on the first character category and the first character category annotation information; if the first target character category loss is less than the first specified character category loss, the feature classification model to be trained is used as the first feature classification model.

[0144] Optionally, also include:

[0145] If the first target character category loss is not less than the first designated character category loss, obtaining a second sample text image, and obtaining second character category labeling information corresponding to the characters in the second sample text image;

[0146] For the second sample text image, obtaining a second sample text image feature map corresponding to the second sample text image, and obtaining second character feature information corresponding to characters in the second sample text image in the second sample text image feature map;

[0147] Inputting the second sample text image feature map and the second character feature information into the feature classification model to be trained to obtain a second character category corresponding to the character in the second sample text image;

[0148] According to the second character category and the second character category labeling information, the first target character category loss corresponding to the second character category and the second character category labeling information is obtained, and so on, until the first feature classification model is obtained.

[0149] Optionally, the second feature classification model obtaining unit 404 is specifically used to input the first character category annotation information and the first sample text image feature map corresponding to the first sample text image into the first feature classification model to obtain the first character category corresponding to the character in the first sample text image; obtain the second target character category loss corresponding to the first character category and the first character category annotation information according to the first character category and the first character category annotation information; if the second target character category loss is less than the second specified character category loss, use the feature classification model to be trained as the second feature classification model.

[0150] Optionally, the method further includes: if the second target character category loss is not less than the second specified character category loss, inputting the second character category labeling information and the second sample text image feature map corresponding to the second sample text image into the first feature classification model to obtain the second character category corresponding to the character in the second sample text image;

[0151] According to the second character category and the second character category labeling information, obtaining a second target character category loss corresponding to the second character category and the second character category labeling information;

[0152] According to the second character category and the second character category labeling information, a second target character category loss corresponding to the second character category and the second character category labeling information is obtained, and so on, until the second feature classification model is obtained.

[0153] Third embodiment

[0154] Corresponding to the method for obtaining a feature classification model provided in the first embodiment of the present application, the third embodiment of the present application further provides an electronic device. Since the third embodiment is basically similar to the first embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the first embodiment. The device embodiment described below is merely illustrative.

[0155] Please refer to Figure 5 , which is a schematic diagram of an electronic device provided in an embodiment of the present application.

[0156] The electronic device includes: a processor 501;

[0157] and a memory 502 for storing a program of a method for obtaining a feature classification model. After the device is powered on and the program of the method for obtaining a feature classification model is run by the processor, the following steps are performed:

[0158] Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image;

[0159] For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map;

[0160] According to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map;

[0161] The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0162] It should be noted that the detailed description of the electronic device provided in the third embodiment of the present application can refer to the application scenarios of the feature classification model obtaining method provided in the present application and the relevant description of the feature classification model obtaining method provided in the first embodiment, which will not be repeated here.

[0163] Fourth embodiment

[0164] Corresponding to the method for obtaining a feature classification model provided in the first embodiment of the present application, the fourth embodiment of the present application further provides a storage medium. Since the fourth embodiment is substantially similar to the first embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the first embodiment. The device embodiment described below is merely illustrative.

[0165] The storage medium stores a computer program, which is executed by a processor to perform the following steps:

[0166] Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image;

[0167] For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map;

[0168] According to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map;

[0169] The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

[0170] It should be noted that the detailed description of the storage medium provided in the fourth embodiment of the present application can refer to the relevant description of the feature classification model acquisition method provided in the first embodiment of the present application, and will not be repeated here.

[0171] Fifth embodiment

[0172] Corresponding to the method for obtaining a feature classification model provided in the first embodiment of the present application, the fifth embodiment of the present application further provides a text recognition method, and the relevant parts can be referred to the partial description of the first embodiment. The following description of the fifth embodiment is merely illustrative.

[0173] Please refer to Figure 6 , which is a flowchart of a method for obtaining a feature classification model provided in the fifth embodiment of the present application. Figure 6 The method for obtaining the feature classification model shown includes: steps S601 to S603.

[0174] Step S601: obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image.

[0175] In the fifth embodiment of the present application, the so-called text image is an image including one or more text characters, wherein the characters include letters, numbers, and characters with complex structures based on radicals, such as Chinese, Korean, and Japanese. The so-called target text image is a pre-selected text image for text recognition.

[0176] The so-called feature map is an image generated after feature extraction of an image and is composed of image features of the image.

[0177] Step S602: Obtain a target feature classification model, where the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to a specified text image.

[0178] The so-called feature classification model is obtained based on the character feature information corresponding to the characters in the sample text image in the sample text image feature map, the character category annotation information corresponding to the characters in the sample text image, and the first feature classification model. The first feature classification model is a model for obtaining the character category corresponding to the characters in the text image based on the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map.

[0179] Step S603: Perform text recognition processing on the target text image using the target feature classification model to obtain a text recognition result of the target text image.

[0180] In the fifth embodiment of the present application, the process of using the target feature classification model to perform text recognition processing on the target text image and obtain the text recognition result of the target text image is as follows: 1. Input the target text image feature map into the target feature classification model to obtain a two-dimensional dense prediction result map corresponding to the target text image feature map, and the two-dimensional dense prediction result map is an image used to represent the image features corresponding to different pixels in the target text image feature map; 2. Obtain the text recognition result according to the two-dimensional dense prediction result map. Among them, the specific implementation method of obtaining the text recognition result according to the two-dimensional dense prediction result map is: obtaining the text recognition result according to the two-dimensional dense prediction result map, including: first, according to the pixel-level character category prediction result in the two-dimensional dense prediction result map, the neighboring characters of the same category are merged in the neighborhood to obtain the character category of the characters in the target text image, and the character coordinates corresponding to the characters in the target text image are obtained. Secondly, according to the ordinate of the character coordinates, the characters in the target text image are clustered to obtain the ordinate cluster corresponding to the characters in the target text image. That is, through the ordinate of the character coordinates, DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering is performed on the characters in the target text image according to the ordinate value. The specific implementation method is: set the cluster radius to 2, the minimum number of points to 1, perform iterative clustering, and obtain the ordinate cluster. Again, according to the ordinate cluster, sort the characters in the same ordinate cluster from left to right according to the abscissa to obtain the single-line text recognition result corresponding to the text recognition result. Finally, according to the order of different ordinate clusters from top to bottom, the single-line text recognition results are connected to obtain the text recognition result.

[0181] In the fifth embodiment of the present application, a text recognition method is provided. First, a target text image to be recognized is obtained, and a target text image feature map corresponding to the target text image is obtained; then, a target feature classification model is obtained, the feature classification model is a model for obtaining a character category corresponding to a character in a specified text image according to a specified text image feature map corresponding to the specified text image, the feature classification model is obtained according to character feature information corresponding to the characters in the sample text image in the sample text image feature map, character category annotation information corresponding to the characters in the sample text image, and a first feature classification model, the first feature classification model is a model for obtaining a character category corresponding to the characters in the text image according to the text image feature map corresponding to the text image and the character feature information corresponding to the characters in the text image in the text image feature map; finally, the target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image. In the text recognition method provided in the fifth embodiment of the present application, when obtaining a target feature classification model for obtaining a character category corresponding to a character in a specified text image according to a specified text image feature map corresponding to the specified text image, it is not necessary to perform character pixel-level annotation on the characters in the sample text image, thereby reducing the cost of text recognition.

[0182] Sixth embodiment

[0183] Corresponding to the application scenario of the text recognition method provided by the present application and the text recognition method provided by the first embodiment, the sixth embodiment of the present application also provides a text recognition device. Since the device embodiment is basically similar to the application scenario and the fifth embodiment, the description is relatively simple, and the relevant parts refer to the application scenario and the partial description of the first embodiment. The device embodiment described below is only illustrative.

[0184] Please refer to Figure 7 , which is a schematic diagram of a text recognition device provided in the sixth embodiment of the present application.

[0185] The text recognition device provided in the sixth embodiment of the present application includes:

[0186] The target text image feature map obtaining unit 701 is used to obtain a target text image to be recognized and obtain a target text image feature map corresponding to the target text image;

[0187] A target feature classification model obtaining unit 702 is used to obtain a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to a specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to a text image and character feature information corresponding to characters in the text image in the text image feature map;

[0188] The text recognition result obtaining unit 703 is used to perform text recognition processing on the target text image using the target feature classification model to obtain a text recognition result of the target text image.

[0189] Optionally, the text recognition result obtaining unit 703 is specifically used to input the target text image feature map into the target feature classification model to obtain a two-dimensional dense prediction result map corresponding to the target text image feature map, wherein the two-dimensional dense prediction result map is an image used to represent image features corresponding to different pixels in the target text image feature map; and obtain the text recognition result based on the two-dimensional dense prediction result map.

[0190] Optionally, obtaining the text recognition result according to the two-dimensional dense prediction result graph includes:

[0191] Performing neighborhood merging on adjacent characters of the same category according to the pixel-level character category prediction results in the two-dimensional dense prediction result image to obtain the character category of the characters in the target text image, and obtaining the character coordinates corresponding to the characters in the target text image;

[0192] Clustering the characters in the target text image according to the ordinates of the character coordinates to obtain ordinate clusters corresponding to the characters in the target text image;

[0193] According to the ordinate cluster, sorting characters in the same ordinate cluster from left to right according to the abscissa to obtain a single-line text recognition result corresponding to the text recognition result;

[0194] The single-line text recognition results are connected according to the order of the different vertical coordinate clusters from top to bottom to obtain the text recognition result.

[0195] Seventh embodiment

[0196] Corresponding to the text recognition method provided in the fifth embodiment of the present application, the seventh embodiment of the present application further provides an electronic device. Since the seventh embodiment is basically similar to the fifth embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the fifth embodiment. The device embodiment described below is only illustrative.

[0197] Please refer to Figure 5 , which is a schematic diagram of an electronic device provided in an embodiment of the present application.

[0198] The electronic device includes: a processor 501;

[0199] and a memory 502 for storing a program of the text recognition method. After the device is powered on and the program of the text recognition method is run by the processor, the following steps are performed:

[0200] Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image;

[0201] Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map;

[0202] The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

[0203] It should be noted that the detailed description of the electronic device provided in the seventh embodiment of the present application can refer to the relevant description of the text recognition method provided in the fifth embodiment of the present application, which will not be repeated here.

[0204] Eighth embodiment

[0205] Corresponding to the text recognition method provided in the first embodiment of the present application, the eighth embodiment of the present application further provides a storage medium. Since the eighth embodiment is basically similar to the fifth embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the fifth embodiment. The device embodiment described below is only illustrative.

[0206] The storage medium stores a computer program, which is executed by a processor to perform the following steps:

[0207] Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image;

[0208] Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map;

[0209] The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

[0210] It should be noted that the detailed description of the storage medium provided in the eighth embodiment of the present application can refer to the relevant description of the text recognition method provided in the fifth embodiment of the present application, and will not be repeated here.

[0211] Ninth embodiment

[0212] Corresponding to the application scenario embodiment of the model training method provided in the present application and the feature classification model acquisition method provided in the first embodiment, the ninth embodiment of the present application also provides a feature classification model acquisition system. Since the ninth embodiment is basically similar to the embodiment corresponding to the application scenario and the first embodiment, the description is relatively simple. For relevant parts, please refer to the embodiment corresponding to the application scenario and the partial description of the first embodiment. The ninth embodiment described below is merely illustrative.

[0213] Please refer to Figure 8 , which is a schematic diagram of a feature classification model acquisition system provided in the ninth embodiment of the present application.

[0214] The feature classification model acquisition system provided in the ninth embodiment of the present application includes: a platform server 801 and a user terminal 802.

[0215] In the ninth embodiment of the present application, the so-called platform server 801 refers to a computing device that provides services for the software platform or application platform installed on the user terminal 802 for executing the feature classification model acquisition method provided in the present application, and is generally a server or server cluster in specific implementation. The so-called user terminal 802 refers to a computing device installed with a software platform or application platform for executing the feature classification model acquisition method provided in the present application, and is generally a smart phone, tablet computer, personal computer, etc. in specific implementation. After the user terminal 802 uploads the sample text image to the platform server 801, the platform server 801 adopts the SAAS (Software-as-a-Service) working mode to obtain the feature classification model and provide it to the user terminal 802.

[0216] The platform server 801 is used to obtain the sample text image sent by the user terminal 802, and obtain the character category annotation information corresponding to the characters in the sample text image, wherein the character category annotation information is the annotation information for the character categories corresponding to the characters in the sample text image; for the sample text image, obtain the sample text image feature map corresponding to the sample text image, and obtain the character feature information corresponding to the characters in the sample text image in the sample text image feature map; obtain a first feature classification model for the character categories corresponding to the characters in the text image based on the character category annotation information, the sample text image feature map and the character feature information; perform model training on the first feature classification model based on the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character categories corresponding to the characters in the specified text image based on the specified text image feature map corresponding to the specified text image; and provide the second feature classification model to the user terminal 802.

[0217] The user terminal 802 is used to send the sample text image to the platform server 801 ; and obtain the second feature classification model sent by the platform server 801 .

[0218] Although the present application is disclosed as above in the preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

[0219] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0220] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash memory (Flash RAM). The memory is an example of a computer-readable medium.

[0221] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.

[0222] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A feature classification model acquisition system, characterized in that: include: Platform server and user terminal; The platform server is used to obtain the sample text image sent by the user terminal, and obtain character category annotation information corresponding to the characters in the sample text image, wherein the character category annotation information is annotation information of the character category corresponding to the characters in the sample text image; For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map; Obtaining a first feature classification model of a character category corresponding to a character in the text image according to the character category annotation information, the sample text image feature map, and the character feature information; Performing model training on the first feature classification model according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image; and providing the second feature classification model to the user terminal; The user terminal is used to send the sample text image to the platform server; and obtain the second feature classification model sent by the platform server.

2. A method for obtaining a feature classification model, characterized in that: include: Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image; For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map; Obtaining a first feature classification model of a character category corresponding to a character in the text image according to the character category annotation information, the sample text image feature map, and the character feature information; The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

3. The method for obtaining a feature classification model according to claim 2, characterized in that: The obtaining of the sample text image and obtaining character category labeling information corresponding to the characters in the sample text image includes: Obtain an initial sample text image; Performing image quality enhancement processing on the initial sample text image to obtain the sample text image; Character categories are annotated for the characters in the sample text image to obtain the character category annotated information.

4. The method for obtaining a feature classification model according to claim 3, characterized in that: The performing image quality enhancement processing on the initial sample text image to obtain the sample text image includes: performing image noise reduction processing on the sample text image to obtain the sample text image.

5. The method for obtaining a feature classification model according to claim 2, characterized in that: The step of obtaining a sample text image feature map corresponding to the sample text image includes: inputting the sample text image into a pre-trained target feature extractor, performing image feature extraction on the sample text image, and obtaining the sample text image feature map.

6. The method for obtaining a feature classification model according to claim 2 or 5, characterized in that: The method of obtaining character feature information corresponding to the characters in the sample text image in the sample text image feature map includes: inputting the sample text image and the sample text image feature map into a target attention module to obtain the character feature information, and the target attention module is used to obtain the character feature information corresponding to the characters in the text image in the text image feature map based on the text image and the text image feature map corresponding to the text image.

7. The method for obtaining a feature classification model according to claim 6, characterized in that: Also includes: Obtain a target sample text image; For the target sample text image, obtaining a target sample text image feature map corresponding to the target sample text image and character feature information corresponding to characters in the target sample text image in the target sample text image feature map; The target attention module is obtained according to the target sample text image, the target sample text image feature map, and character feature information corresponding to characters in the target sample text image in the target sample text image feature map.

8. The method for obtaining a feature classification model according to claim 2 or 5, characterized in that: The obtaining of the sample text image and obtaining character category annotation information corresponding to the characters in the sample text image includes: obtaining a first sample text image and obtaining first character category annotation information corresponding to the characters in the first sample text image; The method of obtaining, for the sample text image, a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to the characters in the sample text image in the sample text image feature map, includes: obtaining, for the first sample text image, a first sample text image feature map corresponding to the first sample text image, and obtaining first character feature information corresponding to the characters in the first sample text image in the first sample text image feature map.

9. The method for obtaining a feature classification model according to claim 8, characterized in that: The method of obtaining, according to the character category labeling information, the sample text image feature map, and the character feature information, a first feature classification model for obtaining a character category corresponding to a character in the text image based on the text image feature map corresponding to a text image and character feature information corresponding to a character in the text image in the text image feature map, comprises: Inputting the first sample text image feature map and the first character feature information into a feature classification model to be trained to obtain a first character category corresponding to the character in the first sample text image; According to the first character category and the first character category labeling information, obtaining a first target character category loss corresponding to the first character category and the first character category labeling information; If the first target character category loss is less than the first specified character category loss, the feature classification model to be trained is used as the first feature classification model.

10. The method for obtaining a feature classification model according to claim 9, characterized in that: Also includes: If the first target character category loss is not less than the first designated character category loss, obtaining a second sample text image, and obtaining second character category labeling information corresponding to the characters in the second sample text image; For the second sample text image, obtaining a second sample text image feature map corresponding to the second sample text image, and obtaining second character feature information corresponding to characters in the second sample text image in the second sample text image feature map; Inputting the second sample text image feature map and the second character feature information into the feature classification model to be trained to obtain a second character category corresponding to the character in the second sample text image; According to the second character category and the second character category labeling information, the first target character category loss corresponding to the second character category and the second character category labeling information is obtained, and so on, until the first feature classification model is obtained.

11. The method for obtaining a feature classification model according to claim 10, characterized in that: The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining a character category corresponding to a character in the specified text image according to the specified text image feature map corresponding to the specified text image, including: Inputting the first character category labeling information and the first sample text image feature map corresponding to the first sample text image into the first feature classification model to obtain the first character category corresponding to the characters in the first sample text image; According to the first character category and the first character category labeling information, obtaining a second target character category loss corresponding to the first character category and the first character category labeling information; If the second target character category loss is less than the second specified character category loss, the first feature classification model is used as the second feature classification model.

12. The method for obtaining a feature classification model according to claim 11, characterized in that: Also includes: If the second target character category loss is not less than the second specified character category loss, inputting the second character category labeling information and the second sample text image feature map corresponding to the second sample text image into the first feature classification model to obtain the second character category corresponding to the character in the second sample text image; According to the second character category and the second character category labeling information, obtaining a second target character category loss corresponding to the second character category and the second character category labeling information; According to the second character category and the second character category labeling information, a second target character category loss corresponding to the second character category and the second character category labeling information is obtained, and so on, until the second feature classification model is obtained.

13. A device for obtaining a feature classification model, characterized in that: include: A character category annotation information obtaining unit, used to obtain a sample text image, and obtain character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image; A character feature information obtaining unit, used for obtaining, for the sample text image, a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map; A first feature classification model obtaining unit, for obtaining, based on the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to a text image and character feature information corresponding to a character in the text image in the text image feature map; The second feature classification model acquisition unit is used to perform model training on the first feature classification model according to the character category annotation information and the sample text image feature map, and obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

14. An electronic device, characterized in that: include: processor; as well as The memory is used to store a program of a method for obtaining a feature classification model. After the device is powered on and the program of the method for obtaining a feature classification model is run by the processor, the following steps are performed: Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image; For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map; According to the character category labeling information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and character feature information corresponding to the characters in the text image in the text image feature map; The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

15. A storage medium, characterized in that: A program for obtaining a feature classification model is stored, and the program is executed by a processor to perform the following steps: Obtaining a sample text image, and obtaining character category annotation information corresponding to characters in the sample text image, wherein the character category annotation information is annotation information of character categories corresponding to characters in the sample text image; For the sample text image, obtaining a sample text image feature map corresponding to the sample text image, and obtaining character feature information corresponding to characters in the sample text image in the sample text image feature map; According to the character category annotation information, the sample text image feature map and the character feature information, a first feature classification model is obtained for obtaining a character category corresponding to a character in the text image according to the text image feature map corresponding to the text image and character feature information corresponding to the characters in the text image in the text image feature map; The first feature classification model is trained according to the character category annotation information and the sample text image feature map to obtain a second feature classification model for obtaining the character category corresponding to the character in the specified text image according to the specified text image feature map corresponding to the specified text image.

16. A text recognition method, characterized in that: include: Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image; Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map; The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

17. The text recognition method according to claim 16, characterized in that: The step of performing text recognition processing on the target text image by using the target feature classification model to obtain a text recognition result of the target text image includes: Inputting the target text image feature map into the target feature classification model to obtain a two-dimensional dense prediction result map corresponding to the target text image feature map, wherein the two-dimensional dense prediction result map is an image used to represent image features corresponding to different pixels in the target text image feature map; The text recognition result is obtained according to the two-dimensional dense prediction result graph.

18. The text recognition method according to claim 17, characterized in that: The step of obtaining the text recognition result according to the two-dimensional dense prediction result graph includes: Performing neighborhood merging on adjacent characters of the same category according to the pixel-level character category prediction results in the two-dimensional dense prediction result image to obtain the character category of the characters in the target text image, and obtaining the character coordinates corresponding to the characters in the target text image; Clustering the characters in the target text image according to the ordinates of the character coordinates to obtain ordinate clusters corresponding to the characters in the target text image; According to the ordinate cluster, sorting characters in the same ordinate cluster from left to right according to the abscissa to obtain a single-line text recognition result corresponding to the text recognition result; The single-line text recognition results are connected according to the order of different vertical coordinate clusters from top to bottom to obtain the text recognition result.

19. A text recognition device, characterized in that: include: A target text image feature map obtaining unit, used to obtain a target text image to be recognized, and obtain a target text image feature map corresponding to the target text image; a target feature classification model obtaining unit, used for obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image according to a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained according to character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image according to a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map; The text recognition result obtaining unit is used to perform text recognition processing on the target text image using the target feature classification model to obtain the text recognition result of the target text image.

20. An electronic device, characterized in that: include: processor; as well as The memory is used to store a program of the text recognition method. After the device is powered on and the program of the text recognition method is run by the processor, the following steps are performed: Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image; Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map; The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

21. A storage medium, characterized in that: A program for storing a text recognition method is executed by a processor to perform the following steps: Obtaining a target text image to be recognized, and obtaining a target text image feature map corresponding to the target text image; Obtaining a target feature classification model, wherein the feature classification model is a model for obtaining character categories corresponding to characters in a specified text image based on a specified text image feature map corresponding to the specified text image, wherein the feature classification model is obtained based on character feature information corresponding to characters in a sample text image in the sample text image feature map, character category annotation information corresponding to characters in the sample text image, and a first feature classification model, wherein the first feature classification model is a model for obtaining character categories corresponding to characters in the text image based on a text image feature map corresponding to the text image and character feature information corresponding to characters in the text image in the text image feature map; The target feature classification model is used to perform text recognition processing on the target text image to obtain a text recognition result of the target text image.

Citation Information

Patent Citations

  • A method and apparatus for generating an image recognition model

    CN109214386A

  • A character recognition model training method and device

    CN109697442A