Method, apparatus, and server for recognizing text characters
The proposed text character recognition method uses a model with low and high layer convolutional networks and fire modules to address overlapping text issues, achieving precise and efficient character identification.
Patent Information
- Application Number
- CN202110300713.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-03-22
AI Technical Summary
The prior art is difficult to accurately identify overlapping text characters in complex scenarios, resulting in low recognition accuracy.
A preset character recognition model including low-level convolutional network layer, high-level convolutional network layer and fire module is adopted, and a cross-layer connection is set up between the low-level convolutional network layer and the high-level convolutional network layer. Through multi-scale feature extraction and feature fusion, the recognition accuracy is improved.
In complex scenarios such as overlapping text characters, efficient and accurate text character recognition is achieved, reducing recognition errors and improving the accuracy of character recognition.
Smart Images

Figure CN112883956B_ABST
Abstract
Description
Technical Field
[0001] This specification belongs to the technical field of artificial intelligence, and particularly relates to a method, apparatus, and server for identifying text characters. Background Art
[0002] In many data processing scenarios, the directly obtained system data is often image data containing text characters. At this time, the system needs to first perform text character recognition (e.g., OCR recognition) on the above image data to extract the text characters contained in the image data; and then perform specific data processing based on the extracted text characters.
[0003] However, for some relatively complex recognition scenarios, for example, the text characters in the image overlap due to certain reasons, making it difficult to identify the characters. Based on the existing methods, it is often difficult to accurately identify the real text characters in the image.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] This specification provides a method, apparatus, and server for identifying text characters, which can be applicable to complex recognition scenarios such as overlapping text characters, and accurately and efficiently identify and determine the text characters in the image.
[0006] This specification provides a method for identifying text characters, including:
[0007] Obtain a target image to be processed; wherein, the target image contains target text characters to be recognized;
[0008] Call a preset character recognition model to process the target image to obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer;
[0009] Determine the target text characters according to the processing result.
[0010] In one embodiment, the low-level convolutional network layer includes: a first convolutional layer, a second convolutional layer, and a third convolutional layer; the high-level convolutional network layer includes: a fourth convolutional layer and a fifth convolutional layer; the fire module includes a first fire module and a second fire module.
[0011] In one embodiment, the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are sequentially connected in series; and a first fire module and a second fire module are sequentially connected in series between the fourth convolutional layer and the fifth convolutional layer.
[0012] In one embodiment, a cross-layer connection is provided between the third convolutional layer and the fifth convolutional layer; and / or, a cross-layer connection is provided between the first convolutional layer and the fourth convolutional layer.
[0013] In one embodiment, before obtaining the target image to be processed, the method further includes:
[0014] Construct an initial model; wherein, the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial fire module, and a cross-layer connection is also provided between the initial low-level convolutional network layer and the initial high-level convolutional network layer;
[0015] Obtain a sample image; wherein, the sample image includes overlapping text characters;
[0016] According to the sample image, establish a training set and a test set; and label the sample images in the training set to obtain the labeled training set;
[0017] Use the labeled training set and the test set to train the initial model to obtain a preset character recognition model that meets the requirements.
[0018] In one embodiment, obtaining a sample image includes:
[0019] Collect first picture data including text characters;
[0020] Perform augmentation processing on the first picture data to obtain second picture data;
[0021] According to the text characters, segment the second picture data to obtain a plurality of third picture data;
[0022] Select a sample image including overlapping text characters from the plurality of third picture data.
[0023] In one embodiment, the method further includes:
[0024] Determine the receptive field range of the feature;
[0025] According to the receptive field range of the feature, adjust the size parameters of the convolution kernels used in the initial low-level convolutional network layer and the initial high-level convolutional network layer.
[0026] In one embodiment, after obtaining the sample image, the method further includes:
[0027] According to the sample image, calculate the average value and variance of the sample image;
[0028] Perform batch normalization processing on the sample image according to the mean value and variance of the sample image.
[0029] In one embodiment, the target image includes at least one of the following: a picture containing a bill, a picture containing a certificate, and a picture containing a contract.
[0030] This specification also provides a method for recognizing text characters, including:
[0031] Obtain a target image to be processed; wherein, the target image contains target text characters to be recognized;
[0032] Call a preset character recognition model to process the target image to obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module;
[0033] Determine the target text characters according to the processing result.
[0034] This specification also provides a method for establishing a preset character recognition model, including:
[0035] Construct an initial model; wherein, the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial fire module, and a cross-layer connection is also set between the initial low-level convolutional network layer and the initial high-level convolutional network layer;
[0036] Obtain sample images; wherein, the sample images contain overlapping text characters;
[0037] Establish a training set and a test set according to the sample images; and label the sample images in the training set to obtain a labeled training set;
[0038] Use the labeled training set and the test set to train the initial model to obtain a preset character recognition model that meets the requirements.
[0039] This specification also provides a text character recognition device, including:
[0040] An acquisition module, configured to acquire a target image to be processed; wherein, the target image contains target text characters to be recognized;
[0041] A call module, configured to call a preset character recognition model to process the target image to obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also set between the low-level convolutional network layer and the high-level convolutional network layer;
[0042] A determination module, configured to determine the target text character according to the processing result.
[0043] This specification also provides a server, including a processor and a memory for storing processor-executable instructions. When the processor executes the instructions, relevant steps of the text character recognition method are implemented.
[0044] This specification also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed, relevant steps of the text character recognition method are implemented.
[0045] This specification provides a method, device, and server for text character recognition. Based on this method, before specific implementation, a preset character recognition model including at least a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and with a cross-layer connection set between the low-level convolutional network layer and the high-level convolutional network layer can be pre-trained and established; during specific implementation, after obtaining a target image to be processed, the above preset character recognition model can be called to process the target image to obtain a corresponding processing result; then, according to the above processing result, the target text character included in the target image is recognized and determined. Thus, by calling the above preset character recognition model that supports multi-scale feature extraction and has good effects, it can effectively be applicable to complex recognition scenarios where characters are difficult to identify, such as overlapping text characters, and accurately and efficiently recognize and determine the text characters in the image, reducing recognition errors and improving the accuracy of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the embodiments of this specification, the drawings required for use in the embodiments will be briefly introduced below. The drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 FIG. is a schematic diagram of an embodiment of the structural composition of a system applying the text character recognition method provided by the embodiments of this specification;
[0048] Figure 2 FIG. is a schematic flowchart of the text character recognition method provided by an embodiment of this specification;
[0049] Figure 3 FIG. is a schematic diagram of an embodiment of applying the text character recognition method provided by the embodiments of this specification in a scenario example;
[0050] Figure 4It is a schematic diagram of an embodiment of applying the text character recognition method provided in the embodiments of this specification in a scenario example;
[0051] Figure 5 It is a schematic flowchart of the text character recognition method provided in an embodiment of this specification;
[0052] Figure 6 It is a schematic flowchart of the method for establishing a preset character recognition model provided in an embodiment of this specification;
[0053] Figure 7 It is a schematic diagram of the structural composition of a server provided in an embodiment of this specification;
[0054] Figure 8 It is a schematic diagram of the structural composition of a text character recognition device provided in an embodiment of this specification;
[0055] Figure 9 It is a schematic diagram of an embodiment of applying the text character recognition method provided in the embodiments of this specification in a scenario example. Detailed implementation manners
[0056] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this specification.
[0057] The embodiments of this specification provide a text character recognition method. The text character recognition method can be specifically applied to a system including a server and a terminal device. Specifically, reference can be made to Figure 1 As shown, the terminal device and the server can be connected by wire or wirelessly to perform specific data interaction.
[0058] In this embodiment, the server can specifically include a background server disposed on one side of the service data processing platform, which can implement functions such as data transmission and data processing. Specifically, the server can be, for example, an electronic device with data operation, storage functions, and network interaction functions. Or, the server can also be a software program running in the electronic device to provide support for data processing, storage, and network interaction. In this embodiment, the number of the servers is not specifically limited. The server can specifically be one server, or several servers, or a server cluster formed by several servers.
[0059] In this embodiment, the terminal device may specifically include a front-end device disposed on the user side and capable of implementing functions such as image data acquisition and image data transmission. Specifically, the terminal device may be, for example, a surveillance camera, or a desktop computer, a tablet computer, a laptop computer, a smart phone, etc. provided with a camera. Alternatively, the terminal device may also be a software application that can run on the above-mentioned electronic devices and support invoking the camera of the electronic device to collect image data. For example, it may be a certain APP running on a smart phone.
[0060] In specific implementation, the terminal device may collect a photo containing target text characters to be recognized and extracted as a target image to be processed. For example, a business handler of a certain bank may use a smart phone provided with a camera as the terminal device to photograph the bill provided by the user as the target image.
[0061] Then, the terminal device may send the collected target image to the server in a wired or wireless manner. Correspondingly, the server obtains the above-mentioned target image.
[0062] Furthermore, the server may call a preset character recognition model that has been pre-trained using sample images containing overlapping text characters to process the target image to output a corresponding processing result.
[0063] Among them, the model structure of the above-mentioned preset character recognition model is a model structure obtained by specifically improving for complex recognition scenarios such as overlapping text characters.
[0064] Specifically, the above-mentioned preset character recognition model may at least include a low-level convolutional network layer for extracting low-level features, a high-level convolutional network layer for extracting high-level features, and a fire module that deepens the convolutional network layer both in depth and width to obtain more scales and richer features. And a cross-layer connection is further provided before the low-level convolutional network layer and the high-level convolutional network layer, so that different types of image features, namely low-level features and high-level features, can be fused simultaneously later to more accurately identify the text characters in the image.
[0065] In specific implementation, the server may process the target image by invoking the convolutional network layer in the preset character recognition model to extract features with better effects, higher accuracy, and more scales (which can also be called feature vectors, feature maps, feature matrices, etc.); and then the server may call the classifier in the preset character recognition model based on the above features to obtain a corresponding processing result.
[0066] Furthermore, the server may, based on the above processing result, identify and determine the target text characters contained in the target image; and perform specific target data processing according to the identified target text characters.
[0067] For example, the server of a data processing system of a certain bank can identify the target text characters on a bill from a target image, and then further extract key information such as the payer information, payee information, and drawn amount related to the bill based on the identified target text characters, and implement electronic filing and storage of data for the bill based on the above key information.
[0068] Through the above system, it can be effectively applied to complex recognition scenarios such as overlapping text characters, and accurately and efficiently identify and determine the text characters in the image.
[0069] Refer to Figure 2 As shown, the embodiments of this specification provide a method for recognizing text characters. Among them, this method is specifically applied to the server side. Specifically, this method may include the following content.
[0070] S201: Obtain a target image to be processed; wherein, the target image contains target text characters to be recognized.
[0071] In some embodiments, the above target image may specifically be a picture to be processed that contains target text characters to be recognized. For example, it may be a photo that contains target text characters to be recognized, or a video screenshot that contains target text characters to be recognized, etc.
[0072] In some embodiments, specifically, in the scenario of banking business handling, the above target image may specifically be a picture that contains a bill. Correspondingly, the text characters on the bill (such as payee information, payer information, drawn amount, etc.) may be target text characters to be recognized. In the scenario of document verification, the above target image may be a picture that contains a document. Correspondingly, the text characters on the document (such as name, document number, native place, etc.) may be target text characters to be recognized. In the scenario of contract data filing, the above target image may also be a picture that contains a contract. Correspondingly, the text characters on the contract (such as contract terms, contract signatures, etc.) may be target text characters to be recognized. Of course, it should be noted that the above-listed target images and target text characters to be recognized are only illustrative. Specifically, for specific application scenarios, other types of target images and other types of target text characters may also be introduced.
[0073] Through the above embodiments, the method for recognizing text characters provided in this specification can be extended and applied to multiple different application scenarios to process multiple different target images.
[0074] In some embodiments, taking the scenario of handling banking services as an example, sometimes it is necessary to collect images containing bills (or forms) provided by users, identify and extract the text characters on the bills in the images, and then handle corresponding services for the users according to the identified and extracted text characters, or perform electronic archiving and other related data processing on the bills.
[0075] However, for the above-mentioned images containing bills, sometimes due to the relatively thin paper of the bills used, or the relatively heavy ink of the pen used by the user when filling in, there may be a phenomenon that a certain character on the bill overlaps with another character (there are overlapping text characters), making the characters on the bill difficult to identify. Or, due to factors such as the shooting angle of the image, ambient light, and the camera, the characters on the bill in the captured image become relatively blurred and difficult to identify.
[0076] In view of the above complex recognition scenarios where text characters in the image are difficult to identify, such as overlapping text characters, existing text character recognition methods are prone to errors and have low recognition accuracy. For example, based on existing methods, when performing text character recognition, it may be interfered by overlapping characters, resulting in the inability to recognize the real text characters.
[0077] In some embodiments, in order to be able to apply to the above complex recognition scenarios simultaneously, accurately and efficiently identify and determine the text characters in the image, and improve the recognition accuracy, a preset character recognition model that meets the requirements can be pre-trained and established. How to train and establish this preset character recognition model will be described separately later.
[0078] S202: Invoke the preset character recognition model to process the target image to obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer.
[0079] In some embodiments, it can specifically participate Figure 3 As shown, the preset character recognition model used at least includes the following structures: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module; and a corresponding cross-layer connection can be further provided between the low-level convolutional network layer and the high-level convolutional network layer.
[0080] Among them, the above-mentioned low-level convolutional network layer can specifically be used to extract low-level features, such as low-level features like color and size. The above-mentioned high-level convolutional network layer can specifically be used to extract corresponding high-level features, such as high-level features like texture and grain, based on the above low-level features through non-linear transformation.
[0081] The above-mentioned FIRE module may specifically refer to the FIRE module proposed in the compressed neural network. This module includes two parts: a compression layer and an expansion layer. Specifically, it can be used to deepen the convolutional neural network (e.g., a shallow convolutional network) in both depth and width aspects, enrich the feature expression ability of the convolutional neural network, so as to extract multi-scale and more abundant features, so that subsequent text character recognition and determination can be carried out more precisely and accurately based on the above features.
[0082] The above-mentioned cross-layer connection can specifically be used to input the lower-layer features output by the relatively front low-level convolutional network layer into the relatively back high-level convolutional network layer separated by the middle convolutional network layer, so that in the relatively back high-level convolutional network layer, two types of features, namely the higher-layer features and the lower-layer features input by the previous convolutional network layer, can be fused simultaneously, and more accurate and comprehensive features can be extracted, effectively reducing the influence of the loss of some feature information of the lower-layer features during the processing of the middle convolutional layer on the subsequent text character recognition and extraction.
[0083] In some embodiments, the low-level convolutional network layer may specifically include: a first convolutional layer, a second convolutional layer, and a third convolutional layer; the high-level convolutional network layer may specifically include: a fourth convolutional layer and a fifth convolutional layer; the FIRE module may specifically include a first FIRE module and a second FIRE module.
[0084] Specifically, reference can be made to Figure 3 As shown, the above-mentioned first convolutional layer may specifically be a 5×5 convolutional layer, denoted as: Conv1; the above-mentioned second convolutional layer may specifically be a 3×3 convolutional layer, denoted as: Conv2; the above-mentioned third convolutional layer may specifically be a 1×1 convolutional layer, denoted as: Conv3. The above-mentioned fourth convolutional layer may specifically be a 3×3 convolutional layer, denoted as: Conv4; the above-mentioned fifth convolutional layer may specifically be denoted as: Conv5.
[0085] The above-mentioned first FIRE module (which can be denoted as fire1) and second FIRE module (which can be denoted as fire2) may specifically be two FIRE modules with the same structure. Among them, each FIRE module may include a compression layer and an expansion layer cascaded together.
[0086] Specifically, reference can be made to Figure 4As shown in the figure, the above compression layer may specifically include a 1×1 convolutional layer (denoted as the compression convolutional layer, for example, Kernerl = 1*1, Num = s1, that is, the number of layers is 1×1, and the number of convolutional kernels is s1), which is used to perform a compression operation on the input features. The above expansion layer may specifically include two parallel convolutional layers, namely: a 1×1 convolutional layer (denoted as the first expansion convolutional layer, for example, Kernerl = 1*1, Num = e1, that is, the number of layers is 1×1, and the number of convolutional kernels is e1) and a 3×3 convolutional layer (denoted as the second expansion convolutional layer, for example, Kernerl = 3*3, Num = e3, that is, the number of layers is 3×3, and the number of convolutional kernels is e3). Further, after the above two parallel convolutional layers, there is also a fusion structure for cascading and fusing features, denoted as: contact.
[0087] Among them, the number of convolutional kernels of the above compression convolutional layer can be set to s1, the number of convolutional kernels of the above first expansion convolutional layer can be set to e1, the number of convolutional kernels of the above second expansion convolutional layer can be set to e3, and the following relationship is satisfied: e1 = e3 = 4s1, and s1 is less than the number of image channels.
[0088] Based on the fire module with the above structure, in specific implementation, the input features (for example, features with a size of H*W*M) can be first compressed by the compression convolutional layer of the compression layer, and then the compressed features (for example, features with a size of H*W*s1) are output. Then, the above compressed features are input into the first expansion layer of the expansion layer to output the first intermediate features (for example, features with a size of H*W*e1), and at the same time, the above compressed features are input into the second expansion layer of the expansion layer to output the second intermediate features (for example, features with a size of H*W*e3). Then, the above first intermediate features and second intermediate features are input into the fusion structure for splicing and fusion to obtain the fused expansion features (for example, features with a size of H*W*(e1+e3), where (e1+e3) represents the dimension of the features).
[0089] By introducing and using the fire module with the above structure, the processing efficiency of feature extraction can be improved; at the same time, it can also deepen the convolutional neural network in terms of depth and width, increase the robustness of the network, so that relatively richer and more multi-scale features can be extracted based on this network.
[0090] In some embodiments, further considering that the size of the target image containing the target text characters to be recognized to be processed is usually relatively small, in order to avoid the inability to train a preset character recognition model that meets the requirements due to the disappearance of gradients during the training process. Therefore, only the first fire module and the second fire module are selected for series combination as a whole fire module and applied to the above preset character recognition model.
[0091] In some embodiments, in the above-mentioned preset character recognition model, the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer may be specifically connected in series in sequence; and between the fourth convolutional layer and the fifth convolutional layer, a first Fire module and a second Fire module are also connected in series in sequence.
[0092] Specifically, reference may be made to Figure 3 As shown, in the preset recognition model, the first convolutional layer is connected to the second convolutional layer, the second convolutional layer is connected to the third convolutional layer, the third convolutional layer is connected to the fourth convolutional layer, the fourth convolutional layer is connected to the first Fire module in the Fire module, and the first Fire module is connected to the second Fire module.
[0093] In addition, the above-mentioned preset character recognition model may specifically further include an input layer (denoted as: Input). Among them, the above-mentioned input layer is connected to the first convolutional layer, and the input layer is used to access the target image to be processed (which may be denoted as X).
[0094] The above-mentioned preset character recognition model may specifically further include one or a plurality of fully connected layers connected in series in sequence (for example, fully connected layer 1, fully connected layer 2, and fully connected layer 3 connected in series in sequence), and a Softmax classifier. Through the above structure, the features output by the fifth convolutional layer can be accessed, and based on the above features, the processing result for the target image (which may be denoted as Y) can be processed and output.
[0095] Through the above embodiments, by introducing and connecting the Fire module, the convolutional neural network can be deepened in depth and breadth, and the robustness of the network can be increased, so that relatively richer and more multi-scale features can be extracted subsequently.
[0096] In some embodiments, in order to further improve the recognition accuracy of the preset character recognition model, a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer.
[0097] In some embodiments, specifically in implementation, a cross-layer connection may be provided between the third convolutional layer and the fifth convolutional layer; and / or a cross-layer connection may be provided between the first convolutional layer and the fourth convolutional layer.
[0098] In some embodiments, specifically, reference may be made to Figure 3 As shown, a cross-layer connection is provided between the third convolutional layer (a low-level convolutional network layer) and the fifth convolutional layer (a high-level convolutional network layer), which may be denoted as the first cross-layer connection.
[0099] There are multiple intermediate convolutional layers such as the fourth convolutional layer, the first Fire module, and the second Fire module between the above-mentioned third convolutional layer and the fifth convolutional layer.
[0100] Through the above-mentioned first cross-layer connection, the features input to the fifth convolutional layer not only include the higher-level features from the second Fire module, but also the lower-level features from the third convolutional layer. In this way, the above two types of features can be obtained and feature fusion can be performed in the above-mentioned fifth convolutional layer, and then extraction processing can be performed based on the fused features to make full use of the information of the lower-level features and reduce the impact of the loss of feature information of the lower-level features during the extraction processing of the intermediate convolutional layer on subsequent character recognition, so as to obtain and output relatively more comprehensive and better-effect features, so that the subsequent fully connected layer and Softmax classifier can obtain more accurate processing results based on the above features.
[0101] In some embodiments, in order to further improve the model accuracy of the preset character recognition model, another cross-layer connection can be set between the first convolutional layer and the fourth convolutional layer, which can be denoted as the second cross-layer connection.
[0102] Through the above-mentioned second cross-layer connection, the features input to the fourth convolutional layer can simultaneously include the output features from the first convolutional layer and the output features from the third convolutional layer. Correspondingly, the fourth convolutional layer can obtain and perform feature fusion on the above two types of features, and then perform extraction processing based on the fused features, so as to obtain and output relatively more comprehensive and better-effect features to the connected first Fire module.
[0103] In some embodiments, during specific implementation, a non-linear function can also be used for one or more convolutional layers in the above-mentioned preset character recognition model to increase the expression ability of the convolutional neural network.
[0104] In some embodiments, during specific implementation, corresponding pooling layers can also be connected after one or more convolutional layers in the above-mentioned preset character recognition model to perform pooling processing on the features output by the convolutional layers first and then input them to the next convolutional layer to reduce the subsequent data processing volume.
[0105] Among them, the pooling processing method adopted by the above-mentioned pooling layer can specifically be the maximum pooling method. Correspondingly, through the pooling processing by the above-mentioned pooling layer, the features can be compressed, the main features can be extracted, and overfitting can be reduced.
[0106] Specifically, corresponding pooling layers can be connected after the first convolutional layer, the second convolutional layer, and the fourth convolutional layer respectively to perform pooling processing on the features output by the above three convolutional layers.
[0107] In some embodiments, further, the features output by one or more convolutional layers in the above-mentioned preset character recognition model can be randomly inactivated to reduce the interaction between different convolutional layers and the dependency between the features output by different convolutional layers, so as to extract features with relatively better effects.
[0108] In some embodiments, during specific implementation, the server can input the target image into a preset character recognition model and run the preset character recognition model. Accordingly, the specific low-level convolutional network layer and high-level convolutional network layer in the preset character recognition model can be used to perform corresponding feature extraction processing with the fire module to obtain relatively comprehensive, high-precision, and good-effect features; then the features are input into the fully connected layer and the Softmax classifier for processing; and finally the corresponding processing results are output.
[0109] In some embodiments, in order to obtain a more accurate processing result, during the specific implementation, the target image may be preprocessed to obtain a preprocessed target image with better effect after invalid information is removed. Then, the preset character recognition model is called to process the preprocessed target image to obtain the corresponding processing result. The preprocessing may specifically include noise reduction processing, batch normalization processing, segmentation processing, etc.
[0110] In some embodiments, during specific implementation, the server may first perform noise reduction processing on the acquired target image to remove noise interference in the target image and obtain a denoised, relatively pure target image; further, the position of the text characters in the target image may be first located; then, according to the position of the text characters, the target image may be divided into a plurality of sub-images arranged in sequence, wherein each sub-image contains a text character; and then, the plurality of sub-images arranged in sequence may be input into the input values of the preset character recognition model for processing to obtain a relatively accurate processing result.
[0111] S203: Determine the target text characters according to the processing result.
[0112] In some embodiments, the above processing result may specifically include a pending character with a higher probability identified based on a preset character recognition model, and a score value corresponding to the pending character.
[0113] In some embodiments, during specific implementation, the undetermined characters with the highest score values may be screened out as the target text characters according to the processing results.
[0114] In some embodiments, after determining the target text character according to the processing result, the specific implementation of the method may further include: performing specific data processing according to the target text character.
[0115] Specifically, for example, in the scenario of handling banking business, an electronic file can be created for the bill according to the target text characters recognized in the bill. In addition, the fund data in the corresponding payer account and payee account can be updated according to the key information such as the payer information, payee information, and drawn amount in the target text characters.
[0116] In some embodiments, before obtaining the target image to be processed, a sample image involving a complex recognition scenario such as text characters with overlap can be obtained, and a preset character recognition model that meets the requirements can be established and trained based on the above sample image.
[0117] In some embodiments, before obtaining the target image to be processed, when the method is specifically implemented, the following content may further be included:
[0118] S1: Construct an initial model; wherein, the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial fire module, and a cross-layer connection is also provided between the initial low-level convolutional network layer and the initial high-level convolutional network layer;
[0119] S2: Obtain a sample image; wherein, the sample image includes text characters with overlap;
[0120] S3: According to the sample image, establish a training set and a test set; and label the sample images in the training set to obtain the labeled training set;
[0121] S4: Use the labeled training set and the test set to train the initial model to obtain a preset character recognition model that meets the requirements.
[0122] Through the above embodiments, a preset character recognition model with high accuracy and capable of supporting the recognition of difficult-to-identify text characters such as those with overlap can be pre-trained and established.
[0123] In some embodiments, when specifically constructing the initial model, reference can be made to Figure 3 As shown, an initial fire module including an initial first fire module and an initial second fire module connected in series is introduced and set in the initial model. At the same time, a cross-layer connection is also introduced to connect the initial low-level convolutional network layer and the initial high-level convolutional network layer to obtain an initial model with an improved model structure. Subsequently, a preset character recognition model that meets the requirements can be trained based on the above initial model with an improved model structure.
[0124] In some embodiments, when specifically implementing the above obtaining of the sample image, the following content may be included:
[0125] S1: collecting first image data containing text characters;
[0126] S2: performing expansion processing on the first picture data to obtain second picture data;
[0127] S3: Segment the second image data according to the text characters to obtain a plurality of third image data;
[0128] S4: Filtering out sample images containing overlapping text characters from the plurality of third image data.
[0129] Through the above embodiments, sample images with satisfactory quantity and quality can be obtained.
[0130] In some embodiments, when segmenting, the position of the text character can be first located from the second image data; then the second image data is segmented according to the position of the text character to obtain a plurality of third image data. The third image data thus obtained may be a small image containing only one normal text character, or may be a small image containing only one abnormal text character with overlap.
[0131] Specifically, for example, through searching, a text string as shown below is found in the second image data: "Today's weather is sunny". The position of each text character in the text string in the second image can be determined in sequence, and then the second image data is divided into five small images arranged in sequence according to the position of each text character, namely, a small image containing "today", a small image containing "day", a small image containing "sky", a small image containing "air", and a small image containing "sunny", and then the five small images can be determined as the corresponding third picture data.
[0132] In some embodiments, considering that the collected effective first picture data is often relatively limited, in order to be able to train a preset character recognition model with better effect and higher accuracy, the first picture data may be expanded first. Specifically, the first picture data may be expanded into the second picture data by performing one or more processing such as translation processing, noise addition processing, noise reduction processing, etc. on the first picture data.
[0133] In some embodiments, during specific implementation, each text character in the second image data can be detected first; then the second image data can be segmented accordingly based on the text characters to ensure that each segmented image data only contains a single text character, thereby obtaining third image data that meets the requirements.
[0134] In some embodiments, during specific implementation, for complex recognition scenarios with overlapping text characters, third picture data containing overlapping text characters can be specifically selected from the above-mentioned multiple third picture data as the sample image.
[0135] In some embodiments, after obtaining the sample image, during the specific implementation of the method, the following content may further be included: calculating the average value and variance of the sample image according to the sample image; performing batch normalization processing on the sample image according to the average value and variance of the sample image.
[0136] Through the above embodiments, the image data of the sample image can be unified into the same scale range first, and then subsequent model training can be performed, thereby effectively reducing the error impact on subsequent model training caused by differences in scale and other factors of the image data of different sample images, and improving the training accuracy during subsequent model training.
[0137] In some embodiments, during specific implementation, the average value of the sample image can be calculated according to the following formula:
[0138]
[0139] where β is the average value of the sample image, x i is the image data value of the sample image numbered i, m is the total number of sample images, and i is the number of the sample image.
[0140] In some embodiments, during specific implementation, the variance of the sample image can be calculated according to the following formula:
[0141]
[0142] where γ 2 is the variance of the sample image.
[0143] In some embodiments, during specific implementation, the sample image can be normalized according to the following formula:
[0144]
[0145] where ω i is the image data value after batch normalization of the sample image numbered i, and ε is a small positive number used to avoid division by zero.
[0146] Through the above embodiments, the image data values of different sample images can be normalized and unified into the normal distribution of (0, 1), thereby reducing the impact of excessive differences in the image data of the sample images on model training; at the same time, it can also reduce the amount of calculation involved in the subsequent model training process, accelerate the convergence of the model, and improve the training efficiency.
[0147] In some embodiments, during specific implementation, according to a preset ratio parameter, sample images that meet the preset ratio parameter (for example, 70%, etc.) can be randomly selected from multiple sample images as the training set, and the remaining sample images can be used as the test set.
[0148] In some embodiments, during specific annotation, the true characters of the overlapping text characters included in the sample images in the training set can be annotated on the sample images in the training set, so as to obtain the annotated training set.
[0149] In some embodiments, when specifically training a preset character recognition model, the method may further include the following: determining the receptive field range of the feature; according to the receptive field range of the feature, adjusting the size parameters of the convolution kernels used in the initial low-level convolutional network layer and the initial high-level convolutional network layer.
[0150] Among them, the above-mentioned receptive field can be specifically understood as the area of the input image that the convolutional neural network feature can see.
[0151] During specific adjustment, when it is determined that the receptive field range of the feature of the input convolutional network layer is relatively large, a convolution kernel with a relatively large size can be used as the convolution kernel used in this convolutional network layer to preferentially extract global features. On the contrary, when it is determined that the receptive field range of the feature of the input convolutional network layer is relatively small, a convolution kernel with a relatively small size can be used as the convolution kernel used in this convolutional network layer to preferentially extract local features.
[0152] Through the embodiments, by specifically adjusting the convolution kernel according to the receptive field in the above manner, a preset character recognition model with relatively high accuracy can be trained relatively more efficiently.
[0153] As can be seen from the above, based on the text character recognition method provided in the embodiments of this specification, before specific implementation, a preset character recognition model including at least a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and with a cross-layer connection set between the low-level convolutional network layer and the high-level convolutional network layer can be pre-trained and established; during specific implementation, after obtaining the target image to be processed, the above-mentioned preset character recognition model can be called to process the target image to obtain the corresponding processing result; then according to the processing result, the target text characters included in the target image can be recognized and determined. Thus, by calling the above-mentioned preset character recognition model that supports multi-scale feature extraction and has good effects, complex recognition scenarios such as overlapping text characters can be adapted, and the text characters in the image can be accurately and efficiently recognized and determined, reducing recognition errors and improving the accuracy of character recognition.
[0154] See Figure 5As shown in the figure, the embodiments of the present specification also provide another method for recognizing text characters. Specifically, when this method is implemented, it may include the following content:
[0155] S501: Obtain a target image to be processed; wherein, the target image contains target text characters to be recognized;
[0156] S502: Call a preset character recognition model to process the target image to obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module;
[0157] S503: Determine the target text characters according to the processing result.
[0158] As can be seen from the above, based on the text character recognition method provided by the embodiments of the present specification, before specific implementation, a preset character recognition model including at least a low-level convolutional network layer, a high-level convolutional network layer, and a fire module can be pre-trained and established; during specific implementation, after obtaining the target image to be processed, the above preset character recognition model can be called to process the target image to obtain a corresponding processing result; then, according to the processing result, the target text characters included in the target image can be recognized and determined. Thus, by calling the above preset character recognition model, it is possible to apply to complex recognition scenarios such as overlapping text characters, and accurately and efficiently recognize and determine the text characters in the image.
[0159] Refer to Figure 6 As shown in the figure, the embodiments of the present specification also provide a method for establishing a preset character recognition model. Specifically, when this method is implemented, it may include the following content:
[0160] S601: Construct an initial model; wherein, the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial fire module, and a cross-layer connection is also set between the initial low-level convolutional network layer and the initial high-level convolutional network layer;
[0161] S602: Obtain sample images; wherein, the sample images contain overlapping text characters;
[0162] S603: Establish a training set and a test set according to the sample images; and label the sample images in the training set to obtain a labeled training set;
[0163] S604: Use the labeled training set and the test set to train the initial model to obtain a preset character recognition model that meets the requirements.
[0164] Through the above embodiments, a preset character recognition model with high recognition accuracy and good effect can be trained, which is applicable to complex recognition scenarios such as overlapping text characters.
[0165] The embodiments of this specification also provide a method for establishing a preset character recognition model. Specifically, when the method is implemented, it may include the following: constructing an initial model, where the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial fire module; obtaining sample images, where the sample images include overlapping text characters; establishing a training set and a test set according to the sample images, and annotating the sample images in the training set to obtain an annotated training set; training the initial model using the annotated training set and the test set to obtain a preset character recognition model that meets the requirements.
[0166] The embodiments of this specification also provide a server, including a processor and a memory for storing instructions executable by the processor. Specifically, when the processor executes the instructions, it may perform the following steps: obtaining a target image to be processed, where the target image includes target text characters to be recognized; calling a preset character recognition model to process the target image to obtain a corresponding processing result, where the preset character recognition model at least includes a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer; determining the target text characters according to the processing result.
[0167] To be able to complete the above instructions more accurately, refer to Figure 7 As shown, the embodiments of this specification also provide another specific server. The server includes a network communication port 701, a processor 702, and a memory 703. The above structures are connected by internal cables so that each structure can perform specific data interactions.
[0168] The network communication port 701 can specifically be used to obtain a target image to be processed, where the target image includes target text characters to be recognized.
[0169] The processor 702 can specifically be used to call a preset character recognition model to process the target image to obtain a corresponding processing result, where the preset character recognition model at least includes a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer; determining the target text characters according to the processing result.
[0170] The memory 703 can be specifically used to store corresponding instruction programs.
[0171] In this embodiment, the network communication port 701 can be bound to different communication protocols, so as to send or receive different data. For example, the network communication port can be a port responsible for web data communication, or a port responsible for FTP data communication, or a port responsible for mail data communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.
[0172] In this embodiment, the processor 702 can be implemented in any suitable manner. For example, the processor can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuit (ASIC), programmable logic controller, and embedded microcontroller, etc. This specification does not make any limitations.
[0173] In this embodiment, the memory 703 can include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit without a physical form but with a storage function is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory module, a TF card, etc.
[0174] An embodiment of this specification also provides a computer storage medium for an identification method based on the above text characters. The computer storage medium stores computer program instructions, which when executed implement: obtaining a target image to be processed; wherein the target image contains target text characters to be recognized; calling a preset character recognition model to process the target image to obtain a corresponding processing result; wherein the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer; determining the target text characters according to the processing result.
[0175] In this embodiment, the above storage medium includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards specified by the communication protocol and is used as an interface for network connection communication.
[0176] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained in contrast with other embodiments and will not be elaborated here.
[0177] Refer to Figure 8 As shown, at the software level, an embodiment of this specification also provides an apparatus for recognizing text characters. The apparatus can specifically include the following structural modules:
[0178] An acquisition module 801, which can specifically be used to acquire a target image to be processed; wherein, the target image contains target text characters to be recognized;
[0179] An invocation module 802, which can specifically be used to invoke a preset character recognition model to process the target image and obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer;
[0180] A determination module 803, which can specifically be used to determine the target text characters according to the processing result.
[0181] It should be noted that the units, apparatuses, or modules, etc. illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above apparatuses, they are divided into various modules according to functions and described separately. Of course, when implementing this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The apparatus embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.
[0182] As can be seen above, the text character recognition device provided in the embodiments of this specification can be well adapted to complex recognition scenarios such as overlapping text characters, accurately and efficiently recognize and determine the text characters in the image, reduce recognition errors, and improve the accuracy of character recognition.
[0183] In a specific scenario example, the text character recognition method provided in this specification can be applied to recognize and obtain the text characters filled in by users on bills when handling business at a bank. For the specific implementation process, please refer to the following content.
[0184] Specifically, refer to Figure 9 As shown, a convolutional neural network model (e.g., a preset character recognition model) based on a convolutional neural network that supports the recognition of overlapping handwritten characters on bank bills can be trained first according to the following steps.
[0185] Step 1: Data preparation.
[0186] In this step, various types of bank bills can be collected first, and they can be photographed from multiple angles and under different lighting conditions by a camera to convert various types of paper bills into electronic images (e.g., the first picture data); and supplemented through various image processing methods such as translation and noise (to obtain the second picture data); then the characters in the sample image can be segmented to obtain multiple images containing single characters (e.g., the third picture data); the images containing overlapping characters can be screened out from them (e.g., the sample images); and a part of the overlapping character images can be labeled as the training set, and the other part is not labeled as the test set.
[0187] Step 2: Data preprocessing.
[0188] In this step, considering that the image data obtained in Step 1 may have inconsistent sizes or contain factors such as noise, resulting in large differences in image quality, and the quality of the image directly affects the accuracy of subsequent character recognition. Therefore, the image data can be preprocessed before model training to eliminate the irrelevant information in the image, retain the real information in the image, enhance the detectability of relevant information; and simplify the data to the greatest extent, thereby improving the reliability of feature extraction, image segmentation, matching, and recognition, and obtaining the processed training set and the processed test set.
[0189] Specifically, the obtained image data can be batch-normalized to transform the image into a normal distribution of (0, 1), reducing the subsequent calculation amount and accelerating the convergence of the model network.
[0190] The specific calculation process is as follows, where m represents the number of image samples, and x i represents the data value of the i-th image sample.
[0191] (1) First, the mean value β can be calculated according to the following formula:
[0192] (2) Next, the variance γ can be calculated according to the following formula 2 :
[0193] (3) Then, the image samples can be batch-normalized according to the following formula:
[0194] After preprocessing the image samples in the above manner, the preprocessed image samples are then used for subsequent model training.
[0195] Step 3, model training.
[0196] In this step, when specifically constructing the model, a new convolutional neural network model can be obtained by introducing the fire module in the compressed convolutional neural network. Based on this model, the shallow convolutional neural network can be deepened both in depth and width at the same time, different convolutional features can be cascaded, the multi-scale features of the image can be improved, the feature expression ability of the shallow convolutional neural network is enriched, and the recognition rate of overlapping handwritten characters is improved. Further, when constructing the model, the idea of cross-layer connection is also introduced to fuse the low-level features and high-level features, so that the features extracted by the low-level network layer (for example, the low-level convolutional network layer) can be fully utilized.
[0197] In addition, by using random inactivation for the convolutional layer, the interaction between neurons is reduced, and the dependence between features is reduced to further improve the accuracy of recognizing overlapping handwritten characters.
[0198] The following combines Figure 3 and Figure 4 , and specifically illustrates the process of model training. The processed training set can be used for model training first; then the processed test set is used for testing to obtain the trained model.
[0199] Specifically, as shown in Figure 3 , a convolutional neural network-based support model for recognizing overlapping handwritten characters on banknotes constructed in this scenario example may include: 1 input layer, 5 convolutional layers, 2 fire modules, 3 fully connected layers, 1 output layer, and cross-layer connection.
[0200] Among them, the input layer is mainly used to input the training set processed in step 2. The data in the training set can specifically be an image with a size of n×m×H, where H is the number of channels.
[0201] In this scenario example, the convolutional layer is used to perform local calculations on the input image by using the corresponding convolutional kernel. Among them, convolution can be specifically understood as a special linear operation based on matrices, and its calculation formula is as follows:
[0202] S(i,j) = (X * W)(i,j) = ∑ m ∑ n x(i + m,j + n)w(m , n)
[0203] Among them, X is the input matrix, and W is the convolutional kernel matrix.
[0204] Generally, by using convolutional calculations through low-level convolutional layers (for example, low-level convolutional network layers), low-level features (such as color, size, etc.) can be extracted; through high-level convolutional layers (for example, high-level convolutional network layers), the features extracted by the low-level layers can be further subjected to non-linear transformations to extract high-level features (such as texture, etc.).
[0205] In addition, based on the above convolutional layer, local connection and weight sharing can also be realized. Among them, local connection specifically means that the nodes of a certain convolutional layer are only connected to some neurons in the previous layer, which is used to extract local features; weight sharing specifically means using the same convolutional kernel to convolve the feature map. These two characteristics can reduce the number of parameters and the computational complexity.
[0206] In this scenario example, it is also considered that when different convolutional kernel sizes are used in the convolutional layer, the receptive fields corresponding to neurons in the overlapping handwritten character image are different, that is, the size of the area mapped by the pixel points on the output feature map in the original image is different. Generally, the larger the size of the convolutional kernel, the larger the receptive field corresponding to the original image of the overlapping handwritten character, and the larger the number of parameters of the network. Specifically, when the receptive field is small, the neural network extracts the local features of the overlapping handwritten character image; when the receptive field is large, the neural network extracts the global features of the overlapping handwritten character image. Therefore, the convolutional kernel used in the convolutional layer can be flexibly adjusted according to the receptive field.
[0207] In this scenario example, in order to achieve the purpose of extracting richer features of the overlapping handwritten character image while reducing calculations, convolutional kernels of 5×5, 3×3, and 1×1 can be specifically used. The feature map output by the input layer will first pass through 4 convolutional layers of 5×5, 3×3, 1×1, and 3×3 (i.e., the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer) in sequence, and batch normalization processing is performed on the output of each convolutional layer to make the input of each layer of the convolutional neural network maintain the same distribution.
[0208] In addition, the ReLU non-linear mapping function can also be used for each convolutional layer to increase the expressive power of the neural network. To reduce the number of parameters, specifically, pooling can be performed on the feature maps of overlapping handwritten characters output by the first convolutional layer, the second convolutional layer, and the fourth convolutional layer. Among them, the pooling layer can compress the feature maps by using the maximum pooling method to extract the main features and reduce overfitting.
[0209] In this scenario example, the fire module can specifically include: a compression layer and an expansion layer. Among them, the compression layer can compress the feature maps through a series of 1×1 convolutions; the expansion layer performs convolution operations using 1×1 and 3×3 convolutions respectively, and then cascades the outputs of the two to obtain an output.
[0210] Refer to Figure 4 As shown. Suppose the input feature map of the fire module is H×W×M, and the calculation process when processing with the fire module is as follows: First, the feature map will be processed by the compression layer with a convolution kernel size of 1×1, and the size of the obtained feature map is H×W×s1. Secondly, this feature map can be input into the expansion layer and convolved through a 1×1 convolutional layer and a 3×3 convolutional layer respectively; finally, the two results are cascaded, and the size of the finally obtained feature map is H×W×(e1+e3). Among them, s1, e1, and e3 respectively represent the number of convolution kernels of the corresponding convolutional layers, and can also represent the dimensions of the corresponding output feature maps, and satisfy the following relationship: e1 = e3 = 4s1, s1 < M.
[0211] Based on this fire module, the traditional convolutional neural network can be deepened both in depth and width, increasing the robustness of the network, enabling the network to extract richer features, and improving the accuracy of model recognition.
[0212] In this scenario example, the fully connected layer can be located at the tail of the convolutional neural network. Among them, each neuron is connected to all neurons in the adjacent layer, integrating the discriminative key feature information extracted in the convolutional layer or the pooling layer, and then transmitting it to the classifier (Softmax) for classification processing.
[0213] Specifically, assume that the vector composed of input nodes is x with a dimension of N, and the vector composed of output nodes is y with a dimension of M. Then the calculation formula of the fully connected layer is expressed in the following form:
[0214] y = Wx
[0215] Among them, W is a weight matrix of N×M dimensions.
[0216] In this scenario example, the output layer can use a Softmax classifier, which is mainly used for multi-classification.
[0217] Specifically, the classifier can use the log loss function to calculate the error between the predicted value and the true value, and then use the gradient descent method to update the relevant parameters of the network according to the error to complete the training process of the network model.
[0218] Among them, the formula of the above log loss function can specifically refer to the following formula:
[0219] Loss = -(y i logs i +(1 - y i )log(1 - s i ))
[0220] Among them, y i represents the true value of the i-th category, and the value can only be 0 or 1; s i represents the probability that the input data belongs to a certain category, that is, the prediction result. The calculation formula of s i can specifically refer to the following formula:
[0221]
[0222] Among them, z i represents the output of the i-th neuron.
[0223] In this scenario example, the cross-layer connection specifically refers to that the lower network layer crosses the network layer directly connected to it (for example, the middle network layer) and is directly connected to the higher network layer, so that the lower-layer features and higher-layer features can be fused in the higher network layer, so that the feature information extracted by the lower network layer can be fully utilized to improve the accuracy of overlapping handwritten character recognition.
[0224] Step four, output the result.
[0225] Specifically in use, the test set in step two can be input into the convolutional neural network model trained in step three, and the maximum value of s i calculated by the log loss function in the Softmax classifier is used as the result of image recognition.
[0226] Through the above scenario examples, the recognition method for text characters provided in this specification is verified. First, by introducing the fire module in the compressed convolutional neural network, the shallow convolutional neural network is deepened both in depth and width simultaneously, enabling the cascading of convolutional features at different scales, extracting multi-scale features of the image, enriching the feature expression ability of the shallow convolutional neural network, and improving the accuracy of recognizing overlapping handwritten characters. Second, by introducing the idea of cross-layer connection, the low-level features are fused with the high-level features, enabling the full utilization of the features extracted by the low-level network layer, improving the accuracy of the convolutional neural network in recognizing overlapping handwritten characters on bank bills, and thus enabling the automatic, accurate, and efficient recognition and extraction of text characters that support overlapping handwritten characters, improving the business processing efficiency.
[0227] Although this specification provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The step order listed in the embodiments is only one way among many step execution orders and does not represent the only execution order. When the actual device or client product executes, it can be executed in the method order shown in the embodiments or the drawings or in parallel (such as in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product or device including the said elements is not excluded. Words such as first, second, etc. are used to denote names and do not represent any specific order.
[0228] Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and the structures within the hardware component.
[0229] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0230] From the description of the above embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this specification can essentially be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this specification.
[0231] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. This specification can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0232] Although this specification is depicted through embodiments, those of ordinary skill in the art know that this specification has many variations and changes without departing from the spirit of this specification. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of this specification.
Claims
1. A method for identifying text characters, characterized in that, Including: Obtain a target image to be processed; wherein, the target image contains target text characters to be recognized; the target image is an image collected in a complex recognition scenario with overlapping text characters; Call a preset character recognition model to process the target image to obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is also provided between the low-level convolutional network layer and the high-level convolutional network layer; the low-level convolutional network layer includes: a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in series in sequence; the high-level convolutional network layer includes: a fourth convolutional layer and a fifth convolutional layer; the fire module includes a first fire module and a second fire module connected in series in sequence; the fire module is integrally disposed between the fourth convolutional layer and the fifth convolutional layer; and, a first cross-layer connection is provided between the third convolutional layer and the fifth convolutional layer, and a second cross-layer connection is provided between the fourth convolutional layer and the first convolutional layer; the processing result includes the undetermined characters recognized by the preset character recognition model based on the target image, and the score values corresponding to the undetermined characters; Determine the target text characters according to the processing result; Wherein, the preset character recognition model is a model pre-trained using the segmented third picture data; the third picture data includes small images containing only one normal text character and / or small images containing only one abnormal text character with overlap; during the model training process, random inactivation processing is also performed on the features output by one or more convolutional layers in the preset character recognition model to reduce the interaction between different convolutional layers and reduce the dependence between the features output by different convolutional layers.
2. The method according to claim 1, wherein Before obtaining the target image to be processed, the method further includes: Construct an initial model; wherein, the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial fire module, and a cross-layer connection is also provided between the initial low-level convolutional network layer and the initial high-level convolutional network layer; Obtain a sample image; wherein, the sample image contains overlapping text characters; According to the sample image, establish a training set and a test set; and label the sample images in the training set to obtain a labeled training set; Use the labeled training set and the test set to train the initial model to obtain a preset character recognition model that meets the requirements.
3. The method according to claim 2, characterized in that, Obtaining a sample image includes: Collect first picture data containing text characters; Perform augmentation processing on the first picture data to obtain second picture data; According to the text characters, segment the second picture data to obtain a plurality of third picture data; Select sample images containing overlapping text characters from the plurality of third picture data.
4. The method according to claim 2, wherein The method further includes: Determine the receptive field range of the feature; According to the receptive field range of the feature, the size parameters of the convolution kernels used by the initial low-level convolutional network layer and the initial high-level convolutional network layer are adjusted.
5. The method according to claim 2, wherein After acquiring the sample image, the method further includes: According to the sample image, calculating the average value and variance of the sample image; The sample images are batch-normalized according to the mean and variance of the sample images.
6. The method according to claim 1, wherein The target image includes at least one of the following: a picture containing a bill, a picture containing a certificate, and a picture containing a contract.
7. A method for recognizing text characters, characterized in that, include: Acquire a target image to be processed; wherein the target image contains target text characters to be recognized; the target image is an image acquired in a complex recognition scenario with overlapping text characters; Calling a preset character recognition model to process the target image to obtain a corresponding processing result; wherein the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module; the low-level convolutional network layer includes: a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence; the high-level convolutional network layer includes: a fourth convolutional layer and a fifth convolutional layer; the fire module includes a first fire module and a second fire module connected in sequence; the fire module is arranged as a whole between the fourth convolutional layer and the fifth convolutional layer; and a first cross-layer connection is arranged between the third convolutional layer and the fifth convolutional layer, and a second cross-layer connection is arranged between the fourth convolutional layer and the first convolutional layer; the processing result includes the pending characters recognized by the preset character recognition model based on the target image, and the score value corresponding to the pending characters; Determining the target text character according to the processing result; Among them, the preset character recognition model is a model pre-trained using the third image data obtained by segmentation; the third image data includes a small image containing only one normal text character and / or a small image containing only one abnormal text character with overlap; during the model training process, the features output by one or more convolutional layers in the preset character recognition model are randomly inactivated to reduce the interaction between different convolutional layers and reduce the dependency between the features output by different convolutional layers.
8. A method for establishing a preset character recognition model, characterized in that, include: Construct an initial model; wherein, the initial model at least includes an initial low-level convolutional network layer, an initial high-level convolutional network layer, and an initial Fire module, and a cross-layer connection is also arranged between the initial low-level convolutional network layer and the initial high-level convolutional network layer; the initial low-level convolutional network layer includes: a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in series in sequence; the initial high-level convolutional network layer includes: a fourth convolutional layer and a fifth convolutional layer; the initial Fire module includes a first Fire module and a second Fire module connected in series in sequence; the Fire module is integrally arranged between the fourth convolutional layer and the fifth convolutional layer; and, a first cross-layer connection is arranged between the third convolutional layer and the fifth convolutional layer, and a second cross-layer connection is arranged between the fourth convolutional layer and the first convolutional layer; Obtain a sample image; wherein, the sample image contains overlapping text characters; the sample image includes third picture data obtained by segmentation; the third picture data includes small images containing only one normal text character and / or small images containing only one abnormal text character with overlap; during the model training process, random inactivation processing is also performed on the features output by one or more convolutional layers in a preset character recognition model to reduce the interaction between different convolutional layers and reduce the dependence between the features output by different convolutional layers; According to the sample image, establish a training set and a test set; and label the sample images in the training set to obtain a labeled training set; Use the labeled training set and the test set to train the initial model to obtain a preset character recognition model that meets the requirements; And, during the model training process, random inactivation processing is also performed on the features output by one or more convolutional layers in a preset character recognition model to reduce the interaction between different convolutional layers and reduce the dependence between the features output by different convolutional layers.
9. An apparatus for recognizing text characters, characterized in that, Including: An acquisition module, configured to acquire a target image to be processed; wherein, the target image contains target text characters to be recognized; the target image is an image acquired in a complex recognition scenario of overlapping text characters; A calling module, configured to call a preset character recognition model to process the target image and obtain a corresponding processing result; wherein, the preset character recognition model at least includes: a low-level convolutional network layer, a high-level convolutional network layer, and a fire module, and a cross-layer connection is further arranged between the low-level convolutional network layer and the high-level convolutional network layer; the low-level convolutional network layer includes: a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in series in sequence; the high-level convolutional network layer includes: a fourth convolutional layer and a fifth convolutional layer; the fire module includes a first fire module and a second fire module connected in series in sequence; the fire module as a whole is arranged between the fourth convolutional layer and the fifth convolutional layer; and, a first cross-layer connection is arranged between the third convolutional layer and the fifth convolutional layer, and a second cross-layer connection is arranged between the fourth convolutional layer and the first convolutional layer; the processing result includes a pending character recognized by the preset character recognition model based on the target image, and a score value corresponding to the pending character. A determining module, configured to determine the target text character according to the processing result. Wherein, the preset character recognition model is a model pre-trained by using the third picture data obtained by segmentation; the third picture data includes small images containing only one normal text character and / or small images containing only one abnormal text character with overlap; during the model training process, random inactivation processing is further performed on the features output by one or more convolutional layers in the preset character recognition model to reduce the interaction between different convolutional layers and reduce the dependence between the features output by different convolutional layers.
10. A server, characterized in that, It includes a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, the steps of the method according to any one of claims 1 to 6 are implemented.
11. A computer-readable storage medium, characterized in that, Computer instructions are stored thereon, and when the instructions are executed, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Text recognition method and device, electronic equipment and storage medium
CN110378338A