OCR Recognition Model Training Method, OCR Recognition Method and Related Devices

By performing image block occlusion and reconstruction feature training on unlabeled data, pre-training the OCR recognition model, and using a small amount of labeled data for training the task processing model, the problems of low training efficiency and poor recognition effect of the existing OCR recognition model are solved, and more efficient training and more accurate recognition effect are achieved.

CN114565751BActive Publication Date: 2025-06-27HUI ZE (CHENGDU) NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210192272.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-06-27
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The training efficiency of existing OCR recognition models is low and the recognition effect is poor, mainly due to their dependence on a large amount of labeled data and only considering visual information, resulting in a high misidentification rate.

Method used

By dividing the image samples without labeling data into multiple image blocks, and randomly obstructing partial blocks, preset features are reconstructed using occluded and unoccluded image blocks, the initial feature recognition model is pretrained to obtain the pretrained feature recognition model. Then, a task processing model is built based on the pre-trained model and a small amount of labeled data is used for training to improve the recognition ability of the OCR recognition model.

Benefits of technology

This method does not require a large amount of labeling data, which improves the training efficiency of the OCR recognition model, and reduces the misrecognition rate and improves the recognition effect by extracting more refined visual semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565751B_ABST
    Figure CN114565751B_ABST
Patent Text Reader

Abstract

The present application provides an OCR recognition model training method, an OCR recognition method and related devices. The OCR recognition model training method includes: splitting a first image sample of unlabeled data into a plurality of first image patches, randomly selecting some of the first image patches for occlusion to obtain occluded image patches and unoccluded image patches; using the occluded image patches and the unoccluded image patches, aiming at reconstructing the preset features of the first image sample, pre-training an initial feature recognition model pre-constructed with an encoder and a first decoder; constructing a task processing model based on the encoder and a second decoder in the pre-trained feature recognition model; splitting a second image sample of labeled data into a plurality of second image patches; training the task processing model with a plurality of second image patches and the token sequence included in the second image sample to obtain an OCR recognition model. The present application does not require a large amount of labeled data, has high model training efficiency, and at the same time, the training method enables the OCR recognition model to have high recognition ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of OCR recognition, and particularly to an OCR recognition model training method, an OCR recognition method and related devices. Background Art

[0002] OCR (Optical Character Recognition) is to use optical technology and computer technology to read the text printed or written on paper and convert it into a format that can be accepted by a computer and understood by people. Now it also includes the recognition of text in natural scenes.

[0003] Currently, the mainstream OCR recognition models mainly use the CNN network as the backbone network and CTC as the decoder. Among them, a large number of image samples with labeled data are required for model training when constructing the OCR recognition model.

[0004] However, labeling data requires a large amount of time and resources, which greatly reduces the training efficiency of the OCR recognition model. Moreover, the existing OCR recognition models only consider visual information, are prone to misrecognition, and have poor recognition effects. Summary of the Invention

[0005] In view of this, the present application provides an OCR recognition model training method, an OCR recognition method and related devices, which are used to solve the problems of low training efficiency and poor recognition effect of the existing OCR recognition model. The technical solutions are as follows:

[0006] An OCR recognition model training method includes:

[0007] Cut a first image sample without labeled data into a plurality of first image blocks, and randomly select some first image blocks from the plurality of first image blocks for occlusion to obtain occluded image blocks and unoccluded image blocks, where the number of occluded image blocks is greater than the number of unoccluded image blocks;

[0008] Use the occluded image blocks and unoccluded image blocks to pre-train the initially constructed feature recognition model including an encoder and a first decoder with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model;

[0009] Construct a task processing model based on the encoder and a second decoder in the pre-trained feature recognition model;

[0010] Cut a second image sample with labeled data into a plurality of second image blocks;

[0011] Use the plurality of second image blocks and the word piece sequence included in the second image sample to train the task processing model, and use the obtained model as the OCR recognition model.

[0012] Optionally, using the occluded image patches and the unoccluded image patches, with the aim of reconstructing the preset features of the first image sample, pre-train the initially constructed initial feature recognition model including an encoder and a first decoder to obtain the preset features of the pre-trained feature recognition model, including:

[0013] Input the unoccluded image patches into the encoder included in the initial feature recognition model to obtain the visual semantic information of the unoccluded image patches;

[0014] Input the visual semantic information of the unoccluded image patches and the occluded image patches into the first decoder included in the initial feature recognition model to obtain the corresponding preset features of the reconstructed unoccluded image patches and occluded image patches respectively;

[0015] According to the preset features corresponding to the reconstructed unoccluded image patches and occluded image patches respectively, and the preset features extracted from the first image sample, pre-train the parameters of the initial feature recognition model to obtain the pre-trained feature recognition model.

[0016] Optionally, the preset features include one or more of the following features: Histogram of Oriented Gradients (HOG) feature, Convolutional Neural Network (CNN) feature, Local Binary Pattern (LBP) feature, Haar-like feature, ORB feature, Speeded Up Robust Features (SURF), and Scale-Invariant Feature Transform (SIFT) feature.

[0017] Optionally, use multiple second image patches and the word piece sequence included in the second image sample to train the task processing model, and use the obtained model as the OCR recognition model, including:

[0018] Input multiple second image patches into the task processing model to obtain the OCR recognition results corresponding to the multiple second image patches;

[0019] Based on the OCR recognition results corresponding to the multiple second image patches and the word piece sequence included in the second image sample, train the parameters of the task processing model to obtain the OCR recognition model.

[0020] Optionally, the method for obtaining the image sample includes:

[0021] Obtain the original sample;

[0022] Perform image conversion processing on the original sample to obtain the processed original sample, where the image conversion processing includes one or more of the following processes: random rotation, Gaussian blur, image dilation, image erosion, downsampling, and adding underlines;

[0023] Use the original sample and the processed original sample as the image sample.

[0024] An OCR recognition method, applied to the OCR recognition model as described in any of the above items, includes:

[0025] Obtain the image to be recognized;

[0026] Divide the image to be recognized into multiple image blocks to be recognized;

[0027] Input the multiple image blocks to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

[0028] An OCR recognition model training device, comprising: a first segmentation module, a first training module, a model construction module, a second segmentation module, and a second training module;

[0029] The first segmentation module is used to divide the first image sample of the unlabeled data into multiple first image blocks, and randomly select some first image blocks from the multiple first image blocks for occlusion to obtain occluded image blocks and unoccluded image blocks, wherein the number of occluded image blocks is greater than the number of unoccluded image blocks;

[0030] The first training module is used to use the occluded image blocks and unoccluded image blocks to pre-train the initially constructed feature recognition model including an encoder and a first decoder with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model;

[0031] The model construction module is used to construct a task processing model based on the encoder and the second decoder in the pre-trained feature recognition model;

[0032] The second segmentation module is used to divide the second image sample of the labeled data into multiple second image blocks;

[0033] The second training module is used to train the task processing model with multiple second image blocks and the word piece sequence included in the second image sample, and the obtained model is used as the OCR recognition model.

[0034] An OCR recognition device, applied to the OCR recognition model as described above, comprising: an acquisition module, a third segmentation module, and a recognition module;

[0035] The acquisition module is used to acquire the image to be recognized;

[0036] The third segmentation module is used to divide the image to be recognized into multiple image blocks to be recognized;

[0037] The recognition module is used to input the multiple image blocks to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

[0038] A data processing device, comprising a memory and a processor;

[0039] A memory for storing programs;

[0040] A processor for executing the program to implement the OCR recognition model training method as described in any one of the above, or each step of the above OCR recognition method.

[0041] A readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the OCR recognition model training method as described in any one of the above, or each step of the above OCR recognition method.

[0042] As can be seen from the above technical solutions, the OCR recognition model training method provided by this application first divides the first image samples of unlabeled data into multiple first image blocks, randomly selects some of the first image blocks from the multiple first image blocks for occlusion to obtain occluded image blocks and unoccluded image blocks, and then uses the occluded image blocks and unoccluded image blocks to pre-train the initially constructed feature recognition model including an encoder and a first decoder with the goal of reconstructing the preset features of the first image samples to obtain a pre-trained feature recognition model. Then, a task processing model is constructed based on the encoder and the second decoder in the pre-trained feature recognition model. After that, the second image samples of labeled data are divided into multiple second image blocks, and finally, the task processing model is trained using the multiple second image blocks and the token sequence included in the second image samples, and the obtained model is used as the OCR recognition model. This application can train the initially constructed feature recognition model based on a large number of first image samples of unlabeled data to obtain a pre-trained feature recognition model. Based on the encoder and the second decoder in this pre-trained feature recognition model, a task processing model can be constructed. Then, based on a small number of second image samples of labeled data and the token sequence included in the second image samples, the task processing model is trained, and an OCR recognition model with higher recognition ability can be obtained. Since the training process of this application does not require a large amount of labeled data, the training efficiency of the OCR recognition model is improved. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0044] Figure 1 It is a flowchart showing the OCR recognition model training method provided by an embodiment of this application;

[0045] Figure 2a It is a schematic diagram of the image samples provided by an embodiment of this application;

[0046] Figure 2b Schematic diagram of occluded image blocks and unoccluded image blocks provided by an embodiment of the present application;

[0047] Figure 3 Schematic diagram of the encoder structure provided by an embodiment of the present application;

[0048] Figure 4 Schematic diagram of the decoder structure provided by an embodiment of the present application;

[0049] Figure 5 Schematic diagram of the training process of the initial feature recognition model provided by an embodiment of the present application;

[0050] Figure 6 Schematic diagram of the process of the OCR recognition method provided by an embodiment of the present application;

[0051] Figure 7 Schematic diagram of the recognition process of the OCR recognition model provided by an embodiment of the present application;

[0052] Figure 8 Schematic diagram of the structure of the OCR recognition model training device provided by an embodiment of the present application;

[0053] Figure 9 Schematic diagram of the structure of the OCR recognition device provided by an embodiment of the present application;

[0054] Figure 10 Hardware structure block diagram of the OCR recognition model training device provided by an embodiment of the present application;

[0055] Figure 11 Hardware structure block diagram of the OCR recognition device provided by an embodiment of the present application. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0057] As introduced in the background art, existing OCR recognition models mainly use a CNN network as the backbone network, use CTC as the decoder, and a large number of labeled image samples are required for model training during the construction process.

[0058] However, labeled data requires a large amount of time and resources and cannot utilize large-scale unlabeled data that is cheap and easily obtainable. Moreover, when performing OCR recognition, this kind of OCR recognition model only considers visual information, which is prone to misrecognition, especially prone to misrecognizing characters with similar shapes, resulting in poor recognition effects. Often, it is necessary to cooperate with an additional character-level language model or a dictionary of similar-shaped characters for post-processing (such as error correction, etc.) to improve the overall accuracy. In addition, due to light changes, object occlusion, or paper wrinkles, some words, phrases, and sentences are occluded, which will also greatly reduce the OCR recognition effect.

[0059] In view of this, the present application provides an OCR recognition model training method, an OCR recognition method, and related devices to solve the above problems. To make those skilled in the art better understand the present application, the OCR recognition model training method provided by the present application will be introduced in detail through the following embodiments.

[0060] Please refer to Figure 1 , which shows a schematic flowchart of the OCR recognition model training method provided by an embodiment of the present application. The OCR recognition model training method may include:

[0061] Step S101: Split a first image sample of unlabeled data into multiple first image patches, and randomly select some first image patches from the multiple first image patches for occlusion to obtain occluded image patches and unoccluded image patches.

[0062] Among them, the number of occluded image patches is greater than the number of unoccluded image patches.

[0063] In this embodiment, a first image sample of a large amount of unlabeled data can be used to train an initial feature recognition model. Among them, for each first image sample, in this step, the first image sample can be split into multiple image patches patch of a fixed size. For the convenience of description, the image patches cut out in this step are defined as first image patches.

[0064] Considering that an image has local similarity and translational invariance and has a lot of redundant information (that is, the semantic density of the image is very low), if only a very small proportion of the first image patches are occluded, when training the initial feature recognition model, the feature recognition model is very likely to be lazy and can simply learn by imagining from neighbor information, ultimately resulting in the failure to learn high-level semantic information. Based on this, it is necessary to randomly select some first image patches from the multiple first image patches for occlusion, and make the number of occluded image patches greater than the number of unoccluded image patches, so as to force the feature recognition model to learn the most refined visual semantic information.

[0065] It should be noted that when randomly occluding the first image patch in this step, the occlusion operation is performed according to the ratio of the number of occluded image patches to the number of unoccluded image patches being a set ratio; optionally, in this step, when randomly occluding, the ratio of the number of occluded image patches to the number of unoccluded image patches can be 3 to 2. For example, taking the first image sample as Figure 2a the image sample shown, the effect after random occlusion can be seen in Figure 2b as shown.

[0066] Of course, in the embodiments of the present application, the ratio of the number of occluded image patches to the number of unoccluded image patches can also be other values. For example, the ratio of the number of occluded image patches to the number of unoccluded image patches is 3 to 1, and the present application does not make specific limitations on this.

[0067] Step S102: Using the occluded image patches and the unoccluded image patches, with the goal of reconstructing the preset features of the first image sample, pre-train the initially constructed initial feature recognition model including an encoder and a first decoder to obtain a pre-trained feature recognition model.

[0068] In this embodiment, first, an initial feature recognition model needs to be constructed based on the encoder and the first decoder.

[0069] Among them, the encoder is used to extract the visual semantic information of the input image. Since the semantic density of the image is very low, by using a reasonable training method to train the initial feature recognition model, the trained encoder (i.e., the encoder in the pre-trained feature recognition model) can extract more refined and complete visual semantic information. Therefore, in this step, the unoccluded image patches are input into the encoder for training, so that the visual semantic information of the unoccluded image patches extracted by the encoder in the pre-trained feature recognition model can represent the visual semantic information of the first image sample.

[0070] Optionally, the above encoder can adopt a transformer structure, and this structure can be seen in Figure 3 as shown, that is, the encoder is composed of N multi-head attention layers and a feed-forward neural network layer. Here, the specific value of N can be determined according to the actual situation.

[0071] The first decoder in the above text is used to reconstruct the preset features corresponding to the occluded image patches and the unoccluded image patches based on the visual semantic information of the unoccluded image patches and the occluded image patches. In this embodiment, the reason for choosing to reconstruct the preset features based on the first decoder instead of reconstructing the pixels of the first image sample is that the pixels contain too many unimportant high-frequency details. Letting the constructed feature recognition model fit these high-frequency details easily causes overfitting. However, the preset features have stability. Letting the constructed feature recognition model fit the preset features enables the feature recognition model to learn better and more high-level semantic information and avoid being trapped in the noise details.

[0072] Optionally, the first decoder can also adopt a transformer structure, and this structure can be referred to Figure 4 as shown, that is, the first decoder is composed of N masked multi-head attention layers, multi-head attention layers, and feed forward neural network layers.

[0073] Optionally, the above preset features can be histogram of oriented gradients features (i.e., HOG features). Since the HOG features can be calculated quickly and conveniently based on the first image sample, the HOG features are selected as the preset features in this embodiment, making this embodiment have the advantages of small computational complexity and fast speed. Of course, the preset features can also be others, such as convolutional neural network features (i.e., CNN features) or other traditional features, such as local binary pattern (LBP) features, Haar-like features (i.e., HAAR features), ORB (Oriented Fast and Rotated Brief) features, speeded up robust features (SURF) features, SIFT (scale-invariant feature transform) features, etc. This application does not limit this.

[0074] It should be noted that in view of the fact that there are similar encoder structures and decoder structures in the prior art, that is, in the encoder and the first decoder of this embodiment, some layers are the same as those in the prior art. To improve the training efficiency of the feature recognition model, during the initialization stage of the feature recognition model, the parameters that have been trained in the prior art can be used as the initial parameters of some layers of the encoder and the first decoder in this embodiment, while the initial parameters of the other layers that do not exist in the prior art need to be randomly initialized.

[0075] After the initial feature recognition model is constructed, this step can use the occluded image patches and unoccluded image patches to pre-train the initial feature recognition model with the goal of reconstructing the preset features of the first image sample, and obtain a pre-trained feature recognition model.

[0076] Step S103: Construct a task processing model based on the encoder and the second decoder in the pre-trained feature recognition model.

[0077] After the pre-trained feature recognition model is trained in the previous step, this step can construct a task processing model based on the encoder and the second decoder in the pre-trained feature recognition model.

[0078] It should be noted that when constructing the task processing model, since the number of input image patches of the task processing model may be different from the number of input image patches of the pre-trained feature recognition model, therefore, Figure 3 "N" of the encoder in the shown pre-trained feature recognition model may be different from "N" of the encoder in the task processing model.

[0079] The above-mentioned second decoder is used to determine the OCR recognition result of the image input to the task processing model based on the visual semantic information output by the encoder in the pre-trained feature recognition model.

[0080] Optionally, the second decoder can also adopt Figure 4 the shown transformer structure, and the specific structure can be referred to the introduction in the above step S102, which will not be elaborated here.

[0081] Step S104: Split the second image sample with labeled data into multiple second image patches.

[0082] In this embodiment, a small number of second image samples with labeled data can be used to train the task processing model. Among them, for each second image sample, the second image sample can be split into multiple image patches patch of a fixed size. To distinguish from the above-mentioned first image patch, the image patches cut out in this step are defined as second image patches.

[0083] It should be noted that since the encoder that can extract more refined and complete visual semantic information has been trained in the above step S102, therefore, in this step, only a second image sample with a number less than that of the first image sample needs to be selected, and the task processing model is trained based on the second image patches corresponding to the small number of second image samples, and an OCR recognition model with higher recognition ability can be trained.

[0084] Step S105: Use multiple second image patches and the token sequence included in the second image sample to train the task processing model, and use the obtained model as the OCR recognition model.

[0085] As described above, the task processing model is constructed based on the encoder and the second decoder in the pre-trained feature recognition model. In this step, the encoder in the task processing model is used to obtain the visual semantic information of multiple second image patches, and the second decoder is used to obtain the OCR recognition result of the second image sample based on the visual semantic information of the multiple second image patches.

[0086] Optionally, the process of the second decoder obtaining the OCR recognition result of the second image sample may include: the second decoder obtains each wordpiece in the second image sample based on the visual semantic information of the multiple second image patches, and then predicts the wordpiece sequence in the second image sample based on the obtained wordpieces, and this wordpiece sequence is the OCR recognition result.

[0087] The wordpiece sequence included in the above second image sample refers to the annotation data corresponding to the second image sample. For example, if the second image sample is Figure 2a as shown, the wordpiece sequence included in this second image sample may be "SUBWAY".

[0088] The OCR recognition model training method provided by this application first divides the first image sample without annotation data into multiple first image patches, randomly selects some first image patches from the multiple first image patches for occlusion to obtain occluded image patches and unoccluded image patches, and then uses the occluded image patches and unoccluded image patches to pre-train the pre-constructed initial feature recognition model including an encoder and a first decoder with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model. Then, a task processing model is constructed based on the encoder and the second decoder in the pre-trained feature recognition model. After that, the second image sample with annotation data is divided into multiple second image patches. Finally, the task processing model is trained using the multiple second image patches and the wordpiece sequence included in the second image sample, and the obtained model is used as the OCR recognition model. This application can train the pre-constructed initial feature recognition model based on a large number of first image samples without annotation data to obtain a pre-trained feature recognition model. Based on the encoder and the second decoder in this pre-trained feature recognition model, a task processing model can be constructed. Then, by training the task processing model based on a small number of second image samples with annotation data and the wordpiece sequence included in the second image sample, an OCR recognition model with higher recognition ability can be obtained. Since the training process of this application does not require a large amount of annotation data, it saves human resources and improves the training efficiency of the OCR recognition model.

[0089] In addition, since the encoder can extract more complete and refined visual semantic information of multiple second sub-images, the OCR recognition model can obtain a higher-accuracy OCR recognition result without additional post-processing by a language model.

[0090] To enable those skilled in the art to better understand the training process of the initial feature recognition model in the embodiments of the present application, the following embodiment describes the process of "Step S102. Using the occluded image patches and the unoccluded image patches, aiming at reconstructing the preset features of the first image sample, pre-training the initially constructed initial feature recognition model including an encoder and a first decoder to obtain a pre-trained feature recognition model".

[0091] Optionally, this process may include:

[0092] A1. Input the unoccluded image patches into the encoder included in the initial feature recognition model to obtain the visual semantic information of the unoccluded image patches.

[0093] In this step, in order to avoid introducing too many occluded first image patches into the encoder included in the initial feature recognition model and causing interference to the training of the encoder included in the initial feature recognition model, only the unoccluded image patches may be input into the encoder included in the initial feature recognition model to obtain the visual semantic information of the unoccluded image patches.

[0094] Refer to Figure 5 the schematic diagram of the training process of the initial feature recognition model shown, Figure 5 in which the initial feature recognition model is constructed based on the encoder Encoder and the first decoder Decoder-1. In this step, the six unoccluded image patches shown in Figure 2b may be input into the encoder (Encoder) included in the initial feature recognition model to train the encoder included in the initial feature recognition model based on the six unoccluded image patches.

[0095] A2. Input the visual semantic information of the unoccluded image patches and the occluded image patches into the first decoder included in the initial feature recognition model to obtain the preset features corresponding to the unoccluded image patches and the occluded image patches respectively.

[0096] Still refer to Figure 5 shown. In this step, the visual semantic information of the unoccluded image patches and the occluded image patches may be arranged in a preset order and then input into the first decoder (Decoder-1) to reconstruct the preset features corresponding to the unoccluded image patches and the occluded image patches respectively by using the first decoder.

[0097] A3. According to the preset features corresponding to the reconstructed unoccluded image patches and occluded image patches respectively, and the preset features extracted from the first image sample, pre-train the parameters of the initial feature recognition model to obtain a pre-trained feature recognition model.

[0098] Still refer to Figure 5As shown, the preset features corresponding to the reconstructed unoccluded image patches and occluded image patches can be combined in a preset order to reconstruct the preset features of the first sample image. Then, in this embodiment, the preset features of the reconstructed first sample image can be compared with the preset features extracted from the first sample image (i.e., the extracted preset features are the standard features or true features). Based on the comparison results, the loss and gradient are calculated, and backpropagation is performed to adjust and update the parameters of the initial feature recognition model. After multiple trainings, the optimal parameters can be obtained, and based on these optimal parameters, the pre-trained feature recognition model can be obtained.

[0099] In this embodiment, the initial feature recognition model is trained based on a large number of unlabeled first image samples. During training, by adjusting the proportion of the occluded part and using a first decoder that can output preset features to help the encoder extract refined visual semantic information, the trained pre-trained feature recognition model becomes a good feature extractor, and the trained encoder can extract more refined and complete visual semantic information.

[0100] Using this framework, this application can utilize a large number of cheap and easily available unlabeled data, eliminate the need for a large number of labeled data, thereby saving a large amount of labor costs, and achieving good recognition results.

[0101] To enable those skilled in the art to better understand the training process of the initial OCR recognition model in the embodiments of this application, in the following embodiment, the process of "Step S105. Use multiple second image patches and the word piece sequence included in the second image sample to train the task processing model, and use the obtained model as the OCR recognition model" is described.

[0102] Optionally, this process may include:

[0103] B1. Input multiple second image patches into the task processing model to obtain the OCR recognition results corresponding to the multiple second image patches.

[0104] Specifically, in this step, multiple second image patches can be input into the encoder included in the task processing model to obtain the visual semantic information of the multiple second image patches output by the encoder included in the task processing model. Then, the visual semantic information of the multiple second image patches is input into the second decoder to obtain the OCR recognition results corresponding to the multiple second image patches. The OCR recognition results corresponding to the multiple second image patches are the OCR recognition results of the second image sample.

[0105] In this step, the second decoder (see the Decoder-2 shown in Figure 7 ) can replace the language model for word piece error correction during OCR recognition, so as to output more accurate OCR recognition results.

[0106] Here, the OCR recognition result refers to the result obtained by recognizing the word piece sequence included in the second sample image. Optionally, for the convenience of understanding when the recognition ends, in the OCR recognition result, an end-of-recognition symbol may be added after the recognized word piece sequence to represent the end of the word piece sequence recognition. For example, "[EOS]" is used as the end-of-recognition symbol.

[0107] B2. Train the parameters of the task processing model based on the OCR recognition results corresponding to multiple second image blocks and the word piece sequence included in the second image sample to obtain an OCR recognition model.

[0108] Corresponding to the previous step, if the OCR recognition result includes an end-of-recognition symbol, then in this step when training the model, a start-of-recognition symbol, such as "[BOS]", needs to be added before the "word piece sequence included in the second image sample". It should be noted that "[BOS]" is only an example, and this start-of-recognition symbol can also be others, and the present application does not limit this.

[0109] Optionally, in this step during training, the word piece sequence included in the second image sample can be input into the second decoder, and the cross-entropy loss function is used to supervise the output of the second decoder. When the task processing model makes a prediction, the second decoder starts iteratively predicting the subsequent word piece (wordpiece) from the "[BOS]" symbol and uses the predicted wordpiece as the input for the next time.

[0110] In this embodiment, the task processing model is trained based on a small number of second image samples with labeled data and the labeled data. Since the task processing model is constructed based on the encoder in the pre-trained feature recognition model that has been trained, the second decoder in the OCR recognition model can perform OCR recognition based on more complete visual semantic information. Therefore, even without the participation of an additional language model for error correction, the OCR recognition model still has a good recognition effect.

[0111] It should be noted that the above OCR recognition model can be easily extended to a multilingual model, only by using a multilingual pre-trained model at the second decoder end; in addition, after obtaining the OCR recognition model, if it is necessary to extend to OCR recognition in other scenarios, such as printed text recognition, handwritten text recognition, natural scene text tasks and other scenarios, only by fine-tuning the parameters of the pre-trained OCR recognition model and re-training, it can quickly converge to the required model, and the training efficiency is higher.

[0112] An embodiment of the present application introduces a method for obtaining an image sample.

[0113] Optionally, the method for obtaining an image sample may include:

[0114] C1. Obtain the original sample.

[0115] C2. Perform image conversion processing on the original sample to obtain the processed original sample.

[0116] It should be understood that only by training the model based on various image samples can the trained model be applicable to more scenarios. Based on this, after obtaining the original sample, this step needs to use data augmentation techniques to obtain the processed original sample by performing various image conversion processes on the original sample.

[0117] Optionally, the image conversion processing includes one or more of the following processes: random rotation, Gaussian blur, image dilation, image erosion, downsampling, and adding underlines. This step can randomly select an image conversion processing method for each original sample with equal probability to perform data augmentation on the original sample.

[0118] C3. Use the original sample and the processed original sample as image samples.

[0119] The image samples obtained in this step can be used as the first image samples or, after annotating the data, as the second image samples. This application does not limit this.

[0120] In summary, through the image conversion processing of the original sample in the embodiments of this application, more types of first image samples and second image samples can be obtained. Based on these more types of first image samples and second image samples for model training, the trained model can be made more general.

[0121] The embodiments of this application also provide an OCR recognition method applied to the above OCR recognition model. Next, the OCR recognition method provided by this application will be introduced in detail through the following embodiments.

[0122] Please refer to Figure 6 , which shows the schematic flowchart of the OCR recognition method provided by the embodiments of this application. The OCR recognition method may include:

[0123] Step S601. Obtain the image to be recognized.

[0124] Step S602. Split the image to be recognized into multiple image blocks to be recognized.

[0125] Step S603. Input the multiple image blocks to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

[0126] For specific reference, see Figure 7Schematic diagram of the recognition process of the OCR recognition model shown. Multiple image blocks to be recognized are input into the OCR recognition model to obtain the OCR recognition result "SUBWAY" output by the OCR recognition model.

[0127] In summary, the embodiments of the present application do not require an additional language model for post-processing, nor do they require cumbersome pre-processing, and can achieve a recognition effect equivalent to or better than that of current mainstream OCR methods only based on the OCR recognition model.

[0128] The embodiments of the present application also provide an OCR recognition model training device. The OCR recognition model training device provided by the embodiments of the present application will be described below. The OCR recognition model training device described below can be correspondingly referred to the OCR recognition model training method described above.

[0129] Please refer to Figure 8 , which shows a schematic structural diagram of the OCR recognition model training device provided by the embodiments of the present application. As Figure 8 shown, the OCR recognition model training device may include: a first segmentation module 801, a first training module 802, a model construction module 803, a second segmentation module 804, and a second training module 805.

[0130] The first segmentation module 801 is configured to segment a first image sample of unlabeled data into multiple first image blocks, and randomly select some first image blocks from the multiple first image blocks for occlusion to obtain occluded image blocks and unoccluded image blocks, wherein the number of occluded image blocks is greater than the number of unoccluded image blocks.

[0131] The first training module 802 is configured to use the occluded image blocks and unoccluded image blocks to pre-train an initial feature recognition model including an encoder and a first decoder constructed in advance with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model.

[0132] The model construction module 803 is configured to construct a task processing model based on the encoder and the second decoder in the pre-trained feature recognition model.

[0133] The second segmentation module 804 is configured to segment a second image sample of labeled data into multiple second image blocks.

[0134] The second training module 805 is configured to train the task processing model with multiple second image blocks and the token sequence included in the second image sample, and use the obtained model as the OCR recognition model.

[0135] The OCR recognition model training device provided by this application first divides the first image sample of unlabeled data into multiple first image patches, randomly selects some of the first image patches from the multiple first image patches for occlusion to obtain occluded image patches and unoccluded image patches, and then uses the occluded image patches and unoccluded image patches to pre-train the initially constructed feature recognition model including an encoder and a first decoder with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model. Then, a task processing model is constructed based on the encoder and the second decoder in the pre-trained feature recognition model. After that, the second image sample of labeled data is divided into multiple second image patches, and finally, the task processing model is trained using the multiple second image patches and the word piece sequence included in the second image sample, and the obtained model is used as the OCR recognition model. This application can train the initially constructed feature recognition model based on a large number of first image samples of unlabeled data to obtain a pre-trained feature recognition model. Based on the encoder and the second decoder in this pre-trained feature recognition model, a task processing model can be constructed. Then, based on a small number of second image samples of labeled data and the word piece sequence included in the second image sample, the task processing model is trained to obtain an OCR recognition model with higher recognition ability. Since the training process of this application does not require a large amount of labeled data, it saves human resources and improves the training efficiency of the OCR recognition model.

[0136] In a possible implementation manner, the above-mentioned first training module 802 may include: a first input sub-module, a second input sub-module, and a first training sub-module.

[0137] Among them, the first input sub-module is used to input the unoccluded image patches into the encoder included in the initial feature recognition model to obtain the visual semantic information of the unoccluded image patches.

[0138] The second input sub-module is used to input the visual semantic information of the unoccluded image patches and the occluded image patches into the first decoder included in the feature recognition model to obtain the preset features corresponding to the unoccluded image patches and the occluded image patches respectively;

[0139] The first training sub-module is used to train the parameters of the initial feature recognition model according to the preset features corresponding to the reconstructed unoccluded image patches and occluded image patches respectively, and the preset features extracted from the first image sample to obtain a pre-trained feature recognition model.

[0140] In a possible implementation manner, in the OCR recognition model training device provided by the embodiments of this application, the preset features include one or more of the following features: Histogram of Oriented Gradients (HOG) feature, Convolutional Neural Network (CNN) feature, Local Binary Pattern (LBP) feature, Haar-like feature, Oriented FAST and Rotated BRIEF (ORB) feature, Speeded-Up Robust Features (SURF) feature, and Scale-Invariant Feature Transform (SIFT) feature.

[0141] In a possible implementation, the above-mentioned second training module 805 may include: a third input sub-module and a second training sub-module.

[0142] Among them, the third input sub-module is used to input multiple second image patches into the task processing model to obtain OCR recognition results corresponding to the multiple second image patches.

[0143] The second training sub-module is used to train the parameters of the task processing model based on the OCR recognition results corresponding to the multiple second image patches and the word piece sequence included in the second image sample to obtain an OCR recognition model.

[0144] In a possible implementation, the OCR recognition model training device provided by the embodiments of the present application may further include an image sample acquisition module.

[0145] Specifically, the image sample acquisition module is used to acquire an original sample, perform image conversion processing on the original sample to obtain a processed original sample, and use the original sample and the processed original sample as image samples, where the image conversion processing includes one or more of the following processes: random rotation, Gaussian blur, image dilation, image erosion, downsampling, and adding underlines.

[0146] The embodiments of the present application also provide an OCR recognition device. The OCR recognition device provided by the embodiments of the present application will be described below. The OCR recognition device described below can be correspondingly referred to the OCR recognition method described above.

[0147] Please refer to Figure 9 , which shows a schematic structural diagram of the OCR recognition device provided by the embodiments of the present application. As Figure 9 shown, the OCR recognition device may include: an acquisition module 901, a third segmentation module 902, and a recognition module 903.

[0148] The acquisition module 901 is used to acquire an image to be recognized.

[0149] The third segmentation module 902 is used to segment the image to be recognized into multiple image patches to be recognized.

[0150] The recognition module 903 is used to input the multiple image patches to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

[0151] The OCR recognition device provided by the present application does not require an additional language model for post-processing, nor does it require cumbersome pre-processing, and can achieve an equivalent or better recognition effect than the current mainstream OCR methods only based on the OCR recognition model.

[0152] The embodiments of the present application also provide an OCR recognition model training device. Optionally, Figure 10 shows a hardware structure block diagram of the OCR recognition model training device. Referring to Figure 10 , the hardware structure of the OCR recognition model training device may include: at least one processor 1001, at least one communication interface 1002, at least one memory 1003, and at least one communication bus 1004;

[0153] In the embodiments of the present application, the number of the processor 1001, the communication interface 1002, the memory 1003, and the communication bus 1004 is at least one, and the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004;

[0154] The processor 1001 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;

[0155] The memory 1003 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;

[0156] Among them, the memory 1003 stores a program, and the processor 1001 can call the program stored in the memory 1003. The program is used for:

[0157] Cut a first image sample of unlabeled data into multiple first image blocks, and randomly select some first image blocks from the multiple first image blocks for occlusion to obtain occluded image blocks and unoccluded image blocks, where the number of occluded image blocks is greater than the number of unoccluded image blocks; use the occluded image blocks and unoccluded image blocks to pre-train an initial feature recognition model including an encoder and a first decoder constructed in advance with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model; construct a task processing model based on the encoder and the second decoder in the pre-trained feature recognition model;

[0158] Cut a second image sample of labeled data into multiple second image blocks;

[0159] Train the task processing model with multiple second image blocks and the word piece sequence included in the second image sample, and use the obtained model as the OCR recognition model.

[0160] Optionally, the refined functions and extended functions of the program can be referred to the above description.

[0161] An embodiment of the present application also provides an OCR recognition device. Optionally, Figure 11 shows a hardware structure block diagram of the OCR recognition device. Referring to Figure 11 , the hardware structure of the OCR recognition device may include: at least one processor 1101, at least one communication interface 1102, at least one memory 1103, and at least one communication bus 1104;

[0162] In the embodiment of the present application, the number of the processor 1101, the communication interface 1102, the memory 1103, and the communication bus 1104 is at least one, and the processor 1101, the communication interface 1102, and the memory 1103 complete mutual communication through the communication bus 1104;

[0163] The processor 1101 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;

[0164] The memory 1103 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;

[0165] Among them, the memory 1103 stores a program, and the processor 1101 can call the program stored in the memory 1103. The program is used for:

[0166] Obtain an image to be recognized;

[0167] Cut the image to be recognized into multiple image blocks to be recognized;

[0168] Input the multiple image blocks to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

[0169] Optionally, the refined functions and extended functions of the program can be referred to the above description.

[0170] An embodiment of the present application also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each processing flow in the above OCR recognition model training method, or implements each processing flow in the above OCR recognition method.

[0171] Optionally, the refined functions and extended functions of the program can be referred to the above description.

[0172] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0173] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other.

[0174] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training an OCR recognition model, characterized in that, Including: Segmenting a first image sample of unlabeled data into multiple first image patches, randomly selecting some of the first image patches from the multiple first image patches for occlusion to obtain occluded image patches and unoccluded image patches, wherein the number of the occluded image patches is greater than the number of the unoccluded image patches; using the occluded image patches and the unoccluded image patches, aiming at reconstructing the preset features of the first image sample, pre-training an initially constructed feature recognition model including an encoder and a first decoder to obtain a pre-trained feature recognition model; constructing a task processing model based on the encoder and a second decoder in the pre-trained feature recognition model; Segmenting a second image sample of labeled data into multiple second image patches; Training the task processing model with the multiple second image patches and the word piece sequence included in the second image sample, and using the obtained model as an OCR recognition model; The step of using the occluded image patches and the unoccluded image patches, aiming at reconstructing the preset features of the first image sample, pre-training an initially constructed feature recognition model including an encoder and a first decoder to obtain a pre-trained feature recognition model includes: inputting the unoccluded image patches into the encoder included in the initially constructed feature recognition model to obtain the visual semantic information of the unoccluded image patches; inputting the visual semantic information of the unoccluded image patches and the occluded image patches into the first decoder included in the initially constructed feature recognition model to obtain the preset features respectively corresponding to the reconstructed unoccluded image patches and the occluded image patches; pre-training the parameters of the initially constructed feature recognition model according to the preset features respectively corresponding to the reconstructed unoccluded image patches and the occluded image patches, and the preset features extracted from the first image sample to obtain the pre-trained feature recognition model.

2. The OCR recognition model training method according to claim 1, wherein The preset features include one or more of the following features: Histogram of Oriented Gradients (HOG) feature, Convolutional Neural Network (CNN) feature, Local Binary Pattern (LBP) feature, Haar-like feature, Oriented FAST and Rotated BRIEF (ORB) feature, Speeded Up Robust Features (SURF) and Scale-Invariant Feature Transform (SIFT) feature.

3. The OCR recognition model training method according to claim 1, wherein The step of training the task processing model with the multiple second image patches and the word piece sequence included in the second image sample, and using the obtained model as an OCR recognition model includes: Inputting the multiple second image patches into the task processing model to obtain the OCR recognition results corresponding to the multiple second image patches; Training the parameters of the task processing model based on the OCR recognition results corresponding to the multiple second image patches and the word piece sequence included in the second image sample to obtain the OCR recognition model.

4. The OCR recognition model training method according to claim 1, wherein A method for obtaining an image sample, including: Obtaining an original sample; Performing image conversion processing on the original sample to obtain a processed original sample, wherein the image conversion processing includes one or more of the following processes: random rotation, Gaussian blur, image dilation, image erosion, downsampling, and adding an underline; Using the original sample and the processed original sample as the image sample.

5. An OCR recognition method, characterized in that, Applied to the OCR recognition model according to any one of claims 1 to 4, including: Obtain the image to be recognized; Segment the image to be recognized into a plurality of image blocks to be recognized; Input the plurality of image blocks to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

6. An OCR recognition model training device, characterized in that Including: A first segmentation module, a first training module, a model construction module, a second segmentation module, and a second training module; The first segmentation module is configured to segment a first image sample of unannotated data into a plurality of first image blocks, and randomly select some of the first image blocks from the plurality of first image blocks for occlusion to obtain occluded image blocks and unoccluded image blocks, wherein the number of the occluded image blocks is greater than the number of the unoccluded image blocks; The first training module is configured to use the occluded image blocks and the unoccluded image blocks to pre-train an initial feature recognition model including an encoder and a first decoder with the goal of reconstructing the preset features of the first image sample to obtain a pre-trained feature recognition model; the model construction module is configured to construct a task processing model based on the encoder and the second decoder in the pre-trained feature recognition model; The second segmentation module is configured to segment a second image sample of annotated data into a plurality of second image blocks; The second training module is configured to train the task processing model with the plurality of second image blocks and the token sequence included in the second image sample, and use the obtained model as the OCR recognition model; Specifically, the first training module is configured to: input the unoccluded image blocks into the encoder included in the initial feature recognition model to obtain the visual semantic information of the unoccluded image blocks; input the visual semantic information of the unoccluded image blocks and the occluded image blocks into the first decoder included in the initial feature recognition model to obtain the preset features corresponding to the reconstructed unoccluded image blocks and the occluded image blocks respectively; pre-train the parameters of the initial feature recognition model according to the preset features corresponding to the reconstructed unoccluded image blocks and the occluded image blocks respectively, and the preset features extracted from the first image sample to obtain the pre-trained feature recognition model.

7. An OCR recognition device, characterized in that, Applied to the OCR recognition model according to claim 6, including: an acquisition module, a third segmentation module, and a recognition module; The acquisition module is configured to obtain the image to be recognized; The third segmentation module is configured to segment the image to be recognized into a plurality of image blocks to be recognized; The recognition module is configured to input the plurality of image blocks to be recognized into the OCR recognition model to obtain the OCR recognition result output by the OCR recognition model as the OCR recognition result of the image to be recognized.

8. A data processing device, characterized in that, Including a memory and a processor; The memory is configured to store a program; The processor is configured to execute the program to implement the steps of the OCR recognition model training method according to any one of claims 1 to 4, or the OCR recognition method according to claim 5.

9. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the various steps of the OCR recognition model training method according to any one of claims 1 to 4, or the OCR recognition method according to claim 5.