Character Recognition Model Training Method, Device, and Equipment
By automatically labeling unlabeled images using prediction results during character recognition model training, the problem of low manual labeling efficiency in the prior art is solved, and the model training efficiency and recognition performance are improved.
Patent Information
- Application Number
- CN201910645222.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-07-17
AI Technical Summary
The existing character recognition model training method requires a large number of manual image labels, resulting in inefficient labeling, which in turn affects the model training efficiency.
By selecting unlabeled images to input the trained character recognition model, the image is automatically labeled using the model's prediction results, and the image is trained while labeling during the training process, increasing the number of marked images, and improving the sample size to improve the model performance.
It reduces the workload and cost of image labeling, improves the labeling efficiency, thereby improving the model training efficiency and ensuring the recognition performance of the character recognition model.
Smart Images

Figure CN112241749B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and particularly relates to a method, device and equipment for training a character recognition model. Background Art
[0002] With the development of science and technology, deep learning algorithms have performed excellently in tasks such as classification, detection, and recognition. In character recognition technology, an image is input into a trained character recognition model to recognize the characters in the image through the character recognition model. The prerequisite for implementing this technology is to train a character recognition model using a large number of samples.
[0003] In the existing character recognition model training method, after collecting a large number of images, it is necessary to manually label the tags of each character in the images one by one, and then train the required character recognition model using all the images with labeled tags.
[0004] In the above method, since a large number of samples are required for model training, and there may be a large number of characters in each image, a large number of tags need to be manually labeled, and the labeling efficiency is too low, resulting in low model training efficiency. Summary of the Invention
[0005] In view of this, the present invention provides a method, device and equipment for training a character recognition model, which can improve the labeling efficiency and thus improve the model training efficiency.
[0006] The first aspect of the present invention provides a method for training a character recognition model, including:
[0007] Selecting unlabeled images from an image sample set; the image sample set includes labeled images and unlabeled images;
[0008] Inputting the selected unlabeled images into a character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set;
[0009] For each unlabeled image input into the character recognition model, labeling the unlabeled image in the image sample set according to the predicted character recognition result of the unlabeled image to obtain a labeled image;
[0010] Training a target character recognition model according to the labeled images in the image sample set and the character recognition model.
[0011] According to an embodiment of the present invention, selecting unlabeled images from the image sample set includes:
[0012] If the number of unlabeled images in the image sample set is greater than or equal to the set number, select the set number of unlabeled images from the image sample set;
[0013] If the number of unlabeled images in the image sample set is less than the set number, select all the remaining unlabeled images from the image sample set.
[0014] According to an embodiment of the present invention,
[0015] The predicted character recognition result includes predicted character information of each character in the unlabeled image;
[0016] Labeling the unlabeled image in the image sample set according to the predicted character recognition result of the unlabeled image includes:
[0017] Determine the predicted character labels of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image;
[0018] For each character in the unlabeled image, determine whether the correct character label of the character is predicted. If so, determine the character label as the target label of the character. If not, re-determine a character label as the target label of the character;
[0019] Label the unlabeled image in the image sample set according to the target labels of each character in the unlabeled image to obtain a labeled image.
[0020] According to an embodiment of the present invention, determining whether the correct character label of the character is predicted includes:
[0021] Receive an externally input instruction; the instruction carries indication information that the correct character label of the character is not predicted;
[0022] If the indication information carried by the instruction indicates that the correct character label of the character is not predicted, determine that the correct character label of the character is not predicted. If the indication information carried by the instruction does not indicate that the correct character label of the character is not predicted, determine that the correct character label of the character is predicted.
[0023] According to an embodiment of the present invention,
[0024] The instruction further carries: candidate labels;
[0025] The re-determining a character label as the target label of the character includes:
[0026] Select a candidate label from the candidate labels carried by the instruction and determine the selected candidate label as the target label of the character.
[0027] According to an embodiment of the present invention, training a target character recognition model based on the labeled images in the image sample set and the character recognition model includes:
[0028] Training the character recognition model based on the labeled images in the image sample set;
[0029] Checking whether the set training end condition is currently satisfied. If not, returning the operation of selecting unlabeled images from the image sample set. If so, determining the character recognition model as the target character recognition model.
[0030] According to an embodiment of the present invention, before selecting unlabeled images from the image sample set, it further includes:
[0031] Obtaining images collected from a specified scene, where the style of the characters in the specified scene is a specified style;
[0032] Cropping out character regions from each of the collected images, where each character region contains at least one character;
[0033] Determining the image sample set based on all the cropped character regions.
[0034] A second aspect of the present invention provides a character recognition model training device, including:
[0035] A selection module for selecting unlabeled images from the image sample set; the image sample set includes labeled images and unlabeled images;
[0036] A prediction module for inputting the selected unlabeled images into the character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained based on the labeled images in the image sample set;
[0037] A labeling module for labeling each unlabeled image input into the character recognition model according to the predicted character recognition result of the unlabeled image in the image sample set to obtain labeled images;
[0038] A training module for training a target character recognition model based on the labeled images in the image sample set and the character recognition model.
[0039] According to an embodiment of the present invention, when the selection module selects unlabeled images from the image sample set, it specifically is used for:
[0040] If the number of unlabeled images in the image sample set is greater than or equal to the set number, selecting the set number of unlabeled images from the image sample set;
[0041] If the number of unlabeled images in the image sample set is less than the set number, select all the remaining unlabeled images from the image sample set.
[0042] According to an embodiment of the present invention,
[0043] The predicted character recognition result includes predicted character information of each character in the unlabeled image;
[0044] When the labeling module labels the unlabeled image in the image sample set according to the predicted character recognition result of the unlabeled image, it is specifically used for:
[0045] Determine the predicted character labels of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image;
[0046] For each character in the unlabeled image, determine whether the correct character label is predicted for the character. If so, determine the character label as the target label of the character. If not, re - determine a character label as the target label of the character;
[0047] Label the unlabeled image in the image sample set according to the target labels of each character in the unlabeled image to obtain a labeled image.
[0048] According to an embodiment of the present invention,
[0049] When the labeling module determines whether the correct character label is predicted for the character, it is specifically used for:
[0050] Receive an externally input instruction; the instruction carries indication information that the correct character label is not predicted for the character;
[0051] If the indication information carried by the instruction indicates that the correct character label is not predicted for the character, determine that the correct character label is not predicted for the character. If the indication information carried by the instruction does not indicate that the correct character label is not predicted for the character, determine that the correct character label is predicted for the character.
[0052] According to an embodiment of the present invention,
[0053] The instruction further carries: candidate labels;
[0054] When the labeling module re - determines a character label as the target label of the character, it is specifically used for:
[0055] Select a candidate label from the candidate labels carried by the instruction and determine the selected candidate label as the target label of the character.
[0056] According to an embodiment of the present invention, when the training module trains a target character recognition model based on the labeled images in the image sample set and the character recognition model, it specifically is used for:
[0057] Training the character recognition model based on the labeled images in the image sample set;
[0058] Checking whether the set training end condition is currently met. If not, returning to the operation of the selection module. If so, determining the character recognition model as the target character recognition model.
[0059] According to an embodiment of the present invention, the device further includes:
[0060] An image acquisition module, configured to acquire images collected from a specified scene, where the style of characters in the specified scene is a specified style;
[0061] A region intercepting module, configured to intercept a character region from each of the collected images, where the character region contains at least one character;
[0062] An image sample set determination module, configured to determine the image sample set according to all the intercepted character regions.
[0063] A third aspect of the present invention provides an electronic device, including a processor and a memory; the memory stores a program that can be called by the processor; wherein, when the processor executes the program, it implements the character recognition model training method as described in the foregoing embodiments.
[0064] The embodiments of the present invention have the following beneficial effects:
[0065] In the embodiments of the present invention, a small number of images can be pre-labeled as the labeled images in the image sample set, and a character recognition model can be trained using the labeled images in the image sample set. A large number of unlabeled images can be used as the unlabeled images in the image sample set. When entering the model training process, unlabeled images are selected from the image sample set, and the predicted character recognition results of the selected unlabeled images are determined using the character recognition model. Accordingly, the corresponding unlabeled images in the image sample set are labeled to obtain labeled images, increasing the number of labeled images in the image sample set. Based on a sufficient sample size, a target character recognition model can be trained according to the labeled images in the image sample set and the character recognition model. During the above training process, the unlabeled images in the image sample set can be labeled while training, greatly reducing the workload and cost of image annotation, improving the efficiency of annotation, and further improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flowchart of the character recognition model training method according to an embodiment of the present invention;
[0067] Figure 2 is a structural block diagram of a character recognition model training device according to an embodiment of the present invention;
[0068] Figure 3 is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0069] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0070] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0071] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various devices, the information should not be limited to these terms. These terms are only used to distinguish devices of the same type from each other. For example, without departing from the scope of the present invention, the first device may also be referred to as the second device, and similarly, the second device may also be referred to as the first device. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0072] In order to make the description of the present invention clearer and more concise, some technical terms in the present invention are explained below:
[0073] Neural network: A technology abstracted by imitating the structure of the brain. This technology connects a large number of simple functions in a complex manner to form a network system. This system can fit extremely complex functional relationships and generally can include convolution / transposed convolution operations, activation operations, pooling operations, as well as operations such as addition, subtraction, multiplication, division, channel merging, and element rearrangement. Using specific input data and output data to train the network and adjusting the connections therein can enable the neural network to learn and fit the mapping relationship between the input and the output.
[0074] The following describes the character recognition model training method of the embodiments of the present invention in more detail, but should not be limited thereto. Refer toFigure 1 , in one embodiment, the method for training a character recognition model includes the following steps:
[0075] S100: Select unlabeled images from the image sample set; the image sample set includes labeled images and unlabeled images;
[0076] S200: Input the selected unlabeled images into the character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set;
[0077] S300: For each unlabeled image input into the character recognition model, label the unlabeled image in the image sample set according to the predicted character recognition result of the unlabeled image to obtain labeled images;
[0078] S400: Train a target character recognition model according to the labeled images in the image sample set and the character recognition model.
[0079] The execution subject of the method for training a character recognition model according to an embodiment of the present invention may be an electronic device, and further may be a processor of the electronic device. The electronic device may be a device such as a computer or a mobile terminal, and the specific type is not limited as long as it has a certain data processing ability.
[0080] Preferably, human-computer interaction can be realized on the electronic device. For example, it may have an instruction input device and an information output device. The instruction input device can receive instructions input by the user to perform corresponding operations according to the instructions indicated by the external input instructions. The information output device can present the situation of the electronic device performing operations to the user.
[0081] In step S100, unlabeled images are selected from the image sample set; the image sample set includes labeled images and unlabeled images.
[0082] The image sample set is a set composed of labeled images and unlabeled images. In other words, some images in the image sample set are labeled and some images are unlabeled. Since model training usually requires a large number of samples, the image sample set may include multiple labeled images and multiple unlabeled images, and the specific quantity is not limited.
[0083] Preferably, when step S100 is first executed, the number of labeled images in the image sample set may be less than the number of unlabeled images. In this way, most of the labeling work can be automatically completed by the electronic device.
[0084] The labeled image is an image that has been labeled and can be directly used to train the model. The labeled image can carry at least one character label, and each character label is used to describe the character information of the corresponding character in the labeled image. The character information can include character content, character position information, etc. The labeled image can be obtained by manual annotation in advance, and the specific annotation method is not limited.
[0085] Each image in the image sample set can have one or more than two characters, and the categories of characters can include Chinese characters, letters, numbers, special symbols, etc. Each character in the labeled image can be labeled with a character label, and of course, it can also be labeled with a character label for one or several specific categories of characters (such as Chinese characters), and the specific situation is not limited.
[0086] The unlabeled image is an image that has not been labeled and needs to be labeled before being used to train the model. There can be multiple unlabeled images in the image sample set. When selecting unlabeled images from the image sample set, at least one can be selected, and the specific number selected and the selection method are not limited.
[0087] In step S200, the selected unlabeled images are input into the character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set.
[0088] The character recognition model is trained according to the labeled images in the image sample set. Although the character recognition model already has a certain character recognition ability, due to the insufficient sample size for training the character recognition model, the recognition performance of the character recognition model has not yet reached the set requirements, that is, it is not yet the target character recognition model required ultimately.
[0089] The character recognition model can be trained in the following way in advance: input the labeled images in the image sample set into the initial model, so that the initial model can perform character recognition on the input labeled images, compare the recognition results of the characters with the character labels in the input labeled images, and optimize the initial model according to the comparison results. The optimized initial model is used as the character recognition model. It can be understood that the training method here is only an example, and the specific situation is not limited to this.
[0090] The character recognition model can be specifically built by algorithm frameworks such as YOLO, Faster-RCNN, etc. Of course, the specific framework of the character recognition model is not limited, as long as it can implement character recognition after training.
[0091] The character recognition model has the function of character recognition. The selected unlabeled images are input into the character recognition model, and the character recognition model performs character recognition on each input unlabeled image to obtain the predicted character recognition result. The predicted character recognition result may include the predicted character position information, character content, etc. in the unlabeled image. Of course, the predicted character recognition result may also include other information.
[0092] Specifically, the character recognition model can perform character localization on each input unlabeled image to obtain the character position information in the unlabeled image, and recognize the character corresponding to the character position information in the unlabeled image to obtain the character content. Among them, the character position information can represent the position of the character in the image, and specifically can be the position information of the character detection box.
[0093] In step S300, for each unlabeled image input into the character recognition model, the unlabeled image in the image sample set is labeled according to the predicted character recognition result of the unlabeled image to obtain a labeled image.
[0094] For each unlabeled image input into the character recognition model, there is a corresponding predicted character recognition result. The predicted character recognition result can be used to determine the information required for labeling the unlabeled image, so as to complete the labeling of the unlabeled image. For example, the character position information and character content in the predicted character recognition result can be used as the character label of the unlabeled image. The specific labeling method is not limited, as long as the image can be labeled.
[0095] The unlabeled image in the image sample set is labeled to obtain a labeled image. In other words, the unlabeled image in the image sample set becomes a labeled image. For example, there are two unlabeled images A1 and A2 in the image sample set that are selected and input into the character recognition model. A1 in the image sample set is labeled as B1 according to the predicted recognition result of A1, and A2 in the image sample set is labeled as B2 according to the predicted recognition result of A2. Thus, there are two more labeled images B1 and B2 in the image sample set.
[0096] Therefore, after step S300, the number of labeled images in the image sample set increases compared with that before executing step S300. In other words, the sample size for training increases.
[0097] In step S400, a target character recognition model is trained based on the labeled images in the image sample set. The character recognition model can be trained according to the labeled images in the image sample set; the trained character recognition model can be directly used as the target character recognition model, or the operation of selecting unlabeled images from the image sample set can be returned until the target character recognition model is trained. Of course, there is no limit to the specific manner of training the target character recognition model based on the labeled images in the image sample set and the character recognition model.
[0098] Since the number of labeled images in the image sample set has increased and the amount of samples for training has increased, the recognition performance of the trained model is better, and finally a target character recognition model with performance meeting the requirements is obtained.
[0099] In an embodiment of the present invention, a small number of images can be pre-labeled as the labeled images in the image sample set, and a character recognition model can be trained using the labeled images in the image sample set. A large number of unlabeled images can be used as the unlabeled images in the image sample set. When entering the model training process, unlabeled images are selected from the image sample set, and the predicted character recognition results of the selected unlabeled images are determined using the character recognition model. Accordingly, the corresponding unlabeled images in the image sample set are labeled to obtain labeled images, increasing the number of labeled images in the image sample set. Based on a sufficient amount of samples, a target character recognition model can be trained according to the labeled images in the image sample set and the character recognition model. During the above training process, the unlabeled images in the image sample set can be labeled while training, greatly reducing the workload and cost of image labeling, improving the efficiency of labeling, and thus improving the efficiency of model training.
[0100] In one embodiment, the above method flow can be executed by a character recognition model training device. As Figure 2 shown, the character recognition model training device 100 can include 4 modules: a selection module 101, a prediction module 102, a labeling module 103, and a training module 104. The selection module 101 is used to execute the above step S100, the prediction module 102 is used to execute the above step S200, the labeling module 103 is used to execute the above step S300, and the training module 104 is used to execute the above step S4400.
[0101] In one embodiment, in step S400, training the target character recognition model based on the labeled images in the image sample set and the character recognition model may include the following steps:
[0102] S401: Train the character recognition model according to the labeled images in the image sample set;
[0103] S402: Check whether the set training end condition is currently met. If not, return the operation of selecting unlabeled images from the image sample set. If so, determine the character recognition model as the target character recognition model.
[0104] In step S401, training the character recognition model based on the labeled images in the image sample set can be achieved in the following way: Input the labeled images in the image sample set into the character recognition model so that the character recognition model can perform character recognition on the input labeled images. Compare the recognition results of the characters with the character labels in the input labeled images, and optimize the character recognition model according to the comparison results.
[0105] In step S402, check whether the set training end condition is currently met. If not, return the operation of selecting unlabeled images from the image sample set for a new round of iteration. Through multiple iterations, the number of labeled images in the pattern sample set can be continuously increased, that is, the number of samples for model training is gradually increased, and the recognition performance of the character recognition model is continuously improved until the current meets the set training end condition, and the character recognition model is determined as the target character recognition model.
[0106] There can be multiple training end conditions. For example, in one example, test the recognition performance of the character recognition model. If the recognition performance meets the set requirements, it means that the current has met the training end condition and the iteration can be ended; in another example, if there are no unlabeled images in the image sample set, it means that the current has met the training end condition and the iteration can be ended; in yet another example, an iteration count threshold can be set, calculate the current iteration count, and when the current iteration count reaches the iteration count threshold, it means that the current has met the training end condition and the iteration can be ended.
[0107] The above examples are not restrictive. As long as the current meets the set training end condition, the iteration can be ended without returning an operation.
[0108] Optionally, after multiple iterations, all characters with a relatively high prediction error rate in the labeled images marked during the iteration process can be found, and then the character recognition model can be further optimized using the images containing the characters with a relatively high prediction error rate, so that the character recognition model can reduce the recognition error rate of these characters and improve the recognition accuracy.
[0109] In this embodiment, when the training end condition is not currently met, a new round of iteration is returned. Through continuous iteration, labeled images are continuously added to the image sample set. Thus, more and more labeled images can be used to train the character recognition model, improving the recognition performance of the character recognition model. Correspondingly, using the character recognition model with continuously improved performance to perform character recognition prediction on a new round of unlabeled images can make the labeling of unlabeled images more accurate, further reducing the workload and cost of image labeling and ensuring the performance of the target character recognition model.
[0110] In one embodiment, in step S100, selecting unlabeled images from the image sample set includes the following steps:
[0111] S101: If the number of unlabeled images in the image sample set is greater than or equal to the set number, select the set number of unlabeled images from the image sample set;
[0112] S102: If the number of unlabeled images in the image sample set is less than the set number, select all the remaining unlabeled images from the image sample set.
[0113] During the iteration process, if the number of unlabeled images in the image sample set is greater than or equal to the set number, the set number of unlabeled images can be selected each time and input into the character recognition model for character recognition prediction. If the number of unlabeled images in the image sample set is less than the set number, then all the remaining unlabeled images in the image sample set can be selected and input into the character recognition model for character recognition prediction. The set number is not specifically limited. For example, it can be 5, 6, etc.
[0114] In this embodiment, not selecting all unlabeled images at once can avoid the problem of too high a probability of prediction errors caused by the character recognition model that does not meet the requirements yet performing recognition prediction on all unlabeled images at once. Even if the recognition performance of the character recognition model is relatively low in the first few iterations, the number of incorrect predicted character recognition results will not be excessive and can be corrected in a timely manner. As the recognition performance of the character recognition model gradually improves, the error rate of prediction can be gradually reduced, thereby reducing the workload of correction required in the whole process.
[0115] In one embodiment, the predicted character recognition result includes the predicted character information of each character in the unlabeled image;
[0116] In step S300, labeling the unlabeled image in the image sample set according to the predicted character recognition result of the unlabeled image includes:
[0117] S301: Determine the predicted character labels of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image;
[0118] S302: For each character in the unlabeled image, determine whether the correct character label has been predicted for the character. If so, determine this character label as the target label of the character. If not, re-determine a character label as the target label of the character;
[0119] S303: Label the unlabeled image in the image sample set according to the target labels of the characters in the unlabeled image to obtain a labeled image.
[0120] In step S301, according to the predicted character information in the predicted character recognition result of the unlabeled image, determine the character labels predicted for the characters in the unlabeled image. The predicted character information may include the predicted character position information and the character content. The character label predicted for the character can be used to describe the predicted character information of the character.
[0121] Since the character labels predicted for the characters in the unlabeled image are based on the predicted character recognition result of the unlabeled image, there may be inaccurate situations. Especially when the character recognition model is trained according to a small number of labeled images, the probability of inaccuracy is even greater.
[0122] The inaccurate situations are as follows: The character label predicted for a certain character in the image is incorrect. For example, the predicted character content is not the true content of the character, or the predicted character position information is not the true position information of the character in the image; and, a certain character in the image has not been predicted with a character label. The above two situations are both cases where the correct character label has not been predicted and need to be corrected. Of course, there may be other situations. For example, a character label is predicted for a characterless area in the image. In this case, the character label can be directly discarded.
[0123] Therefore, in step S302, for each character in the unlabeled image, determine whether the correct character label has been predicted for the character. If so, determine this character label as the target label of the character. If not, re-determine a character label as the target label of the character. In this way, ensure that the target labels of the characters in the unlabeled image are all correct.
[0124] To determine whether the correct character label has been predicted for each character, regions corresponding to the character position information in the character label predicted for the character can be intercepted from each image, and the intercepted regions are classified and displayed to the user according to the predicted character content, with the regions of the same character content being displayed each time. The user can judge whether the correct character label has been predicted for the character by viewing the displayed regions, and when the correct character label has not been predicted, input an instruction to the electronic device. The instruction indicates that the correct character label has not been predicted for the character and indicates a correct character label for the character.
[0125] For example, if a "6" appears among a bunch of "8"s shown, it means that the "6" is misrecognized as an "8", and the character content recognition is incorrect. The character content in the predicted character label can be modified from "8" to "6". Of course, this is just an example here. In fact, there will be other situations. As long as it is ensured that each character in the unlabeled image has a correct target label.
[0126] In step S303, based on the target labels of each character in the unlabeled image, the unlabeled image in the image sample set is labeled to obtain a labeled image. For example, directly mark the target labels of each character in the unlabeled image in the image sample set. Of course, the specific labeling method is not limited to this.
[0127] In one embodiment, in step S302, determining whether the correct character label is predicted for the character includes the following steps:
[0128] S3021: Receive an instruction input externally; the instruction carries indication information that the correct character label is not predicted for the character;
[0129] S3022: If the indication information carried by the instruction indicates that the correct character label is not predicted for the character, determine that the correct character label is not predicted for the character; if the indication information carried by the instruction does not indicate that the correct character label is predicted for the character, determine that the correct character label is predicted for the character.
[0130] In this embodiment, the instruction can be input by the user. The verification of the predicted character label is realized through the interaction between the user and the electronic device. The verification work only needs to operate on several characters for which the correct character label is not predicted, and there is no need to operate on all characters. Therefore, the verification work takes less time and effort than the labeling work.
[0131] The instruction carries indication information that the correct character label is not predicted for the character. Based on the indication information, it can be determined that the correct character label is not predicted for the character. There are two situations where the correct character label is not predicted for the character: one situation is that an incorrect character label is predicted, and the other situation is that no character label is predicted. In either case, the correct character label is not predicted for the character.
[0132] In other words, if the indication information carried by the instruction indicates that the correct character label is not predicted for the character, determine that the correct character label is not predicted for the character; if the indication information carried by the instruction does not indicate that the correct character label is not predicted for the character, determine that the correct character label is predicted for the character.
[0133] In one embodiment, the instruction further carries: a candidate label;
[0134] In step S302, re-determining a character label as the target label of the character includes:
[0135] Selecting a candidate label from the candidate labels carried by the instruction and determining the selected candidate label as the target label of the character.
[0136] The candidate label can be a correct candidate label determined by the user, which correctly describes the character information of the character. Of course, the instruction can also carry a lot of candidate labels at one time, and one of them is a correct candidate label of the character.
[0137] When selecting a candidate label from the candidate labels carried by the instruction, a candidate label that correctly describes the character information of the character can be selected, and the selected candidate label is determined as the target label of the character.
[0138] Of course, the instruction can also carry other information, such as selection indication information. The selection indication information can indicate the image where the character is located (which can directly jump to the image) and the content of the character in the image. Of course, this is only an example here and is not a limitation. In this way, a correct candidate label is selected for the character according to the selection indication information as the target character of the character.
[0139] In one embodiment, before selecting the unlabeled image from the image sample set, the following steps are further included:
[0140] S010: Obtaining images collected from a specified scenario, where the style of the characters in the specified scenario is a specified style;
[0141] S020: Cropping out the character regions from each of the collected images, where each character region contains at least one character;
[0142] S030: Determining the image sample set according to all the cropped character regions.
[0143] In step S010, obtaining images collected from a specified scenario, where the style of the characters in the specified scenario is a specified style. The specified scenario can be, for example, an industrial scenario. In an industrial scenario, there are usually some devices for manufacturing, control, etc. There are some characters on these devices, such as production date, operation date, device model, device status parameters, and other characters. Generally, the styles of these characters are specified. Of course, the specified scenario can also be other scenarios, as long as the style of the characters in the scenario is the specified style.
[0144] Since the characters in the specified scenario are in a specified style, using the images collected in the specified scenario to train a character recognition model can, compared to scenarios with diverse styles, train a character recognition model with high accuracy and robustness with a relatively small number of samples. This character recognition model can be used to recognize characters in the images collected from the specified scenario.
[0145] In step S020, a character region is cropped from each of the collected images. The character region contains at least one character, and the specific number of characters is not limited and can be determined according to the actual situation.
[0146] When cropping a character region from each of the collected images, it is necessary to first locate the character region. A pre-trained character detection model can be used to locate the character region. The collected image is input into the character detection model, and the character detection model locates the region where the character (which can be multiple characters with close positions) is located and outputs the region position information. The region corresponding to the region position information in the image is the character region. The above character detection model can be built with YOLO.
[0147] Of course, the method of locating the character region is not limited to this. For example, it can also be achieved by template matching and other methods.
[0148] In step S030, determining the image sample set according to all the cropped character regions can reduce the background of the images in the image sample set, thereby reducing the interference factors during model training. The model only needs to learn from the character regions, which is beneficial to improving the recognition performance of the model.
[0149] Specifically, from all the cropped character regions, a part of the character regions can be selected for annotation and added to the image sample set as labeled images, while the other part of the character regions are added to the image sample set as unlabeled images.
[0150] The present invention also provides a character recognition model training device. In one embodiment, see Figure 2 , the character recognition model training device 100 includes:
[0151] A selection module 101, configured to select unlabeled images from the image sample set; the image sample set includes labeled images and unlabeled images;
[0152] A prediction module 102, configured to input the selected unlabeled images into the character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set;
[0153] The annotation module 103 is configured to, for each unannotated image input to the character recognition model, annotate the unannotated image in the image sample set according to the predicted character recognition result of the unannotated image, so as to obtain an annotated image;
[0154] The training module 104 is configured to train a target character recognition model according to the annotated images in the image sample set and the character recognition model.
[0155] The character recognition model training device according to the embodiment of the present invention can be applied to an electronic device. The electronic device can be a device such as a computer or a mobile terminal, and the specific type is not limited as long as it has a certain data processing ability.
[0156] Preferably, human-computer interaction can be implemented on the electronic device. For example, it can have an instruction input device and an information output device. The instruction input device can receive instructions input by the user to perform corresponding operations according to the instructions indicated by the external input instructions. The information output device can present the situation of the electronic device performing operations to the user.
[0157] The selection module 101 is configured to select an unannotated image from the image sample set; the image sample set includes annotated images and unannotated images.
[0158] The image sample set is a set composed of annotated images and unannotated images. In other words, some images in the image sample set are annotated and some are not. Since model training usually requires a large number of samples, the image sample set can include multiple annotated images and multiple unannotated images, and the specific quantity is not limited.
[0159] Preferably, when the selection module 101 is executed for the first time, the number of annotated images in the image sample set can be less than the number of unannotated images. In this way, most of the annotation work can be automatically completed by the electronic device.
[0160] An annotated image is an image that has been annotated and can be directly used for training the model. An annotated image can carry at least one character label, and each character label is used to describe the character information of the corresponding character in the annotated image. The character information can include character content, character position information, etc. The annotated image can be obtained by manual annotation in advance, and the specific annotation method is not limited.
[0161] Each image in the image sample set can have one or more than two characters, and the categories of characters can include Chinese characters, letters, numbers, special symbols, etc. Each character in the annotated image can be labeled with a character label, or of course, it can also be labeled with a character label of one or several specific categories of characters (such as Chinese characters), and the specific situation is not limited.
[0162] An unlabeled image is an image that has not been labeled and needs to be labeled before being used to train the model. There can be multiple unlabeled images in the image sample set. When selecting unlabeled images from the image sample set, at least one can be selected, and the specific number and selection method are not limited.
[0163] The prediction module 102 is used to input the selected unlabeled images into the character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set.
[0164] The character recognition model is trained according to the labeled images in the image sample set. Although the character recognition model already has a certain character recognition ability, due to the insufficient sample size for training the character recognition model, the recognition performance of the character recognition model has not yet reached the set requirements, that is, it is not yet the final target character recognition model we need.
[0165] The character recognition model can be pre-trained in the following way: input the labeled images in the image sample set into the initial model so that the initial model can perform character recognition on the input labeled images, compare the recognition results of the characters with the character labels in the input labeled images, and optimize the initial model according to the comparison results. The optimized initial model is used as the character recognition model. It can be understood that the training method here is only an example and is not limited specifically.
[0166] Specifically, the character recognition model can be built by algorithm frameworks such as YOLO and Faster-RCNN. Of course, the specific framework of the character recognition model is not limited, as long as it can perform character recognition after training.
[0167] The character recognition model has the function of character recognition. Input the selected unlabeled images into the character recognition model so that the character recognition model can perform character recognition on each input unlabeled image to obtain the predicted character recognition results. The predicted character recognition results can include the predicted character position information, character content, etc. in the unlabeled image. Of course, the predicted character recognition results can also include other information.
[0168] Specifically, the character recognition model can perform character localization on each input unlabeled image to obtain the character position information in the unlabeled image, and recognize the character corresponding to the character position information in the unlabeled image to obtain the character content. Among them, the character position information can represent the position of the character in the image, specifically, it can be the position information of the character detection box.
[0169] The annotation module 103 is used to annotate each unannotated image input to the character recognition model according to the predicted character recognition result of the unannotated image, so as to obtain an annotated image.
[0170] For each unannotated image input to the character recognition model, there is a corresponding predicted character recognition result. The predicted character recognition result can be used to determine the information required for annotating the unannotated image, so as to complete the annotation of the unannotated image. For example, the character position information and character content in the predicted character recognition result can be used as the character label of the unannotated image. The specific annotation method is not limited, as long as the annotation of the image can be completed.
[0171] Annotate the unannotated image in the image sample set to obtain an annotated image. In other words, the unannotated image in the image sample set becomes an annotated image. For example, there are two unannotated images A1 and A2 selected and input to the character recognition model in the image sample set. A1 in the image sample set is annotated as B1 according to the predicted recognition result of A1, and A2 in the image sample set is annotated as B2 according to the predicted recognition result of A2. Thus, there are two more annotated images B1 and B2 in the image sample set.
[0172] Therefore, after the execution of the annotation module 103, the number of annotated images in the image sample set increases compared with that before the execution of the annotation module 103. In other words, the sample size for training increases.
[0173] The training module 104 is used to train a target character recognition model according to the annotated images in the image sample set and the character recognition model.
[0174] The character recognition model can be trained according to the annotated images in the image sample set; the trained character recognition model can be directly used as the target character recognition model, or the operation of selecting unannotated images from the image sample set can be returned until the target character recognition model is trained. Of course, the specific method of training the target character recognition model according to the annotated images in the image sample set and the character recognition model is not limited.
[0175] Since the number of annotated images in the image sample set increases and the sample size for training increases, the recognition performance of the trained model is better, and finally a target character recognition model with performance meeting the requirements is obtained.
[0176] In an embodiment of the present invention, a small number of images can be pre-annotated as the annotated images in the image sample set, and a character recognition model can be trained using the annotated images in the image sample set. A large number of unannotated images can be used as the unannotated images in the image sample set. When entering the model training process, unannotated images are selected from the image sample set, and the predicted character recognition results of the selected unannotated images are determined using the character recognition model. Based on this, the corresponding unannotated images in the image sample set are annotated to obtain annotated images, increasing the number of annotated images in the image sample set. Based on a sufficient sample size, a target character recognition model can be trained using the annotated images in the image sample set and the character recognition model. During the above training process, the annotation of unannotated images in the image sample set can be performed while training, greatly reducing the workload and cost of image annotation, improving the annotation efficiency, and thus improving the model training efficiency.
[0177] In one embodiment, when the training module trains a target character recognition model according to the annotated images in the image sample set and the character recognition model, it specifically is used for:
[0178] Training the character recognition model according to the annotated images in the image sample set;
[0179] Checking whether the set training end condition is currently satisfied. If not, returning to the operation of the selection module. If so, determining the character recognition model as the target character recognition model.
[0180] When the training module trains the character recognition model according to the annotated images in the image sample set, it can be implemented in the following manner: Inputting the annotated images in the image sample set into the character recognition model, so that the character recognition model can perform character recognition on the input annotated images, comparing the recognition results of the characters with the character labels in the input annotated images, and optimizing the character recognition model according to the comparison results.
[0181] The training module checks whether the set training end condition is currently satisfied. If not, returning to the operation of the selection module to select unannotated images from the image sample set for a new round of iteration. Through multiple iterations, the number of annotated images in the pattern sample set can be continuously increased, that is, the sample quantity for model training is continuously and gradually increased, and the recognition performance of the character recognition model is also continuously improved until the current set training end condition is satisfied, and the character recognition model is determined as the target character recognition model.
[0182] There can be multiple training end conditions. For example, in one case, when testing the recognition performance of a character recognition model, if the recognition performance meets the set requirements, it indicates that the current training end condition is satisfied and the iteration can be terminated; in another case, if there are no unlabeled images left in the image sample set, it means that the current training end condition is met and the iteration can end; in yet another case, an iteration count threshold can be set, the current iteration count is calculated, and when the current iteration count reaches the iteration count threshold, it shows that the current training end condition is satisfied and the iteration can be ended.
[0183] The above examples are not restrictive. As long as the current situation meets the set training end condition, the iteration can be ended without the need to return for an operation.
[0184] Optionally, after multiple iterations, the training module can find all the characters with a relatively high prediction error rate among the labeled images marked during the iteration process, and then further optimize the character recognition model using the images containing the characters with a relatively high prediction error rate, so that the character recognition model can reduce the recognition error rate of these characters and improve the recognition accuracy.
[0185] In this embodiment, when the current training end condition is not met, a new round of iteration is returned. Through continuous iteration, more and more labeled images are added to the image sample set, so that the character recognition model can be trained using an increasing number of labeled images, improving the recognition performance of the character recognition model. Correspondingly, using the character recognition model with continuously improved performance to perform character recognition prediction on a new round of unlabeled images can make the labeling of unlabeled images more accurate, further reducing the workload and cost of image labeling and ensuring the performance of the target character recognition model.
[0186] In one embodiment, when the selection module selects unlabeled images from the image sample set, it specifically is used for:
[0187] If the number of unlabeled images in the image sample set is greater than or equal to the set number, select a set number of unlabeled images from the image sample set;
[0188] If the number of unlabeled images in the image sample set is less than the set number, select all the remaining unlabeled images from the image sample set.
[0189] During the iteration process, if the number of unlabeled images in the image sample set is greater than or equal to the set number, a set number of unlabeled images can be selected each time and input into the character recognition model for character recognition prediction. If the number of unlabeled images in the image sample set is less than the set number, then all the remaining unlabeled images in the image sample set can be selected and input into the character recognition model for character recognition prediction. The specific set number is not limited, for example, it can be 5, 6, etc.
[0190] In this embodiment, not selecting all unlabeled images at once can avoid the problem of too high a probability of prediction errors caused by the character recognition model that does not meet the requirements yet performing recognition prediction on all unlabeled images at once. Even if the recognition performance of the character recognition model is low in the first few iterations, the number of incorrect predicted character recognition results will not be too large and can be corrected in time. As the recognition performance of the character recognition model gradually improves, the error rate of prediction can be gradually reduced, thereby reducing the workload of correction required in the whole process.
[0191] In one embodiment,
[0192] The predicted character recognition result includes the predicted character information of each character in the unlabeled image;
[0193] When the labeling module labels the unlabeled image in the image sample set according to the predicted character recognition result of the unlabeled image, it is specifically used for:
[0194] Determine the predicted character labels of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image;
[0195] For each character in the unlabeled image, determine whether the correct character label is predicted for the character. If so, determine the character label as the target label of the character. If not, re-determine a character label as the target label of the character;
[0196] Label the unlabeled image in the image sample set according to the target labels of each character in the unlabeled image to obtain a labeled image.
[0197] The labeling module determines the predicted character labels of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image. The predicted character information may include the predicted character position information and the character content. The predicted character label of the character can be used to describe the predicted character information of the character.
[0198] Since the predicted character labels of the characters in the unlabeled image are based on the predicted character recognition results of the unlabeled image, there may be inaccurate situations. Especially when the character recognition model is trained based on a small number of labeled images, the probability of inaccuracy is even greater.
[0199] The inaccurate situations are as follows: the predicted character label of a certain character in the image is incorrect, for example, the predicted character content is not the true content of the character, or the predicted character position information is not the true position information of the character in the image; and, a certain character in the image is not predicted with a character label. The above two situations are both cases where the correct character label is not predicted and need to be corrected. Of course, there may be other situations, for example, a character label is predicted for a characterless area in the image, and in this case, the character label can be directly discarded.
[0200] Therefore, for each character in the unlabeled image, the labeling module determines whether the character is predicted with the correct character label. If so, it determines the character label as the target label of the character. If not, it re-determines a character label as the target label of the character. In this way, it is ensured that the target labels of the characters in the unlabeled image are all correct.
[0201] To determine whether each character is predicted with the correct character label, regions corresponding to the character position information in the predicted character label of the character can be intercepted from each image, and the intercepted regions are classified and displayed to the user according to the predicted character content, and the regions with the same character content are displayed each time. The user can judge whether the character is predicted with the correct character label by viewing the displayed regions, and when the correct character label is not predicted, input an instruction to the electronic device, and the instruction indicates that the character is not predicted with the correct character label and indicates a correct character label for the character.
[0202] For example, if a "6" appears among a bunch of "8"s shown, it means that "6" is recognized as "8", and the character content recognition is incorrect. The character content in the predicted character label can be modified from "8" to "6". Of course, this is only an example here. In fact, there will be other situations as long as it is ensured that each character in the unlabeled image has a correct target label.
[0203] The labeling module labels the unlabeled image in the image sample set according to the target labels of the characters in the unlabeled image to obtain a labeled image. For example, directly mark the target labels of the characters in the unlabeled image on the unlabeled image in the image sample set. Of course, the specific labeling method is not limited to this.
[0204] In one embodiment,
[0205] When the annotation module determines whether the correct character label is predicted for the character, it is specifically used for:
[0206] Receiving an externally input instruction; the instruction carries indication information that the correct character label has not been predicted for the character;
[0207] If the indication information carried by the instruction indicates that the correct character label has not been predicted for the character, then it is determined that the correct character label has not been predicted for the character; if the indication information carried by the instruction does not indicate that the correct character label has not been predicted for the character, then it is determined that the correct character label has been predicted for the character.
[0208] In this embodiment, the instruction can be input by the user, and the verification of the predicted character label is realized through the interaction between the user and the electronic device. The verification work only needs to be performed on several characters for which the correct character label has not been predicted, and there is no need to operate on all characters. Therefore, the verification work requires less time and effort compared to the annotation work.
[0209] The instruction carries indication information that the correct character label has not been predicted for the character, and based on the indication information, it can be determined that the correct character label has not been predicted for the character. There are two situations where the correct character label has not been predicted for the character: one situation is that an incorrect character label is predicted, and the other situation is that no character label is predicted. In either case, it means that the correct character label has not been predicted for the character.
[0210] In other words, if the indication information carried by the instruction indicates that the correct character label has not been predicted for the character, then it is determined that the correct character label has not been predicted for the character; if the indication information carried by the instruction does not indicate that the correct character label has not been predicted for the character, then it is determined that the correct character label has been predicted for the character.
[0211] In one embodiment,
[0212] The instruction further carries: candidate labels;
[0213] When the annotation module re - determines a character label as the target label for the character, it is specifically used for:
[0214] Selecting a candidate label from the candidate labels carried by the instruction and determining the selected candidate label as the target label for the character.
[0215] The candidate label can be determined by the user and is a correct candidate label for the character, which correctly describes the character information of the character. Of course, the instruction can also carry many candidate labels at once, and one of them is a correct candidate label for the character.
[0216] When selecting a candidate tag from the candidate tags carried by the instruction, a candidate tag that correctly describes the character information of the character can be selected, and the selected candidate tag is determined as the target tag of the character.
[0217] Of course, other information can also be carried in the instruction, such as selection indication information. The selection indication information can indicate the image where the character is located (which can directly jump to the image) and the content of the character in the image. Of course, this is only an example here and is not a limitation. In this way, a correct candidate tag is selected for the character according to the selection indication information as the target character of the character.
[0218] In one embodiment, the character recognition model training device further includes:
[0219] An image acquisition module, configured to acquire an image collected from a specified scene, where the style of the characters in the specified scene is a specified style;
[0220] A region intercepting module, configured to intercept a character region from each acquired image, where the character region contains at least one character;
[0221] An image sample set determining module, configured to determine the image sample set according to all the intercepted character regions.
[0222] The image acquisition module is configured to acquire an image collected from a specified scene, where the style of the characters in the specified scene is a specified style. The specified scene can be, for example, an industrial scene. In an industrial scene, there are usually some devices for manufacturing, control, etc. There are some characters on these devices, such as production date, operation date, device model, device status parameters, and other characters. Generally speaking, the styles of these characters are specified. Of course, the specified scene can also be other scenes, as long as the style of the characters in the scene is the specified style.
[0223] Since the characters in the specified scene are of a specified style, using the images collected from the specified scene to train the character recognition model can train a character recognition model with higher accuracy and robustness with a smaller number of samples compared to a scene with variable styles. The character recognition model can be used to perform character recognition on the images collected from the specified scene.
[0224] The region intercepting module is configured to intercept a character region from each acquired image. The character region contains at least one character, and the specific number of characters is not limited and can be determined according to the actual situation.
[0225] When the region extraction module extracts the character region from each collected image, it is necessary to first locate the character region. The character detection model that has been trained can be used to locate the character region. The collected image is input into the character detection model, and the character detection model locates the region where the characters (which can be multiple characters close to each other) are located and outputs the region position information. The region corresponding to the region position information in the image is the character region. The above character detection model can be built with YOLO.
[0226] Of course, the method of locating the character region is not limited to this. For example, it can also be achieved by template matching and other methods.
[0227] The image sample set determination module is used to determine the image sample set according to all the extracted character regions, which can reduce the background of the images in the image sample set, thereby reducing the interference factors during model training. The model only needs to learn from the character regions, which is beneficial to improving the recognition performance of the model.
[0228] Specifically, the image sample set determination module can select a part of the character regions from all the extracted character regions for annotation and add them to the image sample set as the annotated images, while the other part of the character regions are added to the image sample set as the unannotated images.
[0229] The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0230] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units.
[0231] The present invention also provides an electronic device, including a processor and a memory; the memory stores a program that can be called by the processor; wherein, when the processor executes the program, the character recognition model training method described in the foregoing embodiments is implemented.
[0232] The embodiment of the character recognition model training device of the present invention can be applied to an electronic device. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the electronic device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 3 shown, Figure 3 is a hardware structure diagram of the electronic device where the character recognition model training device 100 of the present invention is located according to an exemplary embodiment. Except for Figure 3In addition to the processor 510, memory 530, interface 520, and non-volatile memory 540 shown, the electronic device where the device 100 is located in the embodiment usually further includes other hardware according to the actual functions of the electronic device, which will not be elaborated herein.
[0233] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for training a character recognition model, characterized in that, Including: Obtain an image collected from a specified scenario, where the style of characters in the specified scenario is a specified style; Extract a character region from each collected image, and the character region contains at least one character; Determine an image sample set according to all the extracted character regions; Select unlabeled images from the image sample set; the image sample set includes labeled images and unlabeled images; wherein, when selecting unlabeled images from the image sample set for the first time, the number of labeled images in the image sample set is less than the number of unlabeled images; Input the selected unlabeled images into a character recognition model to obtain the predicted character recognition results of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set; For each unlabeled image input into the character recognition model, determine the predicted character labels of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image; wherein, the predicted character information includes the predicted character position information and the character content; the predicted character label of a character is used to describe the predicted character information of the character; For each character in the unlabeled image, extract the region corresponding to the character position information in the predicted character label of the character from the unlabeled image; Extract the regions corresponding to the character position information in the predicted character labels of the characters from each unlabeled image, and display the extracted regions classified according to the predicted character content to the user, and display the regions with the same character content each time, so that the user can view the displayed regions to judge whether the correct character label is predicted for the character; If the indication information carried in the externally input instruction does not indicate that the correct character label is not predicted for the character, determine the character label as the target label of the character; if the indication information carried in the instruction indicates that the correct character label is not predicted for the character, select a candidate label from the candidate labels carried in the instruction, and determine the selected candidate label as the target label of the character; Label the unlabeled image in the image sample set according to the target labels of each character in the unlabeled image to obtain a labeled image; Train the character recognition model according to the labeled images in the image sample set; Check whether the set training end condition is currently met. If not, return to the operation of selecting unlabeled images from the image sample set for iterative training. If so, end the iterative training and determine the character recognition model as the target character recognition model; Wherein, after several iterative trainings, for the characters with a relatively high prediction error rate in the labeled images labeled during the iterative process, use the images containing the characters with a relatively high prediction error rate to optimize the character recognition model to obtain the target character recognition model; Wherein, selecting unlabeled images from the image sample set includes: If the number of unlabeled images in the image sample set is greater than or equal to the set number, select the set number of unlabeled images from the image sample set; If the number of unlabeled images in the image sample set is less than the set number, select all the remaining unlabeled images from the image sample set.
2. A character recognition model training device, characterized in that, Including: An image acquisition module, configured to acquire images collected from a specified scene, where the style of characters in the specified scene is a specified style; A region intercepting module, configured to intercept a character region from each of the acquired images, where the character region contains at least one character; An image sample set determination module, configured to determine an image sample set according to all the intercepted character regions; A selection module, configured to select unlabeled images from the image sample set; the image sample set includes labeled images and unlabeled images; wherein, when unlabeled images are selected from the image sample set for the first time, the number of labeled images in the image sample set is less than the number of unlabeled images; A prediction module, configured to input the selected unlabeled images into a character recognition model to obtain a predicted character recognition result of each unlabeled image input into the character recognition model; the character recognition model is trained according to the labeled images in the image sample set; A labeling module, configured to, for each unlabeled image input into the character recognition model, determine the predicted character label of each character in the unlabeled image according to the predicted character information in the predicted character recognition result of the unlabeled image; wherein, the predicted character information includes the predicted character position information and the character content; the predicted character label of a character is used to describe the predicted character information of the character; For each character in the unlabeled image, intercept the region corresponding to the character position information in the predicted character label of the character from the unlabeled image; Intercept the regions corresponding to the character position information in the predicted character labels of the characters from each unlabeled image, and classify and display the intercepted regions to the user according to the predicted character content, and display the regions with the same character content each time, so that the user can view the displayed regions to determine whether the correct character label is predicted for the character; If the indication information carried in the externally input instruction does not indicate that the correct character label is not predicted for the character, determine the character label as the target label of the character; if the indication information carried in the instruction indicates that the correct character label is not predicted for the character, select a candidate label from the candidate labels carried in the instruction, and determine the selected candidate label as the target label of the character; Label the unlabeled image in the image sample set according to the target labels of the characters in the unlabeled image to obtain a labeled image; A training module, configured to train the character recognition model according to the labeled images in the image sample set; Check whether the set training end condition is currently met. If not, return to the operation of selecting unlabeled images from the image sample set for iterative training. If so, end the iterative training and determine the character recognition model as the target character recognition model; Among them, after several iterations of training, the characters with a relatively high prediction error rate in the labeled images marked during the iteration process are used to optimize the character recognition model with the images containing the characters with a relatively high prediction error rate, so as to obtain the target character recognition model; Among them, when the selection module selects unlabeled images from the image sample set, it specifically is used for: If the number of unlabeled images in the image sample set is greater than or equal to the set number, select a set number of unlabeled images from the image sample set; If the number of unlabeled images in the image sample set is less than the set number, select all the remaining unlabeled images from the image sample set.
3. An electronic device, characterized in that, It includes a processor and a memory; the memory stores a program that can be called by the processor; among them, when the processor executes the program, it implements the character recognition model training method as described in claim 1.
Citation Information
Patent Citations
Character recognition and recognition model training methods, devices and systems, and storage medium
CN108875722A
Sample labeling method and device based on multiple models
CN109784391A