A training method and device for an identification model
In the field of text recognition, target scenes and other scenes are determined from each text recognition scenario, and the recognition models of other scenes are used to label and train the images of the target scenes, which solves the problem of inefficient training caused by relying on manual annotation in the prior art, and achieves more efficient recognition model training.
Patent Information
- Application Number
- CN202111579413.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-22
AI Technical Summary
In the prior art, the training of the recognition model relies on manual annotation, resulting in long time periods, high cost and poor training effect.
By determining the target scene and other scenes from each text recognition scene, the recognition model completed by training in other scenes is used as the candidate recognition model, the images of the target scene are marked, and the annotated image is input to the target recognition model for training.
It reduces the generation time and cost of training samples, improves the training efficiency of the target recognition model, and improves the recognition effect of the recognition model.
Smart Images

Figure CN114332873B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method and device for training a recognition model. Background Art
[0002] Text, as a way for humans to record and express information through symbols and pass it on for generations, is widely used in our lives. Text recognition, which can identify text in images as text and improve information input efficiency, is being applied in various fields.
[0003] In the prior art, a commonly used text recognition method is implemented based on a recognition model. Specifically, there are multiple scenarios for text recognition, such as road sign recognition and test paper recognition. Since the characteristics of text in different scenarios vary significantly, a corresponding recognition model must be determined for each scenario. The server that deploys the recognition model can first receive a text recognition request and, based on the text recognition request, determine the image for which text recognition is required and the corresponding scene. Then, the recognition model corresponding to the scene is determined, and the image is input into the recognition model to obtain the recognition result corresponding to the image output by the recognition model. Finally, the server can return the recognition result based on the text recognition request.
[0004] However, the recognition models in the prior art are usually trained based on manually labeled samples. The long time and high cost of manual labeling result in poor training effect of the recognition model. Summary of the Invention
[0005] This specification provides a method and device for training a recognition model to partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This manual provides a method for training a recognition model, including:
[0008] Determining a target scene and other scenes from each character recognition scene, and determining first training samples based on each image corresponding to the target scene;
[0009] For each other scenario, determine the trained recognition model corresponding to the other scenario as a candidate recognition model;
[0010] For each first training sample, the first training sample is input into at least one candidate recognition model, a candidate recognition result of the first training sample output by the at least one candidate recognition model is obtained, and a label of the first training sample is determined;
[0011] Each first training sample is input into the target recognition model to be trained corresponding to the target scene, each target recognition result output by the target recognition model is obtained, and based on the labeling of each first training sample and the target recognition result, the target recognition model corresponding to the target scene is trained.
[0012] Optionally, the candidate recognition result includes the probability that each character included in the first training sample belongs to each preset character;
[0013] Determining the label of the first training sample specifically includes:
[0014] For each candidate recognition result, determine the weight of the candidate recognition result;
[0015] The label of the first training sample is determined according to each candidate recognition result and its weight.
[0016] Optionally, the candidate recognition result includes the probability that each character included in the first training sample belongs to each preset character;
[0017] Before training the target recognition model, the method further includes:
[0018] For each first training sample, based on the label of the first training sample and a preset probability threshold, determining a first training sample whose label is not lower than the probability threshold as a target training sample for training the target recognition model;
[0019] Based on the annotations and target recognition results of each first training sample, the target recognition model corresponding to the target scene is trained, specifically including:
[0020] Based on the annotations of each target training sample and the target recognition results, the target recognition model corresponding to the target scene is trained.
[0021] Optionally, based on the annotations and target recognition results of each first training sample, training a target recognition model corresponding to the target scene specifically includes:
[0022] For each other scene, determine each second training sample and its label according to the image in the other scene;
[0023] Inputting each second training sample into the target recognition model to determine a target recognition result for each second training sample;
[0024] According to the labeling and target recognition results of each first training sample and the labeling and target recognition results of each second training sample, a loss is determined, and model parameters of the target recognition model are adjusted according to the loss.
[0025] Optionally, determining the loss based on the labeling and target recognition results of each first training sample and the labeling and target recognition results of each second training sample specifically includes:
[0026] Determining a first loss based on the labeling and target recognition results of the first training sample;
[0027] Determining a second loss based on the labeling and target recognition results of the second training sample;
[0028] Determine the weights corresponding to the first loss and the second loss respectively;
[0029] A total loss is determined based on the first loss and its weight, and the second loss and its weight.
[0030] Optionally, the method further includes:
[0031] Obtaining labeled images corresponding to the target scene as the third training samples;
[0032] Input each third training sample as an input into the target recognition model and the at least one candidate recognition model, and determine a target recognition result and a candidate recognition result of each third training sample outputted by the target recognition model and the at least one candidate recognition model respectively;
[0033] Determining the accuracy rate corresponding to the target recognition model based on the target recognition results and annotations of each third training sample, and determining the accuracy rate corresponding to each of the at least one candidate recognition models based on the candidate recognition results and annotations of each of the three training samples;
[0034] sorting the target recognition model and the candidate recognition models according to the accuracy rate;
[0035] According to the ranking, each candidate recognition model used to determine the label of the first training sample is re-determined.
[0036] Optionally, the method further includes:
[0037] Determining third training samples and their labels based on the labeled images corresponding to the target scene;
[0038] Inputting each third training sample as input into the target recognition model, and determining the target recognition result of each third training sample output by the target recognition model;
[0039] Determining the accuracy of the target recognition model based on the labels of each third training sample and the target recognition result;
[0040] When the accuracy is higher than a preset accuracy threshold, it is determined that the target recognition model training is completed.
[0041] This specification provides a training device for a recognition model, including:
[0042] A sample determination module, configured to determine a target scene and other scenes from each character recognition scene, and use each image corresponding to the target scene as each first training sample;
[0043] A first determination module is configured to determine, for each other scenario, a trained recognition model corresponding to the other scenario as a candidate recognition model;
[0044] a label determination module, configured to, for each first training sample, take the first training sample as input and input it into at least one candidate recognition model, obtain a candidate recognition result of the first training sample output by the at least one candidate recognition model, and determine a label for the first training sample;
[0045] The training module is used to input each first training sample into the target recognition model to be trained corresponding to the target scene, obtain each target recognition result output by the target recognition model, and train the target recognition model corresponding to the target scene based on the labeling and target recognition results of each first training sample.
[0046] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned recognition model training method.
[0047] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned recognition model training method when executing the program.
[0048] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0049] In the recognition model training method provided in this specification, the target scene and other scenes are determined from each text recognition scene, and each image corresponding to the target scene is used as each first training sample. For each other scene, the trained recognition model corresponding to the other scene is determined as a candidate recognition model. For each first training sample, the first training sample is used as input to at least one candidate recognition model, the label of the first training sample is determined, and each first training sample is input into the target recognition model to be trained to obtain each target recognition result output by the target recognition model. Based on the label and target recognition result of each first training sample, the target recognition model corresponding to the target scene is trained.
[0050] It can be seen from the above method that this method does not require manual labeling of samples, reduces the time and cost of generating training samples, and improves the training efficiency of the target recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0052] Figure 1 A flowchart of the training method for the recognition model provided in this specification;
[0053] Figure 2 A structural diagram for determining loss provided for this statement;
[0054] Figure 3 A training device for the recognition model provided in this manual;
[0055] Figure 4 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0056] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0057] Generally, there are many scenarios for text recognition, such as express delivery receipt recognition, bank card recognition, menu recognition, road sign recognition, and test paper recognition. Different scenarios require different fonts, formats, backgrounds, and other factors. Therefore, a recognition model must be trained for each scenario to recognize images containing text.
[0058] For each scene, training the corresponding recognition model usually requires a large number of manually labeled training samples.
[0059] In order to solve the problem of requiring a large number of manually labeled training samples in the existing technology, there are generally two solutions:
[0060] The first method synthesizes image data based on the characteristics of the text in the target scene image, and trains the model based on the synthesized images and annotations. The disadvantage is that the synthesized data differs significantly from the real data, resulting in poor recognition performance of the trained recognition model. The target scene is the scene for which the recognition model needs to be trained.
[0061] The second type is semi-supervised training, which blurs the unlabeled image, determines the blurred image, inputs the unlabeled image into the model, obtains the recognition result as the label of the blurred image, and trains the model based on the label of the blurred image and the labeled image.
[0062] The disadvantage is that it is difficult to process the original image and it is even more difficult to recognize the blurred image, resulting in poor results.
[0063] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0064] Figure 1 This is a flow chart of the training method for the recognition model provided in this specification, which specifically includes the following steps:
[0065] S100: Determine a target scene and other scenes from each character recognition scene, and determine first training samples based on each image corresponding to the target scene.
[0066] Generally, in the field of text recognition, an image containing text can be recognized through a recognition model to determine the text in the image, and then other steps can be performed based on the determined text.
[0067] Typically, a recognition model is pre-trained based on training samples by a server for training the model. This specification provides a method for training a recognition model, and similarly, the process of training the recognition model can be performed by a server for training the model.
[0068] The training model can be divided into a sample generation phase and a model training phase. In the sample generation phase, samples for training the model can be determined based on model and training requirements. In this specification, the server can first determine training samples for training the recognition model. Since recognition models generally recognize content contained in images in a target scene based on images in the target scene, the server can first determine images containing text in the target scene to determine training samples.
[0069] Based on this, the target scene and other scenes can be determined from the scenes of each character recognition.
[0070] Specifically, the server may first receive a training request, wherein the training request includes a target scene for which a recognition model needs to be determined. The scenes other than the target scene in the text recognition scene are referred to as "other scenes." The target scene and other scenes are scenes in the text recognition scene. The target scene is the scene for which a recognition model needs to be trained, and the other scenes are scenes in the text recognition scene for which relatively accurate recognition models have already been trained.
[0071] Therefore, after determining the target scene, the server may use the images corresponding to the target scene as first training samples, wherein the images corresponding to the target scene are unlabeled images in the target scene.
[0072] S102: For each other scenario, determine a trained recognition model corresponding to the other scenario as a candidate recognition model.
[0073] Different from the existing technology of blurring images containing text, using the recognition results of the original image as the annotation of the blurred image, and then training the recognition model based on the recognition results and annotations of the blurred image, which has too high recognition difficulty and leads to poor recognition effect, this solution proposes a new recognition model training method, which can be based on the recognition models in other scenes in the text recognition scene to train the recognition model in the target scene.
[0074] Based on this, the server can determine the trained recognition model corresponding to each other scene for the other scene.
[0075] Specifically, for each other scene, the server may determine a labeled image corresponding to the other scene.
[0076] Then, the server may determine a second training sample and its label according to the determined labeled image in the other scene.
[0077] Finally, the server may train the recognition model in the other scenario according to the determined second training sample.
[0078] Of course, the recognition model of the other scene can be obtained by other servers in advance based on the labeled images in the other scene, and the model parameters can be stored. Then, when the server determines the candidate recognition model, it can determine the model structure and model parameters of the recognition model corresponding to the other scene from the pre-stored model structures and model parameters of each other scene based on the identification of the other scene.
[0079] In addition, since model training is a stage, for each other scenario, the server can also obtain the corresponding trained recognition model of different stages of the other scenario.
[0080] Specifically, during the training of the recognition model for a particular scene, after the recognition model training is complete, the server that trained the model may continue to train the recognition model. However, the recognition model at any subsequent time can be considered converged, meaning that the recognition model has a higher accuracy rate for images in other scenes. The server can then select the recognition model at any time as a candidate recognition model after the recognition model training is complete.
[0081] S104: For each first training sample, take the first training sample as input and input at least one candidate recognition model, obtain the candidate recognition result of the first training sample output by the at least one candidate recognition model, and determine the label of the first training sample.
[0082] In one or more embodiments provided herein, the corresponding recognition model for each text recognition scenario can learn not only the characteristics of the text in that scenario, but also the characteristics of the text itself. Therefore, for each first training sample in the target scenario, the candidate recognition models corresponding to other scenarios produce results that incorporate the characteristics of the text itself in each first training sample in the target scenario. Therefore, the labels for the first training samples can be determined based on the candidate recognition models for other scenarios, and the target recognition model corresponding to the target scenario can be trained based on each first training sample and its labels.
[0083] Specifically, the server may input each first training sample into at least one candidate recognition model to obtain a candidate recognition result of the first training sample output by the at least one candidate recognition model.
[0084] The at least one candidate recognition model may be multiple candidate recognition models corresponding to one other scenario, or may be candidate recognition models corresponding to multiple other scenarios respectively.
[0085] Then, for each first training sample, the server may use the candidate recognition result of the first training sample as a label of the first training sample.
[0086] Of course, since there may be multiple candidate recognition results for the first training sample, the server may add up the candidate results for the first training sample and use the added result as the label for the first training sample.
[0087] Further, since the candidate recognition results of the determined first training samples are the probabilities that the characters in the first training samples belong to each preset classification, and usually a character belongs to a certain preset classification. For example, the text contained in the image is "hour", rather than the text contained in the image being "small - 80%, individual - 20%" and "hour - 60%, pair - 35%, rare character - 5%", etc. Therefore, the server can determine the state with the highest probability from the candidate recognition results according to the candidate recognition results corresponding to the first training samples and their confidence levels, and use it as the annotation of the first training sample. For example, if the probabilities that the first character contained in the first training sample corresponds to the classifications of hour, pair, and rare character are 60%, 30%, and 5% respectively, it can be determined that the annotation of the first training sample is hour. Of course, the annotation can also be hour - 60%. Then, the server can train the target recognition model of the target scenario based on the target recognition result and annotation of the first training sample.
[0088] Furthermore, since the similarity between the fonts of images in different scenarios is different. For example, the similarity between the images in the express bill recognition scenario and the menu recognition scenario is higher than the similarity between the images in the express bill recognition scenario and the road sign recognition scenario. Then, the similarity between the images of different other scenarios and the images of the target scenario is different, and the accuracy rates of the recognition results of the first training samples in the target scenario by the candidate recognition models of different other scenarios are different. Therefore, weights can be preset for the candidate recognition results corresponding to other scenarios in advance.
[0089] When the server determines the annotation of the first training sample, for each candidate recognition result, it can determine the weight of the candidate recognition result, and perform weighted summation according to each candidate recognition result and its weight to determine the annotation of the first training sample.
[0090] Of course, the server can also multiply each candidate recognition result and its weight, and select the state with the highest probability from the products as the annotation of the first training sample.
[0091] In one or more embodiments provided in this specification, since the confidence levels of the recognition models of other scenarios similar to the text characteristics of the images in the target scenario are relatively high, the server can determine an other scenario and obtain multiple recognition models corresponding to the other scenario as each candidate recognition model.
[0092] Therefore, the server can use the training sample as input, input it into each candidate recognition model, determine the candidate recognition results of the first training sample output by each candidate recognition model, and determine the annotation of the first training sample according to each candidate recognition result.
[0093] S106: Input each first training sample into the target recognition model to be trained corresponding to the target scenario, and obtain the target recognition result output by the target recognition model.
[0094] In one or more embodiments provided in this specification, for model training, it is necessary to train based on the annotation of the training sample and the result of the training sample obtained through the model. Therefore, the server can input each first training sample into the target recognition model of the target scenario to obtain the target recognition result of the first training sample.
[0095] Specifically, the server can use each first training sample as input and input it into the target recognition model to be trained corresponding to the target scenario, obtain the target recognition result of each first training sample output by the target recognition model, determine the first loss based on the annotation and the target recognition result of each first training sample, and adjust the model parameters of the target recognition model according to the first loss to complete the training of the target recognition model. As Figure 2 shown.
[0096] Figure 2 This is the structure diagram for determining the loss provided in this specification. Input the first training sample into candidate recognition model 1 and candidate recognition model 2 to determine the candidate recognition result 1 and candidate recognition result 2 of the first training sample, determine the annotation of the first training sample based on candidate recognition result 1 and candidate recognition result 2, and then input the first training sample into the target recognition model to determine the target recognition result output by the target model. Determine the loss based on the target recognition result and the annotation of the first training sample, and take minimizing the loss as the optimization goal to train the recognition model.
[0097] In addition, since the confidence of the candidate recognition model in the first training sample may not be sufficient, the server can also screen each first training sample based on the confidence of the annotation of each first training sample.
[0098] Specifically, for each first training sample, the server can determine the annotation of the first training sample and the probability that the characters included in the first training sample belong to the annotation. Taking the annotation as "Shi - 60%", the probability that the character "Shi" is included in the first training sample corresponding to this annotation is 60%.
[0099] Then, the server can determine the first training samples with a probability higher than the probability threshold as the target training samples according to the preset probability threshold and the probability that the characters in each first training sample belong to their annotation.
[0100] Finally, the server can take each target training sample as input and input it into the target recognition model corresponding to the target scene to obtain the target recognition result output by the target recognition model, and train the target recognition model based on the target recognition result and its annotation of the target training sample.
[0101] based on Figure 1 The method for training a recognition model is to determine the target scene and other scenes from each text recognition scene, and use the images corresponding to the target scene as the first training samples. For each other scene, the trained recognition model corresponding to the other scene is determined as a candidate recognition model. For each first training sample, the first training sample is used as input to at least one candidate recognition model, and the label of the first training sample is determined. Each first training sample is input into the target recognition model to be trained, and the target recognition results output by the target recognition model are obtained. Based on the label of each first training sample and the target recognition results, the target recognition model corresponding to the target scene is trained. This solution does not require manual labeling, and the model training efficiency is high.
[0102] Furthermore, compared to methods that synthesize image data based on the characteristics of text in an image of the target scene and then train the model based on this image data, this solution uses real data from the target scene to train the model. Therefore, the recognition model trained in this specification is more effective. Compared to methods that blur an image and then train the model based on the recognition results of the blurred image and the original image, the recognition model in this specification does not require recognition of the blurred image, resulting in lower recognition difficulty and better results.
[0103] Furthermore, since the labeling of the first training sample may not be accurate enough, when training the target recognition model, labeled images in other scenes may be used to train the model.
[0104] Specifically, the server may determine, for each other scene, the second training samples and their labels based on the images in the other scene.
[0105] Then, the server may input each second training sample into the target recognition model corresponding to the target scene to determine the target recognition result of each second training sample.
[0106] Finally, the server can determine the first loss based on the labeling and target recognition results of the first training sample, and determine the second loss based on the labeled target recognition results of the second training sample, and then determine the total loss based on the sum of the first loss and the second loss, and then adjust the model parameters of the target recognition model based on the total loss.
[0107] Of course, since the second training sample is determined from images of other scenes, the server may also preset weights for the scene and other scenes to reduce the impact of other scenes on the scene, which may result in a lower accuracy of the object recognition model. The server may then determine a total loss based on the preset first loss and its weight, and the second loss and its weight, and then adjust the model parameters of the object recognition model based on the total loss.
[0108] Furthermore, the weights can be varied according to the number of trainings. Specifically, the weight of the first loss is positively correlated with the number of trainings, and the weight of the second loss is negatively correlated with the number of trainings.
[0109] In addition, during the training process, since the confidence of the recognition results of the candidate recognition models corresponding to other scenarios may not be high, and the confidence of the target recognition model obtained by the annotation training based on the candidate recognition results may be low, the server can also replace the candidate recognition results used to determine the annotations of the first training sample during the training process.
[0110] Specifically, the server may determine each third training sample and its label according to the labeled images in the target scene.
[0111] Secondly, the server may take each third training sample as input into the target recognition model and the candidate recognition model used to determine the annotations of the first training sample, and determine the target recognition results and each candidate recognition result of each third training sample.
[0112] Then, the server may determine the accuracy of the target recognition model and the accuracy of each candidate recognition model based on the target recognition results and labels of each third training sample and the candidate recognition results and labels of each third training sample.
[0113] Finally, the server can sort the accuracy rates and determine the model with the highest accuracy in the sorting as the model used to determine the label of the first training sample. For example, if the accuracy of candidate recognition model 1 is 60%, the accuracy of candidate recognition model 2 is 70%, and the accuracy of the target recognition model is 75%, the server can determine candidate recognition model 2 and the target recognition model as the models used to determine the label of the first training sample, that is, the candidate recognition models.
[0114] After re-determining each candidate recognition model, the server may re-execute step S104 based on each candidate recognition model to determine the label of the first training sample.
[0115] Furthermore, since the annotation of the first training sample used to train the target recognition model is determined by the candidate recognition models of other scenarios, the server can also use the third training sample to determine whether the accuracy of the target recognition model is sufficient when training the target recognition model, that is, whether the training is completed.
[0116] Specifically, the server may take each third training sample as input and input it into the target recognition model, and determine the target recognition result of each third training sample output by the target recognition model.
[0117] Then, the server may determine the accuracy of the target recognition model based on the target recognition result and the annotation of the third training sample.
[0118] Finally, the server may determine whether the accuracy is higher than a preset accuracy threshold. If so, the server may determine that the object recognition model training is complete.
[0119] If not, the server may determine that the target recognition model training is not complete and still needs to be trained.
[0120] The above methods for training the recognition model provided in one or more embodiments of this specification are based on the same idea. This specification also provides a corresponding training device for the recognition model, such as Figure 3 shown.
[0121] Figure 3 The training device for the recognition model provided in this manual includes:
[0122] The sample determination module 200 is configured to determine a target scene and other scenes from each character recognition scene, and use each image corresponding to the target scene as each first training sample.
[0123] The first determination module 202 is configured to determine, for each other scenario, a trained recognition model corresponding to the other scenario as a candidate recognition model.
[0124] The labeling determination module 204 is used to input each first training sample into at least one candidate recognition model, obtain the candidate recognition result of the first training sample output by the at least one candidate recognition model, and determine the label of the first training sample.
[0125] The training module 206 is used to input each first training sample into the target recognition model to be trained corresponding to the target scene, obtain each target recognition result output by the target recognition model, and train the target recognition model corresponding to the target scene based on the labeling and target recognition results of each first training sample.
[0126] Optionally, the candidate recognition results include the probability that each character contained in the first training sample belongs to each preset character. The labeling determination module 204 is used to determine the weight of each candidate recognition result, and determine the label of the first training sample based on each candidate recognition result and its weight.
[0127] Optionally, the candidate recognition result includes the probability that each character contained in the first training sample belongs to each preset character. The annotation determination module 204 is used to determine, for each first training sample, based on the annotation of the first training sample and the preset probability threshold, a first training sample whose annotation is not lower than the probability threshold, as the target training sample for training the target recognition model, and based on the annotation of each target training sample and the target recognition result, the target recognition model corresponding to the target scene is trained.
[0128] Optionally, the annotation determination module 204 is used to determine each second training sample and its annotation for each other scene based on the image in the other scene, input each second training sample into the target recognition model, determine the target recognition result of each second training sample, determine the loss based on the annotation and target recognition result of each first training sample and the annotation and target recognition result of each second training sample, and adjust the model parameters of the target recognition model based on the loss.
[0129] Optionally, the labeling determination module 204 is used to determine a first loss based on the labeling and target recognition results of the first training sample, determine a second loss based on the labeling and target recognition results of the second training sample, determine the weights corresponding to the first loss and the second loss respectively, and determine the total loss based on the first loss and its weight, and the second loss and its weight.
[0130] Optionally, the annotation determination module 204 is used to obtain annotated images corresponding to the target scene as each third training sample, and input each third training sample as input into the target recognition model and the at least one candidate recognition model, determine the target recognition results and candidate recognition results of each third training sample output by the target recognition model and the at least one candidate recognition model respectively, determine the accuracy corresponding to the target recognition model based on the target recognition results and their annotations of each third training sample, and determine the accuracy corresponding to the at least one candidate recognition model based on the candidate recognition results and their annotations of each third training sample, sort the target recognition model and the candidate recognition models according to the accuracy, and re-determine the candidate recognition models used to determine the annotation of the first training sample based on the sorting.
[0131] Optionally, the annotation determination module 204 is used to determine each third training sample and its annotation based on the annotated image corresponding to the target scene, input each third training sample as input into the target recognition model, determine the target recognition results of each third training sample output by the target recognition model, determine the accuracy of the target recognition model based on the annotations of each third training sample and the target recognition results, and when the accuracy is higher than a preset accuracy threshold, it is determined that the target recognition model training is completed.
[0132] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 Provides a training method for the recognition model.
[0133] This manual also provides Figure 4 The schematic structure diagram of the electronic device shown in FIG. Figure 4 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0134] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0135] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0136] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0137] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0138] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0140] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0142] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0143] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0144] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0145] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0146] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0147] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0148] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0149] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for training an identification model, characterized in that, it includes: Determine a target scenario and other scenarios from each text recognition scenario, and determine each first training sample according to each image corresponding to the target scenario; For each other scenario, determine the trained identification model corresponding to this other scenario as a candidate identification model; For each first training sample, use this first training sample as input and input it into at least one candidate identification model to obtain the candidate recognition results of this first training sample output by the at least one candidate identification model, and determine the annotation of this first training sample; Input each first training sample into the target identification model to be trained corresponding to the target scenario to obtain each target recognition result output by the target identification model, and train the target identification model corresponding to the target scenario based on the annotations and target recognition results of each first training sample; The method further includes: Obtain the annotated images corresponding to the target scenario as each third training sample; Use each third training sample as input and input it into the target identification model and the at least one candidate identification model to determine the target recognition results and candidate recognition results of each third training sample respectively output by the target identification model and the at least one candidate identification model; Determine the accuracy rate corresponding to the target identification model according to the target recognition results and their annotations of each third training sample, and determine the accuracy rates corresponding to the at least one candidate identification model respectively according to the candidate recognition results and their annotations of each third training sample; Sort the target identification model and each candidate identification model according to the accuracy rate; According to the sorting, re-determine each candidate identification model used to determine the annotation of the first training sample.
2. The method according to claim 1, characterized in that, The candidate recognition result includes the probability that each character included in the first training sample belongs to each preset character; Determining the annotation of this first training sample specifically includes: For each candidate recognition result, determine the weight of this candidate recognition result; Determine the annotation of this first training sample according to each candidate recognition result and its weight.
3. The method according to claim 1, characterized in that, The candidate recognition result includes the probability that each character included in the first training sample belongs to each preset character; Before training the target identification model, the method further includes: For each first training sample, determine the first training sample whose annotation is not lower than the probability threshold according to the annotation of this first training sample and a preset probability threshold as the target training sample for training the target identification model; Training the target identification model corresponding to the target scenario based on the annotations and target recognition results of each first training sample specifically includes: Training the target identification model corresponding to the target scenario based on the annotations and target recognition results of each target training sample.
4. The method according to claim 1, characterized in that, Training the target identification model corresponding to the target scenario based on the annotations and target recognition results of each first training sample specifically includes: For each other scenario, determine each second training sample and its annotation according to the images in that other scenario; Input each second training sample into the target recognition model to determine the target recognition result of each second training sample; Determine the loss according to the annotation and target recognition result of each first training sample, and the annotation and target recognition result of each second training sample, and adjust the model parameters of the target recognition model according to the loss.
5. The method according to claim 4, wherein, Determining the loss according to the annotation and target recognition result of each first training sample, and the annotation and target recognition result of each second training sample specifically includes: Determine the first loss according to the annotation and target recognition result of the first training sample; Determine the second loss according to the annotation and target recognition result of the second training sample; Respectively determine the weights corresponding to the first loss and the second loss; Determine the total loss according to the first loss and its weight, and the second loss and its weight.
6. The method according to claim 1, wherein, The method further includes: Determine each third training sample and its annotation according to the annotated images corresponding to the target scenario; Input each third training sample as input into the target recognition model to determine the target recognition result of each third training sample output by the target recognition model; Determine the accuracy rate of the target recognition model according to the annotation and target recognition result of each third training sample; When the accuracy rate is higher than a preset accuracy rate threshold, determine that the training of the target recognition model is completed.
7. A computer-readable storage medium, wherein, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 6 above is implemented.
8. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the method described in any one of claims 1 to 6 above is implemented.
Citation Information
Patent Citations
Prediction model training method and apparatus for a new scene
CN109359793A
Identification model training method and device
CN112801229A