A method, apparatus, device, and readable storage medium for training an identification model

By optimizing the sample data size and proportion, the problem of long model training cycle and high resource utilization is solved, and efficient model training is achieved.

CN114693999BActive Publication Date: 2025-07-08阳光保险集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210502256.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-07-08
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

During the training process of existing models, in order to improve robustness and accuracy, the use of large amounts of corpus data leads to problems such as excessive training cycle and excessive resource utilization.

Method used

By obtaining the preset number of basic sample pictures and the preset proportion of special sample pictures, train the basic model, adjust the sample data volume and proportion, to achieve the appropriate data volume, and save training resources and time.

Benefits of technology

On the premise of ensuring the accuracy of the model, training time and resource consumption are significantly reduced and training efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693999B_ABST
    Figure CN114693999B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device, and readable storage medium for training an identification model. The method includes obtaining a preset proportion of special sample pictures from a preset number of basic sample pictures and a preset number of special sample pictures corresponding to documents of a preset type, where the basic sample pictures are intercepted pictures of corpus fields in a basic corpus, and the special sample pictures contain the same type of fields in the documents of the preset type; training a basic model based on the preset number of basic sample pictures and the preset proportion of special sample pictures among the preset number of special sample pictures to obtain an identification model for document picture identification. On the premise of meeting the model accuracy rate, this method can achieve the effect of saving resources and training time for training the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of model training. Specifically, it relates to a method, device, equipment, and readable storage medium for training an identification model. Background Art

[0002] With the progress of technology, the training of various models is constantly being optimized. Model training plays an increasingly important role in the field of technology. For example, the training of common single-document models plays a great role in the text recognition of certificates.

[0003] However, currently, during the process of model training, in order to make the model more robust and accurate, a large amount of corpus data is often used. As a result, problems such as too long a model training cycle and too high a resource occupancy occur.

[0004] Therefore, during the process of model training, how to save the resources for model training is a technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method for training an identification model. Through the technical solutions of the embodiments of this application, it is possible to achieve the effect of saving the resources and training time for training the model on the premise of meeting the model accuracy.

[0006] In a first aspect, the embodiments of this application provide a method for training an identification model, including: obtaining a preset proportion of special sample images from a preset number of basic sample images corresponding to a preset type of single document and a preset number of special sample images, where the basic sample images are intercepted images of corpus fields in a basic corpus, and the special sample images contain the same type of fields in the preset type of single document; training a basic model based on a preset number of basic sample images and a preset proportion of special sample images from a preset number of special sample images to obtain an identification model for single-document image recognition.

[0007] In the above process, by training the basic model with a preset number of basic sample images and a preset proportion of special sample images from a preset number of special sample images that are set in advance, the time for preparing sample data can be saved, the most suitable amount of data for model training can be set, and the resources for model training can be saved.

[0008] In one embodiment, after training the basic model based on a preset number of basic sample images and a preset proportion of special sample images from a preset number of special sample images to obtain an identification model for single-document image recognition, it further includes:

[0009] Input the image of the document to be recognized into the recognition model to obtain the recognition result, where the image of the document to be recognized is a document of a preset type or a document of a non-preset type;

[0010] Calculate the recognition accuracy of each type of field in the image of the document to be recognized based on the recognition result;

[0011] Based on the recognition accuracy of each type of field, increase the number of special sample images corresponding to each type of field when the recognition accuracy is lower than the preset range to obtain the increased special sample images;

[0012] Use the increased special sample images to train the recognition model to obtain a document recognition model for recognizing the image of the document to be recognized.

[0013] In the above process, when using the recognition model, the corresponding training samples can be prepared for the recognition model according to the accuracy of each type of field recognized by the recognition model in the new document image, which can improve the recognition accuracy of the model for new documents.

[0014] In one embodiment, before obtaining the preset number of basic sample images corresponding to the documents of the preset type and the preset proportion of the special sample images among the preset number of special sample images, it further includes:

[0015] Reduce the number of basic sample images used to train the test model, and determine the minimum number of basic sample images during the training of the test model when the recognition accuracy of the test model remains within the preset range;

[0016] Use the minimum number of basic sample images as the preset number of basic sample images.

[0017] In the above process, while ensuring the accuracy of model training, reduce the data volume of the basic sample images for model training until the minimum data volume of the basic sample images required for model training is determined, which can minimize the data volume of the basic sample images used for model training and thus improve the time for model training.

[0018] In one embodiment, before obtaining the preset number of basic sample images corresponding to the documents of the preset type and the preset proportion of the special sample images among the preset number of special sample images, it further includes:

[0019] Reduce the number of special sample images used to train the test model, and determine the minimum number of special sample images during the training of the test model when the recognition accuracy of the test model remains within the preset range;

[0020] Use the minimum number of special sample images as the preset number of special sample images.

[0021] In the above process, while ensuring the accuracy of model training, the amount of special sample picture data for model training is reduced until the minimum amount of special sample picture data required for model training is determined, which can minimize the amount of special sample picture data used in model training and thus improve the time for model training.

[0022] In one embodiment, after using the minimum number of special sample pictures as the preset number of special sample pictures, it further includes:

[0023] Under the condition that the number of the preset number of special sample pictures remains unchanged and the recognition accuracy of the test model is within the preset range, adjust the proportion of each type of special sample picture in the preset number of special sample pictures, and determine the special sample pictures corresponding to the proportion when the recognition accuracy of the test model is the highest during the training and testing of the model;

[0024] Use the special sample pictures corresponding to the proportion as the special sample pictures with a preset proportion in the preset number of special sample pictures.

[0025] In the above process, while ensuring that the accuracy of model training is within the preset range and the amount of special sample image data remains unchanged, adjusting the proportion of special sample pictures for model training can improve the time for model training.

[0026] In a second aspect, an embodiment of the present application provides a method for identifying document pictures, including obtaining the document picture to be identified, inputting the document picture to be identified into a pre-trained recognition model for document picture recognition, and obtaining the recognition result corresponding to the document picture to be identified, where the recognition model is trained from a basic model based on a preset number of basic sample pictures and a preset proportion of special sample pictures in a preset number of special sample pictures.

[0027] In the above process, training the basic model with a preset number of basic sample pictures and a preset proportion of special sample pictures in a preset number of special sample pictures set in advance can save the time for preparing sample data, can set the most suitable data volume for model training, and save the resources for model training.

[0028] In a third aspect, an embodiment of the present application provides a device for training a recognition model, including:

[0029] Optionally, the device further includes:

[0030] A second training module, which is used after the training module trains a basic model based on a preset number of basic sample images and a preset proportion of special sample images among a preset number of special sample images to obtain a recognition model for document image recognition, inputs a document image to be recognized into the recognition model to obtain a recognition result, where the document image to be recognized is a document of a preset type or a document of a non-preset type;

[0031] Calculate the recognition accuracy of each type of field in the document image to be recognized based on the recognition result;

[0032] Based on the recognition accuracy of each type of field, increase the number of special sample images corresponding to each type of field when the value of the recognition accuracy is lower than the preset range to obtain the increased special sample images;

[0033] Use the increased special sample images to train the recognition model to obtain a document recognition model for recognizing the document image to be recognized.

[0034] Optionally, the device further includes:

[0035] A first determination module, which is used to reduce the number of basic sample images for training a test model before the acquisition module acquires a preset number of basic sample images corresponding to a document of a preset type and a preset proportion of special sample images among a preset number of special sample images, and determine the minimum number of basic sample images during training the test model when the recognition accuracy of the test model remains within the preset range;

[0036] Use the minimum number of basic sample images as the preset number of basic sample images.

[0037] Optionally, the device further includes:

[0038] A second determination module, which is used to reduce the number of special sample images for training a test model before the acquisition module acquires a preset number of basic sample images corresponding to a document of a preset type and a preset proportion of special sample images among a preset number of special sample images, and determine the minimum number of special sample images during training the test model when the recognition accuracy of the test model remains within the preset range;

[0039] Use the minimum number of special sample images as the preset number of special sample images.

[0040] Optionally, the device further includes:

[0041] An adjustment module, configured to, after the second determination module uses the least number of special sample pictures as the preset number of special sample pictures, and when the number of the preset number of special sample pictures remains unchanged and the recognition accuracy of the test model is within the preset range, adjust the proportion of each type of special sample picture in the preset number of special sample pictures, and determine the special sample pictures with the corresponding proportion when the recognition accuracy of the test model is the highest when training the test model;

[0042] Use the special sample pictures with the corresponding proportion as the special sample pictures with the preset proportion in the preset number of special sample pictures.

[0043] Fourthly, an apparatus for recognizing document pictures according to an embodiment of the present application includes:

[0044] An acquisition module, configured to acquire the document pictures to be recognized;

[0045] A recognition module, configured to input the document pictures to be recognized into a recognition model for document picture recognition that is pre-trained, and obtain a recognition result corresponding to the document pictures to be recognized, where the recognition model is trained on a basic model based on a preset number of basic sample pictures and a preset proportion of special sample pictures in the preset number of special sample pictures.

[0046] Fifthly, an electronic device according to an embodiment of the present application includes a processor and a memory, where the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the method provided in the first aspect above are run.

[0047] Sixthly, an electronic device according to an embodiment of the present application includes a processor and a memory, where the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the method provided in the second aspect above are run.

[0048] Seventhly, a readable storage medium according to an embodiment of the present application stores a computer program thereon, and when the computer program is executed by a processor, the steps in the method provided in the first aspect above are run.

[0049] Eighthly, a readable storage medium according to an embodiment of the present application stores a computer program thereon, and when the computer program is executed by a processor, the steps in the method provided in the second aspect above are run.

[0050] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0052] Figure 1 It is a flowchart of a method for training an identification model provided by an embodiment of the present application;

[0053] Figure 2 It is a flowchart of a method for identifying document images provided by an embodiment of the present application;

[0054] Figure 3 It is a schematic block diagram of an apparatus for training an identification model provided by an embodiment of the present application;

[0055] Figure 4 It is a schematic block diagram of an apparatus for identifying document images provided by an embodiment of the present application;

[0056] Figure 5 It is a schematic structural diagram of an apparatus for training an identification model provided by an embodiment of the present application;

[0057] Figure 6 It is a schematic structural diagram of an apparatus for identifying document images provided by an embodiment of the present application. Detailed Description of the Embodiments

[0058] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0059] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0060] This application is applied to the scenario of model training. The model is specifically used for the recognition of document images. Specifically, in the scenario of model training, the model can be trained according to a preset number of sample data that has been tested in advance.

[0061] However, in the current process of model training, in order to make the model more robust and have a higher accuracy rate, a large amount of corpus data is often used. As a result, problems such as too long a model training cycle and too high a resource occupancy occur. Based on the above process, the sample image data is divided into basic sample image data and special sample image data. The size of the synthesized data of the basic sample images in traditional model training is generally 2 million. According to the types of documents to be recognized, the types and quantities of special sample images are designed. Generally, for example, special sample images of dates: data volume of 150,000, special sample images of prices: data volume of 100,000, special sample images of numbers: data volume of 100,000, special sample images of English numbers: data volume of 100,000, special sample images of capital numbers: data volume of 100,000, special sample images of tax rates: data volume of 30,000, special sample images of slashes: data volume of 30,000, special sample images of license plate numbers: data volume of 100,000, special sample images of addresses: data volume of 200,000, special sample images of vehicle identification numbers: data volume of 100,000, special sample images of business licenses: data volume of 35,000, special sample images of short numbers: 100,000, other special sample images: 100,000. The total data volume of the above basic sample images and special sample images is 3.245 million. The character accuracy rate of the model obtained by using the above samples is 97.77%, and the field accuracy rate is 90.90%. A total of 183,000 iteration steps are required. It can be seen that a large amount of time and resources are needed.

[0062] Therefore, in this application, a preset proportion of special sample images is obtained from the preset number of basic sample images corresponding to the preset types of documents and the preset number of special sample images. Among them, the basic sample images are intercepted images of corpus fields in the basic corpus, and the special sample images contain the same type of fields in the preset types of documents; the basic model is trained based on the preset number of basic sample images and the preset proportion of special sample images in the preset number of special sample images to obtain an identification model for document image recognition. The effect of saving the resources for model training is achieved.

[0063] In the embodiments of this application, the execution entity can be the device for training and recognizing the model in the system for training and recognizing the model. In practical applications, the device for training and recognizing the model can be electronic devices such as terminal devices and servers, which are not limited herein.

[0064] Next, in combination with Figure 1 A detailed description will be given of the method for training and recognizing the model in the embodiments of this application.

[0065] Please refer to Figure 1 ,Figure 1 This is a flowchart of a method for training an identification model provided by an embodiment of the present application. As Figure 1 shown, the method for training the identification model includes:

[0066] Step 110: Obtain a preset proportion of special sample images from a preset number of basic sample images corresponding to documents of a preset type and a preset number of special sample images.

[0067] In the above process, by setting in advance a preset number of basic sample images and a preset proportion of special sample images among the preset number of special sample images, an appropriate number of sample data can be provided for the training of the model, saving time for the model training.

[0068] Among them, the basic sample images are intercepted images of corpus fields in the basic corpus. For example, they are generally news images. In news, ten fields are grouped together to form an image with a height of 32 and a width of 400. The special sample images contain the same type of fields in the documents of the preset type and are images composed of some fixed fields. For example, they are composed entirely of arrays of 0-9, or can also be composed of 26 English letters, etc. The documents of the preset type can be identity cards, driving licenses, vehicle registration certificates, bank cards, marriage certificates, new car invoices, used car invoices, business licenses, etc. The present application is not limited to this.

[0069] In addition, when preparing the sample images, the identity card can be divided into irregular and regular fields according to the field information. Irregular fields: such as name. Regular fields: such as gender, date of birth, address, and ID number, etc. Because gender only includes male and female, the date of birth and ID number generally only include numbers, and the address is generally in the form of xx Province xx City xx District xx Building xx Number, etc., so they are regular fields. The new car invoice is similar to the identity card and can also be divided into irregular and regular fields; Irregular fields: such as the name of the purchaser, the name of the seller, and the invoicer, etc.; Regular fields: such as the taxpayer identification number, address, and phone number, etc. The present application is not limited to this.

[0070] Specifically, before executing step 110, the following steps can also be adopted:

[0071] Step S1101: Reduce the number of basic sample images used for training the test model, and determine the minimum number of basic sample images when training the test model under the condition that the recognition accuracy rate of the test model remains within the preset range.

[0072] Step S1102: Use the minimum number of basic sample images as the preset number of basic sample images.

[0073] In the above process, while ensuring the accuracy of model training, reduce the amount of basic sample image data for model training until the minimum amount of basic sample image data required for model training is determined, which can minimize the amount of basic sample image data used in model training and thus improve the time for model training.

[0074] Among them, the test model is used to test the number of sample images required by the model recognition model. The preset range of the accuracy rate can be set manually. For example: 97.5%-98%, or it can be set according to the accuracy rate of training the model with a large number of historical image samples. For example, if the accuracy rate of training the model with a large number of historical image samples is 97.68%, then the range value amplitude based on 97.68% can be set. For example, if the upper and lower amplitudes are 0.2%, then the preset range of the accuracy rate is 97.88%-97.48%. This application is not limited to this. Based on the traditional sample data volume, this application reduces the basic sample image data volume from 2 million to 1 million. The character recognition accuracy rate of the recognition model is 97.78%, the field accuracy rate is 91.03%, and the number of iteration steps is 126,000 steps. It can be seen that the accuracy rate remains basically unchanged, but the training time is reduced to 69% of the previous time. When the basic sample image data volume is reduced from 2 million to 500,000, the character recognition accuracy rate of the recognition model is 97.75%, the field accuracy rate is 90.95%, and the number of iteration steps is 97,000 steps. It can be seen that the accuracy rate also remains basically unchanged, but the training time is reduced to 53% of the previous time. When the basic sample image data volume is reduced from 2 million to 400,000, the number of iteration steps has exceeded 183,000 steps, and the character recognition accuracy rate of the recognition model is 95.17%, and the field accuracy rate is 87.67%. It can be seen that the approximate balance point between the data volume of the basic sample image and the recognition accuracy rate of the model is 500,000. Designing the basic sample image as 2 million is not very reasonable, which seriously increases the training time of the model. Therefore, the preset number of basic sample images for model training is 500,000.

[0075] Specifically, before performing step 110, the following steps can also be adopted:

[0076] Step S11001: Reduce the number of special sample images used to train the test model, and determine the minimum number of special sample images during the training of the test model under the condition that the recognition accuracy rate of the test model remains within the preset range;

[0077] Step S11002: Use the minimum number of special sample images as the preset number of special sample images.

[0078] In the above process, while ensuring the accuracy of model training, the amount of special sample image data for model training is reduced until the minimum amount of special sample image data required for model training is determined, which can minimize the amount of special sample image data used in model training and thus improve the time for model training.

[0079] For example, the number of basic sample images is 500,000, and the number of special sample images is 1,245,000. Based on this, the number of special sample images is reduced to 819,000. The character accuracy rate of the recognition model is 97.68%, the field accuracy rate is 90.75%, and the number of iteration steps is 75,000. It can be seen that when the number of special sample images is reduced to 819,000, the model accuracy remains basically unchanged, and the training time is reduced to 40% of the previous time. When the number of special sample images is reduced to 500,000, the character accuracy rate of the recognition model is 95.37%, the field accuracy rate is 87.64%, and the number of iteration steps is 173,000. When the number of special sample images is reduced to 500,000, the model accuracy decreases significantly. Therefore, the number of special sample images with a preset quantity for model training is 500,000.

[0080] Further, after performing step S11002, the following steps can also be adopted:

[0081] Step S110021: Under the condition that the quantity of the special sample images with a preset quantity remains unchanged and the recognition accuracy rate of the test model is within the preset range, adjust the proportion of each type of special sample image in the special sample images with a preset quantity, and determine the special sample images corresponding to the proportion when the recognition accuracy rate of the test model is the highest during the training of the test model;

[0082] Step S110022: Use the special sample images corresponding to the proportion as the special sample images with a preset proportion in the special sample images with a preset quantity.

[0083] In the above process, while ensuring that the accuracy of model training is within the preset range and the amount of special sample image data remains unchanged, adjusting the proportion of the special sample images for model training can improve the time for model training.

[0084] For example, when the number of basic sample pictures is 500,000 and the number of special sample pictures is 500,000, through analysis, it is found that the number of digital fields has decreased significantly without changing the total amount of sample picture data. Other fields such as address fields, slash fields, and large digital character fields have not decreased. Increase the number of digital special sample pictures and English digital special sample pictures by 50,000 each, decrease the number of address special sample pictures by 70,000, decrease the number of slash special sample pictures by 15,000, and decrease the number of large digital character special sample pictures by 15,000. In this case, the character accuracy rate of the recognition model is 97.70% and the field accuracy rate is 90.65%. Thus, it can be seen that without changing the total amount of sample picture data, adjusting the proportion of special sample picture categories can improve the model accuracy rate, and the training time is also reduced to 31% of the previous time. Another example is when the number of special sample pictures is reduced to 400,000, the character accuracy rate of the recognition model is 95.45% and the field accuracy rate is 88.23%. Through analysis, it is found that the data of each field has decreased significantly. Therefore, it can be known that by adjusting the proportion of special sample pictures, the accuracy rate within the preset range cannot be achieved.

[0085] Step 120: Train the basic model based on a preset proportion of special sample pictures among a preset number of basic sample pictures and a preset number of special sample pictures to obtain a recognition model for document picture recognition.

[0086] In the above process, training the basic model based on a preset proportion of special sample pictures among a preset number of basic sample pictures and a preset number of special sample pictures can save the time for determining model data and the time for model training.

[0087] Further, after executing Step 120, the following steps can also be adopted:

[0088] Step S1201: Input the document picture to be recognized into the recognition model to obtain a recognition result, where the document picture to be recognized is a document of a preset type or a document of a non-preset type.

[0089] Step S1202: Calculate the recognition accuracy rate of each type of field in the document picture to be recognized based on the recognition result.

[0090] Step S1203: Based on the recognition accuracy rate of each type of field, increase the number of special sample pictures corresponding to each type of field when the recognition accuracy rate is lower than the preset range to obtain the increased special sample pictures.

[0091] Step S1204: Use the increased special sample pictures to train the recognition model to obtain a document recognition model for recognizing the document picture to be recognized.

[0092] In the above process, when using the recognition model, according to the accuracy rate of each type of field recognized by the recognition model in the new document image, corresponding training samples can be prepared for the recognition model, which can improve the accuracy rate of the model for new document recognition and play a role in quickly training a corresponding recognition model for new documents.

[0093] Among them, the document image to be recognized can be a document of the preset type or other new documents that are not of the above types. Taking a new value-added invoice as an example, adjusting the font and background of the invoice image can make the recognition accuracy of the model higher. Through analysis, it is known that the recognition accuracy of some irregular fields such as military service status and religious belief is significantly lower than that of some regular fields such as date and ID number, resulting in a decrease in the overall accuracy. Therefore, a certain number of special sample images can be designed, and the font and background of the invoice image are adjusted. The character accuracy of the trained model is 97.93%, and the field accuracy is 91.07%. The accuracy of the model is significantly improved. It can be seen that the above method for training the document recognition model has good support for the scalability of new document recognition. Without changing the accuracy of the model, it can effectively reduce the training resources and training time. In addition, emphasizing this training method can be conveniently and quickly copied to the OCR recognition of other documents, has replicability, and can be iterated quickly.

[0094] The foregoing described the method for training the recognition model through Figure 2 figures. Next, a method for recognizing a document image will be described in combination with Figure 2 figures.

[0095] Please refer to Figure 2 , Figure 2 which is a flowchart of a method for recognizing a document image provided by an embodiment of the present application. As shown in Figure 2 , the method for recognizing a document image includes:

[0096] Step 210: Obtain the document image to be recognized.

[0097] Step 220: Input the document image to be recognized into a recognition model for document image recognition that has been trained in advance, and obtain a recognition result corresponding to the document image to be recognized.

[0098] In the above process, by training the basic model with a preset proportion of special sample images among the preset number of basic sample images and the preset number of special sample images set in advance, the time for preparing sample data can be saved, the most suitable data volume can be set for the training of the model, and the resources for model training can be saved.

[0099] Among them, the recognition model is obtained by training the basic model with a preset proportion of special sample pictures among a preset number of basic sample pictures and a preset number of special sample pictures, and the recognition result includes the fields on the document

[0100] The foregoing has described the method for recognizing document pictures. Next, a device for training a recognition model will be described in conjunction with Figure 2 Describe the device for training the recognition model Figure 3

[0101] Please refer to Figure 3 For the schematic block diagram of a device 300 for training a recognition model provided in an embodiment of the present application, the device 300 may be a module, a program segment, or code on an electronic device. The device 300 corresponds to the above Figure 1 Method embodiment, and can execute Figure 1 Each step involved in the method embodiment. The specific functions of the device 300 can be seen in the following description. To avoid repetition, the detailed description is appropriately omitted here

[0102] Optionally, the device 300 includes:

[0103] An acquisition module 310, configured to acquire a preset proportion of special sample pictures among a preset number of basic sample pictures and a preset number of special sample pictures corresponding to a preset type of document. The basic sample pictures are pictures synthesized from corpus fields in a basic corpus, and the special sample pictures contain the same type of fields in the preset type of document

[0104] A training module 320, configured to train the basic model with a preset proportion of special sample pictures among a preset number of basic sample pictures and a preset number of special sample pictures to obtain a recognition model for document picture recognition

[0105]

[0106] Optionally, the device further includes: A second training module, configured to, after the training module trains the basic model with a preset proportion of special sample pictures among a preset number of basic sample pictures and a preset number of special sample pictures to obtain a recognition model for document picture recognition, input the document picture to be recognized into the recognition model to obtain a recognition result, where the document picture to be recognized is a document of a preset type or a document of a non-preset type

[0107] Calculate the recognition accuracy of each type of field in the document picture to be recognized based on the recognition result

[0108] Based on the recognition accuracy of each type of field, increase the number of special sample pictures corresponding to each type of field when the recognition accuracy is lower than the preset range to obtain the increased special sample pictures

[0109] ​Train the recognition model using the increased special sample images to obtain a document recognition model for recognizing the document images to be recognized.

[0110] Optionally, the apparatus further includes:

[0111] A first determination module, configured to reduce the number of basic sample images for training the test model before the acquisition module acquires a preset proportion of special sample images from the preset number of basic sample images and the preset number of special sample images corresponding to the preset types of documents, and determine the minimum number of basic sample images for training the test model when the recognition accuracy of the test model remains within a preset range;

[0112] Use the minimum number of basic sample images as the preset number of basic sample images.

[0113] Optionally, the apparatus further includes:

[0114] A second determination module, configured to reduce the number of special sample images for training the test model before the acquisition module acquires a preset proportion of special sample images from the preset number of basic sample images and the preset number of special sample images corresponding to the preset types of documents, and determine the minimum number of special sample images for training the test model when the recognition accuracy of the test model remains within a preset range;

[0115] Use the minimum number of special sample images as the preset number of special sample images.

[0116] Optionally, the apparatus further includes:

[0117] An adjustment module, configured to, after the second determination module uses the minimum number of special sample images as the preset number of special sample images, and when the number of the preset number of special sample images remains unchanged and the recognition accuracy of the test model is within a preset range, adjust the proportion of each type of special sample image in the preset number of special sample images, and determine the special sample images corresponding to the proportion when the recognition accuracy of the test model is the highest during training of the test model;

[0118] Use the special sample images corresponding to the proportion as the preset proportion of the special sample images in the preset number of special sample images.

[0119] The foregoing has Figure 3 described the apparatus for training the recognition model. Next, in combination with Figure 4 describe the apparatus for document image recognition.

[0120] Please refer to Figure 4, is a schematic block diagram of a document image recognition device 400 provided in an embodiment of the present application. The device 400 may be a module, a program segment, or code on an electronic device. The device 400 corresponds to the above Figure 1 method embodiment and is capable of executing Figure 1 each step involved in the method embodiment. The specific functions of the device 400 can be referred to the descriptions below. To avoid repetition, the detailed descriptions are appropriately omitted here.

[0121] Optionally, the device includes:

[0122] An acquisition module 410, configured to acquire the document image to be recognized;

[0123] A recognition module 410, configured to input the document image to be recognized into a pre-trained recognition model for document image recognition, and obtain a recognition result corresponding to the document image to be recognized, where the recognition model is obtained by training a basic model with a preset proportion of special sample images in a preset number of basic sample images and a preset number of special sample images.

[0124] Please refer to Figure 5 , is a schematic structural diagram of a device 500 for training a recognition model provided in an embodiment of the present application. The device may include a memory 510 and a processor 520. Optionally, the device may further include: a communication interface 530 and a communication bus 540. The device corresponds to the above Figure 1 method embodiment and is capable of executing Figure 1 each step involved in the method embodiment. The specific functions of the device can be referred to the descriptions below.

[0125] Specifically, the memory 510 is configured to store computer-readable instructions.

[0126] The processor 520 is configured to process the readable instructions stored in the memory and is capable of executing Figure 2 each step of method embodiments 110 to 120.

[0127] The communication interface 530 is configured to communicate signaling or data with other node devices. For example: for communicating with a server or a terminal, or for communicating with other device nodes. The embodiments of the present application are not limited to this.

[0128] The communication bus 540 is configured to implement the direct connection communication between the above components.

[0129] Among them, the communication interface 530 of the device in the embodiment of the present application is used to communicate signaling or data with other node devices. The memory 510 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 510 can also be at least one storage device located far from the aforementioned processor. The memory 510 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 520, the electronic device executes the above Figure 1 shown method process. The processor 520 can be used on the device 300 and is used to execute the functions in the present application. Exemplarily, the above-mentioned processor 520 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The embodiments of the present application are not limited thereto.

[0130] Please refer to Figure 6 which is a schematic block diagram of a document image recognition device 600 provided in an embodiment of the present application. The device may include a memory 610 and a processor 620. Optionally, the device may further include: a communication interface 630 and a communication bus 640. The device corresponds to the above Figure 2 method embodiment and is capable of executing Figure 2 each step involved in the method embodiment. The specific functions of the device can be seen in the following description.

[0131] Specifically, the memory 610 is used to store computer-readable instructions.

[0132] The processor 620 is used to process the readable instructions stored in the memory and is capable of executing Figure 2 the method implementation step 210.

[0133] The communication interface 630 is used to communicate signaling or data with other node devices. For example: used for communication with a server or a terminal, or for communication with other device nodes. The embodiments of the present application are not limited thereto.

[0134] The communication bus 640 is used to implement the direct connection communication between the above components.

[0135] Among them, the communication interface 630 of the device in the embodiment of the present application is used for signaling or data communication with other node devices. The memory 610 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 610 can also be at least one storage device located far from the aforementioned processor. The memory 610 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 620, the electronic device executes the above Figure 2 shown method process. The processor 620 can be used on the device 400 and is used to execute the functions in the present application. Exemplarily, the above-mentioned processor 620 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The embodiments of the present application are not limited thereto.

[0136] The embodiment of the present application also provides a readable storage medium. When the computer program is executed by a processor, it executes the method process executed by the electronic device in the method embodiment as Figure 1 shown.

[0137] The embodiment of the present application also provides a readable storage medium. When the computer program is executed by a processor, it executes the method process executed by the electronic device in the method embodiment as Figure 2 shown.

[0138] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method, and will not be elaborated herein.

[0139] In summary, the embodiment of the present application provides a method, device, electronic device, and readable storage medium for training an identification model. The method includes obtaining a preset proportion of special sample pictures from a preset number of basic sample pictures and a preset number of special sample pictures corresponding to a preset type of document. The basic sample pictures are intercepted pictures of the corpus fields in the basic corpus, and the special sample pictures contain the same type of fields in the preset type of documents. Training a basic model based on a preset number of basic sample pictures and a preset proportion of special sample pictures among the preset number of special sample pictures to obtain an identification model for document picture recognition. On the premise of meeting the model accuracy rate, this method can achieve the effect of saving resources and training time for training the model.

[0140] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0141] In addition, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0142] If the described function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0143] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0144] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0145] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. A method for training an identification model, characterized in that Applied to document image recognition, including: Obtain a preset proportion of special sample images from the preset number of basic sample images and the preset number of special sample images corresponding to documents of a preset type, where the basic sample images are cropped images of corpus fields in a basic corpus, and the special sample images contain the same type of fields in the documents of the preset type; Train a basic model based on the preset number of basic sample images and the preset proportion of special sample images in the preset number of special sample images to obtain a recognition model for document image recognition. Before obtaining the preset proportion of special sample images from the preset number of basic sample images and the preset number of special sample images corresponding to documents of a preset type, the method further includes: reducing the number of basic sample images used to train the test model, and determining the minimum number of basic sample images during training of the test model when the recognition accuracy of the test model remains within a preset range, where the test model is used to test the number of sample images required by the model for recognizing the recognition model, and the preset range is set manually or according to the accuracy of training the model using historical image samples; using the minimum number of basic sample images as the preset number of basic sample images.

2. The method according to claim 1, wherein After training the basic model based on the preset number of basic sample images and the preset proportion of special sample images in the preset number of special sample images to obtain a recognition model for document image recognition, the method further includes: Input the document image to be recognized into the recognition model to obtain a recognition result, where the document image to be recognized is a document of a preset type or a document of a non-preset type; Calculate the recognition accuracy of each type of field in the document image to be recognized based on the recognition result; Based on the recognition accuracy of each type of field, increase the number of special sample images corresponding to each type of field when the recognition accuracy is lower than the preset range to obtain the increased special sample images; Use the increased special sample images to train the recognition model to obtain a document recognition model for recognizing the document image to be recognized.

3. The method according to claim 1 or 2, characterized in that Before obtaining the preset proportion of special sample images from the preset number of basic sample images and the preset number of special sample images corresponding to documents of a preset type, the method further includes: Reduce the number of special sample images used to train the test model, and determine the minimum number of special sample images during training of the test model when the recognition accuracy of the test model remains within a preset range; Use the minimum number of special sample images as the preset number of special sample images.

4. The method according to claim 3, wherein After using the minimum number of special sample images as the preset number of special sample images, the method further includes: Under the condition that the number of the preset number of special sample pictures remains unchanged and the recognition accuracy rate of the test model is within a preset range, adjust the proportion of each type of special sample picture in the preset number of special sample pictures, and determine the special sample pictures with the corresponding proportion when the recognition accuracy rate of the test model is the highest during the training of the test model; Use the special sample pictures with the corresponding proportion as the special sample pictures with a preset proportion in the preset number of special sample pictures.

5. A method for recognizing document images, characterized in that, Including: Obtain the document pictures to be recognized; Input the recognized document pictures into a pre-trained recognition model for document picture recognition to obtain the recognition result corresponding to the recognized document pictures. The recognition model is trained based on a preset number of basic sample pictures and a preset proportion of special sample pictures in the preset number of special sample pictures. Among them, reduce the number of basic sample pictures used for training the test model. Under the condition that the recognition accuracy rate of the test model remains within a preset range, determine the minimum number of basic sample pictures during the training of the test model. The test model is used to test the number of sample pictures required by the model for recognition. The preset range is set manually or according to the accuracy rate of training the model using historical picture samples; use the minimum number of basic sample pictures as the preset number of basic sample pictures.

6. An apparatus for training an identification model, characterized in that Including: An acquisition module, configured to acquire a preset number of basic sample pictures corresponding to a preset type of document and a preset proportion of special sample pictures in the preset number of special sample pictures. The basic sample pictures are pictures synthesized from corpus fields in a basic corpus, and the special sample pictures contain the same type of fields in the preset type of document; A training module, configured to train a basic model based on the preset number of basic sample pictures and a preset proportion of special sample pictures in the preset number of special sample pictures to obtain a recognition model for document picture recognition, Before the acquisition module acquires a preset number of basic sample pictures corresponding to a preset type of document and a preset proportion of special sample pictures in the preset number of special sample pictures, it is also used to: reduce the number of basic sample pictures used for training the test model. Under the condition that the recognition accuracy rate of the test model remains within a preset range, determine the minimum number of basic sample pictures during the training of the test model. The test model is used to test the number of sample pictures required by the model for recognition. The preset range is set manually or according to the accuracy rate of training the model using historical picture samples; use the minimum number of basic sample pictures as the preset number of basic sample pictures.

7. An apparatus for recognizing document images, characterized in that, Including: An acquisition module, configured to acquire the document pictures to be recognized; An identification module, configured to input the to-be-identified document image into a pre-trained identification model for document image identification to obtain an identification result corresponding to the to-be-identified document image, where the identification model is obtained by training a basic model with a preset proportion of special sample images among a preset number of basic sample images and a preset number of special sample images. Reducing the number of basic sample images for training the test model, and determining the minimum number of basic sample images during training the test model when the identification accuracy of the test model remains within a preset range, where the test model is used to test the number of sample images required by the model to be identified, and the preset range is set manually or according to the accuracy of training the model using historical image samples; using the minimum number of basic sample images as the preset number of basic sample images.

8. An electronic device, characterized in that, Comprising: A memory and a processor, where the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the method according to any one of claims 1-4 or 5 are run.

9. A computer-readable storage medium, characterized in that, Comprising: A computer program, when the computer program runs on a computer, causing the computer to execute the method according to any one of claims 1-4 or 5.

Citation Information

Patent Citations

  • Electronic device, bill information identification method and computer readable storage medium

    CN107766809A

  • Neural network training and image processing method and device, equipment and storage medium

    CN113792734A