Reinforcement learning method, device and equipment for certificate extraction large model and readable medium
Through the reinforcement learning method of document extraction, the reward function is constructed using the constraints of the fields to be extracted, which simplifies the training process, reduces resource consumption, improves the recognition accuracy and efficiency of document information extraction, and solves the problems of high OCR dependence and high training cost of multimodal large models in the existing technology.
Patent Information
- Application Number
- CN202510501687.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art relies on high OCR accuracy and poor general usage in document extraction, high training cost of multimodal large models and high resource consumption, and complex reinforcement learning methods, resulting in identification accuracy dependent on data quality and resource consumption.
The certificate extraction model is used for reinforcement learning, and the parameters are updated based on the constraints of the fields to be extracted, the sample image is omitted, and the training method is simplified.
It realizes efficient document information extraction, reduces the quality of data labeling and resource consumption, improves recognition accuracy and training efficiency, and simplifies the training process.
Smart Images

Figure CN120354970A_ABST
Abstract
Description
Background Art
[0002] Document extraction refers to extracting text information from document images, such as extracting text information like name, gender, document number, validity period, etc. from identity cards, driver's licenses, passports, business licenses and other documents, and outputting them in a structured and standardized manner.
[0003] Traditional document extraction relies on optical character recognition (OCR) technology to perform text recognition on document images, and then combines template matching, rule engines, or deep learning models, etc. for text processing to extract information. However, the foregoing solutions are highly dependent on the accuracy of OCR, and are also limited by templates, rules, or the annotation accuracy and data volume of model training data, resulting in poor versatility.
[0004] Currently, it is also possible to extract information from images using multi-modal large models. Without relying on OCR technology, it can improve versatility based on a large amount of pre-training, and also has a certain adaptability to changes in document types and templates. However, the training and adjustment of multi-modal large models rely on sample pairs of document images and extraction results, require high annotation accuracy, the recognition accuracy depends on data quality, and the commonly used reinforcement learning methods are complex and consume a large amount of training resources. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a reinforcement learning method for a document extraction large model, a reinforcement learning device for a document extraction large model, an electronic device, and a computer-readable medium. This method uses a large model for document extraction, performs reinforcement learning on unannotated sample document images, has a simple training method and low annotation cost, and the trained document extraction large model has high extraction accuracy for document images and does not depend on the accuracy of OCR technology.
[0006] According to a first aspect of the present disclosure, a reinforcement learning method for a document extraction large model is provided. The method may include: constructing an extraction instruction based on a sample document image and fields to be extracted; inputting the sample document image and the extraction instruction into a policy model to obtain a text extraction result output by the policy model; wherein, for one sample document image, there is a group of multiple text extraction results, and each text extraction result includes a field content extracted by the policy model for each field to be extracted; determining a reward value for each text extraction result using a reward function; the reward function is constructed based on the constraint conditions of the fields to be extracted; determining the relative advantage of each text extraction result within the group according to the reward value; and updating the parameters of the policy model based on the relative advantage to obtain a document extraction large model.
[0007] In an exemplary embodiment, the constraint conditions include at least one of content constraint conditions and output constraint conditions.
[0008] In an exemplary embodiment, the content constraint conditions include at least one of single-field value constraint conditions and multi-field association constraint conditions.
[0009] In an exemplary embodiment, the output constraint conditions include at least one of output format constraint conditions and output field constraint conditions.
[0010] In an exemplary embodiment, weights are set for the constraint conditions of each field to be extracted based on the importance level, and a reward function is used to determine the reward value of each text extraction result, including: using the reward function to determine the initial reward value of each text extraction result; and performing weighted processing on the initial reward value according to the weights corresponding to the constraint conditions to obtain the reward value.
[0011] In an exemplary embodiment, the sample document image includes multiple document types, and a reward function is used to determine the reward value of each text extraction result, including: determining the corresponding target reward function according to the document type corresponding to the text extraction result; and using the target reward function to determine the reward value of the corresponding text extraction result.
[0012] In an exemplary embodiment, parameter update of the policy model is performed based on relative advantage, including: performing parameter update of the policy model based on relative advantage, and regularizing with the information divergence between the reference model and the policy model; the reference model is the original policy model.
[0013] According to a second aspect of the present disclosure, there is provided a reinforcement learning device for a document extraction large model, which may include: an instruction construction module for constructing extraction instructions based on the sample document image and the fields to be extracted; a model processing module for inputting the sample document image and the extraction instructions into the policy model to obtain the text extraction results output by the policy model; wherein, on one sample document image, there is a set of multiple text extraction results, and each text extraction result includes a field content extracted by the policy model for each field to be extracted; a reward calculation module for using a reward function to determine the reward value of each text extraction result; the reward function is constructed based on the constraint conditions of the fields to be extracted; a group calculation module for determining the relative advantage of each text extraction result within the group according to the reward value; and a model update module for performing parameter update of the policy model based on relative advantage.
[0014] In an exemplary embodiment, the constraint conditions include at least one of content constraint conditions and output constraint conditions.
[0015] In an exemplary embodiment, the content constraint conditions include at least one of single-field value constraint conditions and multi-field association constraint conditions.
[0016] In an exemplary embodiment, the output constraint conditions include at least one of output format constraint conditions and output field constraint conditions.
[0017] In an exemplary embodiment, a weight is set for the constraint condition of each field to be extracted based on the importance level. The reward calculation module is specifically configured to determine an initial reward value for each text extraction result by using a reward function; and perform weighted processing on the initial reward value according to the weight corresponding to the constraint condition to obtain a reward value.
[0018] In an exemplary embodiment, the sample document image includes multiple document types. The reward calculation module is specifically configured to determine a corresponding target reward function according to the document type corresponding to the text extraction result; and use the target reward function to determine the reward value of the corresponding text extraction result.
[0019] In an exemplary embodiment, the model update module is specifically configured to update the parameters of the policy model based on the relative advantage and regularize with the information divergence between the reference model and the policy model; the reference model is the original policy model.
[0020] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0021] a processor; and
[0022] a memory for storing a computer program of the processor;
[0023] wherein the processor is configured to implement the reinforcement learning method for the document extraction large model as described in the first aspect by executing the computer program.
[0024] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the reinforcement learning method for the document extraction large model as described in the first aspect is implemented.
[0025] According to a fifth aspect of the present disclosure, there is provided a computer program product, which when running on an electronic device, causes the electronic device to implement the reinforcement learning method for the document extraction large model as described in the first aspect when executed.
[0026] The present disclosure provides a reinforcement learning method for a large model of document extraction, a reinforcement learning device for a large model of document extraction, an electronic device, and a computer-readable medium. In this method, during reinforcement learning, an extraction instruction can be constructed based on a sample document image and fields to be extracted and input into a policy model to obtain a text extraction result output by the policy model. Among them, for one sample document image, there is a set of multiple text extraction results, and each text extraction result includes a type of field content extracted by the policy model for all fields to be extracted respectively. On this basis, a reward function constructed based on the constraint conditions of the fields to be extracted is used to determine the reward value of each text extraction result, and the relative advantage of each text extraction result within the group is determined according to the reward value, and the parameters of the policy model are updated based on this relative advantage to obtain a large model for document extraction. This method uses a large model for document extraction and performs reinforcement learning on unlabeled sample document images, omitting the labeling cost of samples. The large model for document extraction obtained by training not only realizes end-to-end efficient recognition, but also the recognition accuracy does not depend on the quality of data labeling. It is simpler than common reinforcement learning methods and also reduces the consumption of training resources.
[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0029] Figure 1 It is one of the flowcharts of the steps of the reinforcement learning method for the large model of document extraction provided by the embodiment of the present disclosure.
[0030] Figure 2 It is the second flowchart of the steps of the reinforcement learning method for the large model of document extraction provided by the embodiment of the present disclosure.
[0031] Figure 3 It is a schematic diagram of the implementation architecture of the reinforcement learning of the large model of document extraction provided by the embodiment of the present disclosure.
[0032] Figure 4 It is a block diagram of the structure of the reinforcement learning device for the large model of document extraction provided by the embodiment of the present disclosure.
[0033] Figure 5 It is a schematic diagram of the structure of an electronic device provided by the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0035] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0036] It should be noted that all the data obtained in the present disclosure are accessed, collected, stored, and applied to subsequent analysis and processing after clearly informing the user or the relevant data owner of information such as the content of data collection, data usage, processing methods, etc., and with the consent and authorization of the user or the relevant data owner. Moreover, the means for the user or the relevant data owner to access, correct, and delete the data, as well as the methods for revoking consent and authorization, can be provided.
[0037] Large Multimodal Models (LMMs) combine visual and language understanding capabilities and have certain performance in multimodal tasks that simultaneously process image and text information. In current applications, due to their end-to-end information extraction capabilities, they can directly extract information from images, have strong cross-language capabilities learned through large-scale pre-training, and relatively stable recognition accuracy, which can reduce the impact of general typesetting changes and image quality changes in document information extraction. However, the recognition accuracy of LMMs is restricted by training costs and depends on high-quality, large-scale multimodal image-text pair data, resulting in high data costs and large resource consumption for training. Reducing training costs may lead to a decrease in recognition accuracy and poor robustness of the large model, making it necessary to further manually proofread the recognition results and resulting in low task efficiency. LMMs often use the RLHF (Reinforcement Learning from Human Feedback) method for training, which requires additional training of a reward model, making the training method complex and further increasing the consumption of training resources.
[0038] Embodiments of the present disclosure provide a reinforcement learning method for a document extraction large model. A reward function is constructed using the constraint conditions of the fields to be extracted in the document image, and the parameters of the fields to be extracted by the policy model are updated based on the relative advantages of multiple text extraction results for each group on unlabeled sample document images, thereby realizing the reinforcement learning of the policy model and obtaining a document extraction large model. This method effectively reduces the dependence of model training performance on the scale of sample data and the quality of annotations, simplifies the training method, and saves training resources. The specific description is as follows:
[0039] Figure 1 One of the flowcharts of the steps of the reinforcement learning method for the document extraction large model provided by the embodiments of the present disclosure. As Figure 1 shown, the method may include the following steps 101 to 105.
[0040] In step 101, an extraction instruction is constructed based on the sample document image and the fields to be extracted.
[0041] In the embodiments of the present disclosure, the sample document image may include real document images collected on the basis of authorization, or may include synthetic document images obtained by data augmentation based on real document images. The fields to be extracted are information indicating the extraction content on the document image. For example, taking the sample document image as Li Si's ID card, the fields to be extracted may be "name" indicating the extraction of the corresponding information "Li Si" for "name", "gender" indicating the extraction of the corresponding information "male" for "gender", and the fields to be extracted may also be "ethnic group", "date of birth", "ID number", "valid date", and so on.
[0042] In the embodiments of the present disclosure, the extraction instruction is used to prompt the large model to extract the field content corresponding to the field to be extracted from the input document image. The extraction instruction can indicate to the large model to perform extraction on the sample document image, can prompt the fields to be extracted, and can also further indicate the output method. For example, when indicating the sample document image, the document type, region, issuing agency, etc. to which the sample document image belongs can be indicated; when indicating the fields to be extracted, the fields to be extracted can be provided in the form of a list or a set; when indicating the output method, it can include the output content, output format, etc. On this basis, in order to further assist the large model in understanding the reasoning task, the role to be played can also be indicated to the large model, such as "You are a professional information reviewer", "You are an identity information verification personnel", "You are responsible for reviewing the validity of document information", etc., and the specific reasoning task can be described, such as "You need to extract the field content corresponding to the field to be extracted from the given document image", "You need to extract the corresponding information from the document image according to the prompt of the field to be extracted", etc.
[0043] In step 102, the sample document image and the extraction instruction are input into the policy model, and the text extraction result output by the policy model is obtained; among them, for one sample document image, there is a set of multiple text extraction results corresponding to it, and each text extraction result includes a kind of field content extracted by the policy model for all fields to be extracted respectively.
[0044] In the embodiments of the present disclosure, the policy model is a multi-modal large model with the ability to extract text information from images in reinforcement learning. The policy model can select the weights of an open-source multi-modal large model as the basis for reinforcement learning. When selecting the policy model, its model performance and operating cost can be comprehensively considered. For example, an open-source multi-modal large model with a pre-training scenario including multi-modal information extraction and having a certain basic information extraction ability can be selected. For specific document image extraction requirements, when the performance of the open-source multi-modal large model is lower than the basic expectation, cold start on synthetic document images can also be considered.
[0045] On this basis, the sample document image and the constructed extraction instruction can be input into the policy model to prompt the policy model to extract the text extraction result corresponding to the field to be extracted from the sample document image through the extraction instruction. On this basis, a set of multiple text extraction results output by the policy model on one sample document image can be obtained. Among this set of multiple text extraction results, each text extraction result can include a kind of field content extracted for all fields to be extracted respectively, and can be obtained by the policy model sampling the same sample document image multiple times.
[0046] In step 103, a reward function is used to determine the reward value of each text extraction result; the reward function is constructed based on the constraint conditions of the fields to be extracted.
[0047] In the embodiments of the present disclosure, the constraint conditions of the fields to be extracted are used to constrain the field content filled and output corresponding to the fields to be extracted. The constraint conditions for filling the field content of different fields to be filled may be different, such as different constraints on the character type and the number of characters of the field content; there may be associated constraints on the field content filled between different fields to be filled, such as the consistency and continuity of the field content of different fields to be extracted, etc. Based on the constraint conditions of the fields to be extracted, a corresponding reward function can be constructed to make the reward value corresponding to the text extraction result that meets the constraint conditions higher, so as to reward the text extraction result that meets the constraint conditions and punish the text extraction result that does not meet the constraint conditions. Those skilled in the art can automatically judge whether the text extraction result meets the constraint conditions of the fields to be extracted through code programming, and generate the reward value corresponding to the text extraction result based on the judgment result. The reward values for different constraint conditions can be adjusted according to the requirements of the fields to be extracted. When the importance level of the fields to be extracted is high and the accuracy requirement is high, the reward value can be made relatively higher; on the contrary, when the importance level is low and the accuracy requirement is low, the reward value can be made relatively lower.
[0048] In step 104, the relative advantage of each text extraction result within the group is determined according to the reward value.
[0049] In the embodiments of the present disclosure, the relative advantage may be the advantage of each text extraction result within the group relative to multiple text extraction results. This relative advantage establishes a relative baseline for output advantage evaluation based on a group of multiple text extraction results. The higher the relative advantage, the higher the output quality of the text extraction result relative to multiple text extraction results within the group; on the contrary, the lower the output quality. Among them, the reward value of each text extraction result can be compared with the average value of the reward values of multiple text extraction results to determine the relative advantage of each text extraction result within the group; or normalization processing based on the average value and standard deviation can also be used to determine the relative advantage of each text extraction result within the group. The embodiments of the present disclosure do not make specific limitations on this.
[0050] In step 105, the parameter of the policy model is updated based on the relative advantage to obtain the large model for certificate extraction.
[0051] On this basis, the parameter of the policy model can be updated based on the relative advantage. According to the task requirements, the policy model can be updated towards the text extraction result with a higher relative advantage. After the parameter is updated, when the performance of the policy model or the number of times of parameter update reaches the requirements of reinforcement learning training, it can be considered that the policy model converges, thereby obtaining the large model for certificate extraction. This large model for certificate extraction reduces the annotation cost and resource consumption of the reinforcement learning of the multi-modal large model, and improves the training efficiency on the basis of ensuring the recognition accuracy.
[0052] In an alternative method embodiment of the present disclosure, the constraint conditions include at least one of content constraint conditions and output constraint conditions.
[0053] In the embodiments of the present disclosure, the constraint conditions for the fields to be extracted may include content constraint conditions for constraining from the field content, or may include output constraint conditions for constraining from the perspective of the text extraction result output by the policy model.
[0054] In an alternative method embodiment of the present disclosure, the content constraint conditions include at least one of single-field value constraint conditions and multi-field association constraint conditions.
[0055] In the embodiments of the present disclosure, when the constraint condition is a content constraint condition, it may be a single-field value constraint condition constructed based on the internal constraint of the field content for the field to be extracted, or may be a multi-field association constraint condition constructed based on the association constraint of the field content between different fields to be extracted.
[0056] The single-field value constraint condition, for example, may be whether the field content of a single field to be extracted conforms to the requirements of a regular expression, whether it conforms to the character type requirements, whether it conforms to the character digit requirements, whether it conforms to the enumerated option requirements, etc. Exemplarily, the single-field value constraint condition may include whether the ID number is a pure number or a combination of numbers and letters, whether the ID number meets the requirement of 18 digits, whether the last digit verification code of the ID number conforms to the previous digit operation rule, whether the gender enumerated option conforms to "male" or "female", whether the date is a parsable valid date format of year, month, and day, etc.
[0057] The multi-field association constraint condition, for example, may be a constraint on the consistency of the field content of multiple fields to be extracted, such as the consistency association of the field content of the ID number, date of birth, place of birth, gender, etc. to be extracted; or, it may also be a continuity constraint, such as the chronological sequence of the field content of the date of birth, document issuance date, document expiration date, etc. to be extracted.
[0058] In an alternative method embodiment of the present disclosure, the output constraint conditions include at least one of output format constraint conditions and output field constraint conditions.
[0059] In the embodiments of the present disclosure, when the constraint condition is an output constraint condition, it may be an output format constraint condition for the policy model to output the text extraction result after extracting the field content of the field to be extracted, or may be an output field constraint condition such as field name, field quantity, field order, etc. in the text extraction result output by the policy model after extracting the field content of the field to be extracted.
[0060] Output format constraint conditions, for example, they can be constraint conditions on the format of the text extraction results output by the policy model. Exemplarily, the output result can be in the form of key-value pairs or in the format of a JSON string {field name: field value} to meet subsequent parsing requirements.
[0061] Output field constraint conditions, for example, they can be constraint conditions on the number of fields, field names, field order, etc. included in the text extraction results output by the policy model to avoid the situation of missing fields.
[0062] In an optional method embodiment of the present disclosure, weights are set for the constraint conditions of each field to be extracted based on the importance level. The foregoing step 103 may include the following steps A1 to step A2.
[0063] In step A1, a reward function is used to determine the initial reward value of each text extraction result.
[0064] In step A2, the initial reward value is weighted according to the weight corresponding to the constraint condition to obtain the reward value.
[0065] In the embodiments of the present disclosure, weights can also be set for the constraint conditions of the fields to be extracted based on the importance level according to the reinforcement learning requirements of the document extraction model. Exemplarily, taking the constraint conditions including content constraint conditions and output constraint conditions, the content constraint conditions including single-field value constraint conditions and multi-field association constraint conditions, and the output constraint conditions including output format constraint conditions and output field constraint conditions as an example, the output format condition can be set as the first weight, the multi-field association constraint condition can be set as the second weight, the single-field value constraint condition corresponding to the key field to be extracted can be set as the third weight, and the single-field value constraint condition corresponding to other fields to be extracted can be set as the fourth weight, and the first weight, the second weight, the third weight to the fourth weight decrease in turn. Thus, when calculating the reward value of the text extraction result, after determining the initial reward value based on the field content extracted for each field to be extracted, the initial reward value is weighted respectively for each field to be extracted according to the weight corresponding to the constraint condition to obtain the final reward value of the text extraction result.
[0066] For example, the text extraction result is {field name 1: field value 1, field name 2: field value 2}. Based on the constraint condition 1 corresponding to field name 1, the field value 1 is judged to determine the initial reward value 1 corresponding to field name 1; and, based on the constraint condition 2 corresponding to field name 2, the field value 2 is judged to determine the initial reward value 2 corresponding to field name 2, and the initial reward value of the text extraction result is determined by the initial reward value 1 and the initial reward value 2.
[0067] Further, Constraint Condition 1 corresponds to a third weight, Constraint Condition 2 corresponds to a fourth weight, and the third weight is greater than the fourth weight. Then, the initial reward value is weighted based on the third weight and the fourth weight to obtain a reward value that pays more attention to the initial reward value 1.
[0068] In an optional method embodiment of the present disclosure, the sample document image includes multiple document types, and the foregoing step 103 may include the following steps B1 to B2.
[0069] In step B1, according to the document type corresponding to the text extraction result, the corresponding target reward function is determined.
[0070] In step B2, the target reward function is used to determine the reward value corresponding to the text extraction result.
[0071] In the embodiments of the present disclosure, the document type may correspond to a field name, the number of fields, a layout distribution, etc. It may describe different types of documents, such as different document types like ID cards, driver's licenses, passports, etc., or different versions of the same document, such as passports in different countries, languages, or different issuance versions. When the sample document image includes multiple document types, the field names, the number of fields, the layout distribution, etc. of the fields to be extracted in the sample document images of different document types may be different, so the constraint conditions for the fields to be extracted may also be different. On this basis, for different document types, a set of one or more reward functions can be constructed according to one or more corresponding constraint conditions. Thus, when determining the reward value, the corresponding set of reward functions can be flexibly called according to the document type of the input sample document image, so as to meet the reinforcement learning requirements for multiple document types.
[0072] Specifically, among the different reward functions corresponding to multiple document types, the target reward function can be determined according to the document type corresponding to the text extraction result output by the policy model. The target reward function may include one or more reward functions according to the specific situation of the document type. Among them, determining the reward value corresponding to the text extraction result using the target reward function can refer to the relevant descriptions of the foregoing step 103 or steps A1 to A2. To avoid repetition, it will not be elaborated here.
[0073] Figure 2 This is the second step flowchart of the reinforcement learning method for the document extraction large model provided by the embodiments of the present disclosure. As Figure 2 shown, the method may include the following steps 201 to 205.
[0074] In step 201, an extraction instruction is constructed based on the sample document image and the fields to be extracted.
[0075] In the embodiments of the present disclosure, step 201 may refer to the relevant description of the foregoing step 101. To avoid repetition, it will not be elaborated here.
[0076] For example, based on the sample document image 1 of Country A and the fields to be extracted [name, gender, age, place of birth], the following extraction instructions can be constructed:
[0077] "You are a professional information reviewer and need to extract the corresponding field content from the given sample document image 1. The following is a [sample document] <image> from [Country A], and you need to extract the field content corresponding to the specified fields. The text extraction result should be output in JSON format, ensuring that the JSON string is parsable. The fields to be extracted are [field list (name, gender, age, place of birth)]".
[0078] In step 202, the sample document image and the extraction instructions are input into the policy model to obtain the text extraction result output by the policy model; among them, for one sample document image, there is a set of multiple text extraction results, and each text extraction result includes a kind of field content extracted by the policy model for all fields to be extracted respectively.
[0079] In the embodiment of the present disclosure, step 202 can refer to the relevant description of the foregoing step 102 correspondingly. To avoid repetition, it will not be elaborated here.
[0080] For example, based on the above extraction instructions, the multimodal policy model can convert the input sample document image 1 into a feature vector and splice it to the <image> placeholder position of the foregoing extraction instructions for subsequent reasoning.
[0081] The policy model extracts the field content corresponding to the fields to be extracted [name, gender, age, place of birth] on the sample document image 1, and can sample the sample document image 1 four times and output a set of text extraction results as follows:
[0082] [Li Si, male, 20, City A];
[0083] [Li Si, morning, 20, City A];
[0084] [Li Si, male, 2O, City A];
[0085] [Li Si, morning, 2O, City A].
[0086] In step 203, a reward function is used to determine the reward value of each text extraction result; the reward function is constructed based on the constraint conditions of the fields to be extracted.
[0087] In the embodiment of the present disclosure, step 203 can refer to the relevant description of the foregoing step 103 correspondingly. To avoid repetition, it will not be elaborated here.
[0088] For example, taking the text extraction result [Li Si, male, 20, City A] as an example, four reward functions can be used to make judgments respectively. Among them, the field content "Li Si" extracted from the field "Name" conforms to the constraint condition of its character type "Chinese character", the field content "male" extracted from the field "Gender" conforms to the constraint condition of its enumeration options "male, female", the field content "20" extracted from the field "Age" conforms to the constraint condition of its character type "numeric character", and the field content "City A" extracted from the field "Place of Birth" conforms to the constraint condition of its enumeration option "administrative region". Therefore, the reward value of this text extraction result can be 4.
[0089] Similarly, taking the text extraction result [Li Si, early, 20, City A] as an example, it can be seen that the field contents extracted from the fields "Name", "Age" and "Place of Birth" conform to the corresponding constraint conditions, but the field content "early" extracted from the field "Gender" does not conform to the constraint condition of its enumeration option. Therefore, the reward value of this text extraction result can be 3. And so on, the reward value of the text extraction result [Li Si, male, 20, City A] can be 3, and the reward value of the text extraction result [Li Si, early, 20, City A] can be 2.
[0090] In step 204, the relative advantage of each text extraction result within the group is determined according to the reward value.
[0091] In the embodiments of the present disclosure, step 204 can refer to the relevant description of the foregoing step 104. To avoid repetition, it will not be elaborated here.
[0092] For example, the reward values of the foregoing group of four text extraction results are [4, 3, 3, 2]. Based on this, normalization is performed, and its mean value is 3, and the standard deviation is The relative advantage after normalization is
[0093]
[0094] In step 205, the parameters of the policy model are updated based on the relative advantage, and regularization is performed with the information divergence between the reference model and the policy model to obtain the large document extraction model; the reference model is the original policy model.
[0095] In the embodiments of the present disclosure, the parameter update based on the relative advantage in step 205 to drive the iteration of the policy model can refer to the relevant description of the foregoing step 105. To avoid repetition, it will not be elaborated here.
[0096] On this basis, the reference model is the original policy model, that is, the policy model without reinforcement learning. The information divergence can also be called the KL divergence (Kullback-Leibler divergence), relative entropy, and is used to measure the difference degree between two probability distributions. In the embodiments of the present disclosure, the information divergence between the reference model and the policy model is regularized during parameter update to control the difference degree between the output probability distribution of the policy model and the output probability distribution of the reference model, reduce the probability of overfitting, stabilize the reinforcement learning training of the policy model, and prevent it from deviating excessively. Those skilled in the art can dynamically adjust the coefficient of the information divergence according to the performance of the reference model.
[0097] Figure 3 It is a schematic diagram of the reinforcement learning implementation architecture of the document extraction large model in the embodiments of the present disclosure. As Figure 3 shown, the architecture includes a policy model, a reference model, and a reward function. The reward function is constructed based on the constraint conditions of the fields to be extracted for the field content. In this reinforcement learning architecture, online learning can be performed, and the document extraction large model can be continuously updated in combination with the guidance of the reward function.
[0098] In the implementation process of this learning structure, after inputting the sample document image and the corresponding extraction instruction into the policy model, the policy model can perform multiple samplings on the same sample document image to obtain a set of text extraction results; determine a set of reward values corresponding to the set of text extraction results through the reward function based on the constraint conditions; on this basis, perform a group calculation operation, normalize the reward values within the group to convert them into relative advantages, and use the relative advantages to drive the parameter update of the policy model. At the same time, the information divergence between the reference model and the policy model is regularized to prevent the policy model from deviating excessively.
[0099] Among them, the reward function can include overall constraints in multiple aspects such as single-field value taking of the fields to be extracted, multi-field association, and field content output. Using relative advantages as the driving signal for parameter update avoids the annotation cost of the field content in the sample document image; when performing reinforcement learning based on multiple document types, different reward functions can also be dynamically adapted to flexibly support a wider range of application scenarios. When expanding the document types of the sample document images, only the corresponding reward functions need to be supplemented, without complicated annotation work.
[0100] The reinforcement learning method for the document extraction large model provided by the present disclosure can construct an extraction instruction based on a sample document image and a field to be extracted in reinforcement learning and input it into a policy model to obtain a text extraction result output by the policy model; wherein, a group of multiple text extraction results corresponds to one sample document image; on this basis, a reward function constructed based on the constraint conditions of the field to be extracted is used to determine the reward value of each text extraction result, and the relative advantage of each text extraction result within the group is determined according to the reward value, and the policy model is updated with parameters based on this relative advantage to obtain the document extraction large model. This method uses a large model for document extraction and performs reinforcement learning on unannotated sample document images, omitting the annotation cost of samples. The trained document extraction large model not only realizes end-to-end efficient recognition, but also the recognition accuracy does not depend on the quality of data annotation. It is simpler than the commonly used reinforcement learning method and also reduces the consumption of training resources.
[0101] Figure 4 Also provided is a structural block diagram of a reinforcement learning device 400 for a document extraction large model. As Figure 4 shown, the device 400 may include: an instruction construction module 401 for constructing an extraction instruction based on a sample document image and a field to be extracted; a model processing module 402 for inputting the sample document image and the extraction instruction into a policy model to obtain a text extraction result output by the policy model; wherein, a group of multiple text extraction results corresponds to one sample document image, and each text extraction result includes a kind of field content extracted by the policy model for all fields to be extracted respectively; a reward calculation module 403 for determining the reward value of each text extraction result by using a reward function; the reward function is constructed based on the constraint conditions of the field to be extracted; a group calculation module 404 for determining the relative advantage of each text extraction result within the group according to the reward value; a model update module 405 for updating the parameters of the policy model based on the relative advantage.
[0102] In an exemplary embodiment, the constraint conditions include at least one of content constraint conditions and output constraint conditions.
[0103] In an exemplary embodiment, the content constraint conditions include at least one of single-field value-taking constraint conditions and multi-field association constraint conditions.
[0104] In an exemplary embodiment, the output constraint conditions include at least one of output format constraint conditions and output field constraint conditions.
[0105] In an exemplary embodiment, weights are set for the constraint conditions of each field to be extracted based on the importance degree. The reward calculation module 403 is specifically configured to determine the initial reward value of each text extraction result by using a reward function; and perform weighted processing on the initial reward value according to the weights corresponding to the constraint conditions to obtain the reward value.
[0106] In an exemplary embodiment, the sample document image includes multiple document types. The reward calculation module 403 is specifically configured to determine a corresponding target reward function according to the document type corresponding to the text extraction result, and use the target reward function to determine the reward value of the corresponding text extraction result.
[0107] In an exemplary embodiment, the model update module 405 is specifically configured to update the parameters of the policy model based on the relative advantage, and regularize with the information divergence between the reference model and the policy model; the reference model is the original policy model.
[0108] The reinforcement learning device of the large model for document extraction provided by the present disclosure can construct an extraction instruction based on the sample document image and the fields to be extracted in reinforcement learning and input it into the policy model to obtain the text extraction result output by the policy model. Among them, a group of multiple text extraction results corresponds to one sample document image. On this basis, the reward function constructed based on the constraint conditions of the fields to be extracted is used to determine the reward value of each text extraction result, and the relative advantage of each text extraction result in the group is determined according to the reward value, and the parameters of the policy model are updated based on the relative advantage to obtain the large model for document extraction. This method uses a large model for document extraction and performs reinforcement learning on unlabeled sample document images, omitting the annotation cost of the samples. The trained large model for document extraction not only realizes end-to-end efficient recognition, but also the recognition accuracy does not depend on the quality of data annotation. It is simpler than the commonly used reinforcement learning method and also reduces the consumption of training resources.
[0109] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units for embodiment.
[0110] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0111] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0112] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is further provided.
[0113] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a device, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.
[0114] Next, refer to Figure 5 to describe the electronic device 500 according to this embodiment of the present disclosure. Figure 5 The electronic device 500 shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0115] As Figure 5 shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one of the above processing units 510, at least one of the above storage units 520, and a bus 530 connecting different system components (including the storage unit 520 and the processing unit 510).
[0116] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0117] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 5201 and / or a cache storage unit 5202, and may further include a read-only storage unit (ROM) 5203.
[0118] The storage unit 520 may also include a program / utility 5204 having a set (at least one) of program modules 5205. Such program modules 5205 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0119] The bus 530 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0120] The electronic device 500 may also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 500, and / or may communicate with any device that enables the electronic device 500 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the display unit 540 and an input / output (I / O) interface 550 connected to the display unit 540. Also, the electronic device 500 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 560. As shown in the figure, the network adapter 560 communicates with other modules of the electronic device 500 through the bus 530. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0121] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0122] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above methods of this specification is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0123] In an embodiment of the present disclosure, there is also provided a program product for implementing the above method. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited to this. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0124] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0125] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0126] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0127] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partially on the user's device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0128] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.
[0129] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A reinforcement learning method for a large model of document extraction, characterized in that, The method includes: Constructing an extraction instruction based on a sample document image and fields to be extracted; Inputting the sample document image and the extraction instruction into a policy model to obtain a text extraction result output by the policy model; wherein, for one sample document image, a group of multiple text extraction results corresponds, and each text extraction result includes a field content extracted by the policy model for each of all the fields to be extracted; Determining a reward value for each text extraction result by using a reward function; the reward function is constructed based on the constraint conditions of the fields to be extracted; Determining the relative advantage of each text extraction result within the group according to the reward value; Updating the parameters of the policy model based on the relative advantage to obtain a large document extraction model.
2. The method according to claim 1, wherein The constraint conditions include at least one of content constraint conditions and output constraint conditions.
3. The method according to claim 2, characterized in that, The content constraint conditions include at least one of single-field value-taking constraint conditions and multi-field association constraint conditions.
4. The method according to claim 2, characterized in that, The output constraint conditions include at least one of output format constraint conditions and output field constraint conditions.
5. The method according to claim 1, wherein Weights are set for the constraint conditions of each field to be extracted based on the importance degree. The step of determining a reward value for each text extraction result by using a reward function includes: Determining an initial reward value for each text extraction result by using the reward function; Performing weighted processing on the initial reward value according to the weight corresponding to the constraint condition to obtain the reward value.
6. The method according to claim 1, wherein The sample document image includes multiple document types. The step of determining a reward value for each text extraction result by using a reward function includes: Determining a corresponding target reward function according to the document type corresponding to the text extraction result; Determining the reward value corresponding to the text extraction result by using the target reward function.
7. The method according to claim 1, wherein The step of updating the parameters of the policy model based on the relative advantage includes: Updating the parameters of the policy model based on the relative advantage and performing regularization with the information divergence between a reference model and the policy model; the reference model is the original policy model.
8. An enhanced learning device for a large model of document extraction, characterized in that, The device includes: An instruction construction module for constructing an extraction instruction based on a sample document image and fields to be extracted; A model processing module for inputting the sample document image and the extraction instruction into a policy model to obtain a text extraction result output by the policy model; wherein, for one sample document image, a group of multiple text extraction results corresponds, and each text extraction result includes a field content extracted by the policy model for each of all the fields to be extracted; A reward calculation module for determining a reward value for each text extraction result by using a reward function; the reward function is constructed based on the constraint conditions of the fields to be extracted; A group calculation module for determining the relative advantage of each text extraction result within the group according to the reward value; A model update module for updating the parameters of the policy model based on the relative advantage.
9. An electronic device, characterized in that, It includes: A processor; and A memory for storing a computer program of the processor; Among them, the processor is configured to execute the reinforcement learning method of the document extraction large model according to any one of claims 1 to 7 by executing a computer program.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the reinforcement learning method of the document extraction large model according to any one of claims 1 to 7.