Method and device for optimizing document identification large model, equipment, medium and program
By optimizing the training parameters and layout templates of multimodal large models, the problem of weak recognition ability when the model faces unseen document types or complex typesettings, achieving higher recognition accuracy and stability.
Patent Information
- Application Number
- CN202510150165.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
AI Technical Summary
When faced with unseen document types or complex typesetting, existing multimodal large models have weak recognition capabilities and low recognition accuracy, mainly due to poor training data quality.
By obtaining multiple training parameters to be optimized, including linear and nonlinear parameters, iteratively generate corresponding typesetting templates, and determining the training templates to optimize the general big model so that it can be adjusted for different types of document typesetting characteristics.
It improves the model's adaptability to various document structures and formats, enhances the accuracy and stability of the model in document recognition tasks, and reduces the recognition error rate.
Smart Images

Figure CN119992572A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, applicable to financial technology scenarios, and in particular to an optimization method, device, equipment, medium and program for a large document recognition model. Background Art
[0002] Currently, the training of multimodal large models mainly relies on large-scale annotated data. However, due to the diversity of document types and the complexity of typesetting, it is quite difficult to obtain enough high-quality annotated data. To meet this challenge, existing data augmentation techniques generate additional training data through operations such as rotation, scaling, and cropping to increase the diversity of data.
[0003] However, these existing data augmentation methods still have certain limitations. Although they can generate some additional training data, these data often lack sufficient diversity and cannot fully cover the complex layouts and document types that may appear in reality. This results in the trained multimodal large models having limited recognition capabilities when faced with unseen document types or layouts. Summary of the invention
[0004] Based on this, the present invention provides a method, device, equipment, medium and program for optimizing a large document recognition model to solve the problem of weak recognition ability and low recognition accuracy of a large multimodal model used for document recognition due to poor quality of training data.
[0005] In a first aspect, an embodiment of the present invention provides a method for optimizing a large document recognition model, the method comprising:
[0006] Obtaining a plurality of training parameters to be optimized for optimizing a general large model, wherein the training parameters include nonlinear parameters and linear parameters;
[0007] Iteratively generate first-class typesetting templates corresponding to each linear parameter through linear parameter adjustment rules, and iteratively determine first-class training templates corresponding to each linear parameter by inputting the first-class typesetting templates into the general large model for document recognition;
[0008] By means of nonlinear parameter adjustment rules, all second-class typesetting templates corresponding to each nonlinear parameter are generated, and by inputting all second-class typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition, each second-class training template corresponding to each nonlinear parameter is determined;
[0009] The general large model is optimized using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
[0010] In a second aspect, an embodiment of the present invention provides an optimization device for a large document recognition model, the device comprising:
[0011] A training parameter acquisition module, used to acquire a plurality of training parameters to be optimized for optimizing the general large model, wherein the training parameters include nonlinear parameters and linear parameters;
[0012] A first-class training template determination module, for iteratively generating a first-class typesetting template corresponding to each linear parameter by adjusting a linear parameter rule, and iteratively determining each first-class training template corresponding to each linear parameter by inputting the first-class typesetting template into a universal large model for document recognition;
[0013] A second-category training template determination module is used to generate all second-category typesetting templates corresponding to each nonlinear parameter by adjusting the nonlinear parameter rules, and to determine each second-category training template corresponding to each nonlinear parameter by inputting all second-category typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition;
[0014] The document recognition large model optimization module is used to optimize the general large model using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
[0015] In a third aspect, an embodiment of the present invention provides an electronic device, the electronic device comprising:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a method for optimizing a large document recognition model as described in any embodiment of the present invention.
[0019] In a fourth aspect, a computer-readable storage medium is also provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a method for optimizing a large document recognition model as described in any embodiment of the present invention when executed.
[0020] In a fifth aspect, a computer program product is also provided, the computer program product comprising a computer program, and the computer program, when executed by a processor, implements a method for optimizing a large document recognition model as described in any embodiment of the present invention.
[0021] The technical solution of the embodiment of the present invention describes a method for optimizing a large document recognition model, which obtains multiple training parameters to be optimized (including nonlinear and linear parameters), and generates corresponding typesetting templates based on linear and nonlinear parameter adjustment rules, and then determines the training template to optimize the general large model, so that the model can be adjusted according to the typesetting characteristics of different types of documents, thereby improving the model's adaptability to various document structures and formats. Generating training templates based on specific adjustment rules and optimizing the model using these templates helps the model better learn the features and patterns in the document, so that it can more accurately identify and parse the document content in the document recognition task, and reduce the recognition error rate.
[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 is a flowchart of a method for optimizing a large document recognition model provided according to Embodiment 1 of the present invention;
[0025] Figure 2 is a schematic diagram of the structure of an optimization device for a large document recognition model provided according to Embodiment 2 of the present invention;
[0026] Figure 3 It is a structural schematic diagram of an electronic device according to a method for optimizing a large document recognition model provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] Embodiment 1
[0030] Figure 1 This is a flowchart of a method for optimizing a large document recognition model provided in the first embodiment of the present invention. This embodiment is applicable to the case of generating a training template for training a large multimodal model. The method can be executed by an optimization device for a large document recognition model. The device can be implemented in the form of hardware and / or software. The device can be configured in a document management system in the financial field. Figure 1 As shown, the method includes:
[0031] S110, obtaining a plurality of training parameters to be optimized for optimizing the general large model, wherein the training parameters include nonlinear parameters and linear parameters.
[0032] The above content shows that in the process of optimizing the general large model, it is first necessary to determine which training parameters need to be optimized. These training parameters to be optimized are divided into two categories: non-linear parameters and linear parameters. In an embodiment of the present invention, the training data used to optimize the document recognition large model is automatically generated based on a predefined intelligent typesetting model, and the intelligent typesetting model is constructed based on many parameters related to document typesetting in the actual application process. The typesetting parameters used to determine the document layout are called linear parameters, and the typesetting parameters used to determine the document style are called non-linear parameters. The above linear parameters and non-linear parameters are collectively referred to as training parameters to be optimized.
[0033] Optionally, obtaining multiple training parameters to be optimized contained in a training template for optimizing a general large model may include:
[0034] Obtaining a predefined initial typesetting model, identifying a plurality of training parameters to be optimized in the initial training template, and obtaining an initial parameter value of each of the training parameters to be optimized;
[0035] According to a preset classification standard, each of the training parameters to be optimized is classified into a nonlinear parameter or a linear parameter;
[0036] An initial training template corresponding to the initial typesetting model is generated, and the initial training template is input into the general large model to obtain a general recognition error rate of the general large model for the initial training template.
[0037] The initial typesetting model is also the intelligent typesetting model mentioned in the above-mentioned embodiment of the invention, and the initial parameter values in the initial typesetting model can be specified based on prior results, and the classification criteria are divided according to the layout and style generated by the decision template. The general large model is an initial untrained multimodal large model.
[0038] In the embodiment of the present invention, the generation process of the initial training template can be simply described as generating a set of initial training parameters to be optimized based on prior experience, constructing an initial typesetting model using the initialized linear parameter values and nonlinear parameter values, and the initial typesetting model generating an initial training template based on each predefined parameter value. The initial training template is input into the original universal large model for document recognition. Since the universal large model also learns document recognition, it will output a first recognition result that is significantly different from the initial training template. The universal recognition error rate is also the recognition difference rate between the first recognition result and the initial training template.
[0039] By obtaining the initial typesetting model, identifying the training parameters to be optimized and classifying them as nonlinear or linear parameters, the foundation is laid for subsequent targeted parameter adjustment and model optimization, ensuring that different types of parameters can be processed according to appropriate rules, improving the accuracy and effectiveness of the optimization process. Generating the initial training template and obtaining the general recognition error rate of the general large model for it provides a benchmark for subsequent iterative optimization, making it easier to compare and evaluate the improvement effects during the optimization process, understand the performance of the model in its initial state, and better guide the direction of parameter adjustment.
[0040] S120. Iteratively generate the first type of typesetting templates corresponding to each linear parameter through linear parameter adjustment rules, and iteratively determine the first type of training templates corresponding to each linear parameter by inputting the first type of typesetting templates into the general large model for document recognition.
[0041] In an embodiment of the present invention, linear parameters and nonlinear parameters have different parameter adjustment rules. Specifically, the first type of layout template refers to a training template that can be updated and replaced during each iteration, and the first type of training template is the typesetting template finally determined by each linear parameter, that is, it is the optimal training template generated after the iteration of the current linear parameter for input into the general large model for training.
[0042] Further, by adjusting the linear parameter rules, iteratively generating the first type layout templates corresponding to each linear parameter, and by inputting the first type layout templates into the general large model for document recognition, iteratively determining the first type training templates corresponding to each linear parameter, which may include:
[0043] Obtaining a current linear parameter from the training parameters to be optimized, and obtaining a current initial linear parameter value of the current linear parameter;
[0044] According to a preset initial parameter adjustment method, the current initial linear parameter value is adjusted to obtain an adjusted linear parameter value of the current initial linear parameter value;
[0045] According to the adjusted linear parameter value of the current initial linear parameter value, the initial typesetting model is updated to obtain a current updated linear model, and a current updated linear template corresponding to the current updated linear model is generated;
[0046] Inputting the current updated linear template into the general large model to obtain the current recognition error rate of the general large model for the current updated linear model;
[0047] Calculating the difference between the current recognition error rate and the target recognition error rate, and calculating a new adjusted linear parameter value of the current initial linear parameter value according to the difference, and then updating the current recognition error rate to the new target recognition error rate, wherein the target recognition error rate is initialized to the general recognition error rate;
[0048] Return to execute the operation of updating the initial typesetting model according to the adjustment parameter value of the current initial parameter value until the end iteration condition is met.
[0049] Among multiple training parameters to be optimized, all linear parameters are extracted. At this time, each linear parameter carries a predefined initialization parameter value. Since the current linear parameter lacks a basis for parameter adjustment during the first iteration, the initial parameter adjustment method can be a pre-specified parameter adjustment method, such as random generation. Through the above method, an adjusted linear parameter value different from the initialization parameter value is obtained. At this time, the change in the parameter value causes a change in the initial typesetting model, that is, the currently updated linear model. After the initial typesetting model changes, a new typesetting template corresponding to the adjusted linear parameter value will be generated, that is, the currently updated linear template.
[0050] Since in the above embodiment, the general recognition error rate has been calculated by inputting the initial typesetting template into the general large model, when the current updated linear template is input into the general large model again, the obtained recognition error rate will replace the previous general recognition error rate as the current recognition error rate, and accordingly, the replaced previous general recognition error rate is called the target recognition error rate. The parameter adjustment amount of the current adjusted linear parameter value can be obtained by calculating the difference between the two recognition error rates (the target recognition error rate and the current recognition error rate), and the current adjusted linear parameter value is adjusted based on the parameter adjustment amount to obtain a new adjusted linear parameter value.
[0051] During the iteration process, the recognition error rate used as the current recognition error rate in the previous iteration process will be used as the target recognition error rate, and the newly obtained recognition error rate in this iteration process will be used as the current recognition error rate. The above process is repeated, and the end condition of the iteration can be reaching a certain number of iterations, or the recognition error rate converges to below a fixed threshold.
[0052] The embodiment of the present invention continuously adjusts the current linear parameter value according to the specific linear parameter adjustment steps, generates an updated linear model and template, and adjusts the parameters according to the recognition error rate difference value feedback, forming an iterative optimization process, so that the linear parameters can gradually converge to the optimal value, improve the model's processing ability for linear related features, and thus improve the accuracy and stability of document recognition. By calculating the formula for the parameter adjustment amount, the linear parameters can be reasonably adjusted according to the change of the recognition error rate, so that the model can continuously reduce the recognition error rate during the iteration process until the iteration termination condition is met, and finally a model optimized for the linear parameters is obtained, which improves the performance of the model in processing document features related to the linear parameters.
[0053] Optionally, calculating a new adjustment parameter value of the current initial parameter value according to the difference value may include:
[0054] According to the formula Calculate the current parameter adjustment Δp i ;
[0055] Where k is the adjustment coefficient, dE is the difference value of the recognition error rate, which is obtained by calculating the difference between the current recognition error rate and the target recognition error rate, dp is the change of the current parameter value, which is obtained by calculating the difference between the current parameter value and the adjusted linear parameter value of the current parameter value, and p i is the adjusted linear parameter value of the current parameter value;
[0056] The adjusted parameter value of the current initial parameter value is adjusted according to the current initial parameter value adjustment amount to obtain a new adjusted parameter value of the current initial parameter value.
[0057] In the embodiment of the present invention, a specific method for adjusting the linear parameters is provided, wherein the concept of parameter value adjustment amount is introduced, that is, the adjusted linear parameter value is obtained by adjusting the current parameter value by the parameter adjustment amount. For ease of understanding, taking the first iteration as an example, dE is the difference between the general recognition error rate corresponding to the initial typesetting template and the current recognition error rate corresponding to the current updated linear template obtained by the first adjustment; dp is the difference between the current initial linear parameter value and the adjusted linear parameter value; and p i The adjusted linear parameter value is the current initial linear parameter value obtained by the first adjustment.
[0058] The current initial parameter value adjustment is calculated based on a specific formula, taking into account factors such as the adjustment coefficient, the difference in recognition error rate, and the change in the current initial parameter value. This calculation method can scientifically and reasonably determine the parameter adjustment range based on the performance of the model at different iteration stages, avoiding over-adjustment or under-adjustment, and ensuring the effectiveness and stability of the linear parameter adjustment process. Adjusting the parameter value according to the calculated adjustment amount helps guide the model in the direction of reducing the recognition error rate during the optimization process, allowing the model to better adapt to document features, thereby improving the overall performance of the model in document recognition tasks.
[0059] S130. Generate all second-category typesetting templates corresponding to each nonlinear parameter through nonlinear parameter adjustment rules, and determine the second-category training templates corresponding to each nonlinear parameter by inputting all second-category typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition.
[0060] Similarly, the above process describes the process of using nonlinear parameter adjustment rules to generate the second type of layout template corresponding to each nonlinear parameter. Different from linear parameters, due to the characteristics of nonlinear parameters, this process requires more complex adjustment rules to ensure that the generated layout template can cover all possible values or states of nonlinear parameters.
[0061] Further, all second-class typesetting templates corresponding to each nonlinear parameter are generated through nonlinear parameter adjustment rules, and all second-class typesetting templates corresponding to each nonlinear parameter are input into the general large model for document recognition to determine each second-class training template corresponding to each nonlinear parameter, which may include:
[0062] Obtain the current nonlinear parameter from the training parameters to be optimized, and obtain the current initial nonlinear parameter value and all optional values of the current nonlinear parameter;
[0063] Replacing the current initial nonlinear parameter value with an optional value to obtain a current optional nonlinear parameter value, and updating the initial typesetting model according to the current optional nonlinear parameter value to obtain a current optional nonlinear model, and generating a current optional nonlinear template corresponding to the current optional nonlinear model;
[0064] Inputting the current optional nonlinear template into the universal large model to obtain the current recognition error rate of the universal large model for multiple current optional nonlinear models;
[0065] Return execution to replace the current initial nonlinear parameter value with the optional value until the current recognition error rate is generated for all optional values of the current nonlinear parameter;
[0066] Calculating an adjusted nonlinear parameter value for a current optional nonlinear parameter value based on all current recognition error rates for the current nonlinear parameter;
[0067] The initial typesetting model is updated by using the adjusted nonlinear parameter value of the current initial nonlinear parameter value to obtain a current updated nonlinear model, and a current updated nonlinear template corresponding to the current updated nonlinear model is generated.
[0068] First, for nonlinear parameters, each optional value corresponds to a solution of the current nonlinear parameter. For example, for the nonlinear parameter of font size, based on certain rules or restrictions of the intelligent typesetting model, its optional value includes at least one value used to describe the font size. In the embodiment of the present invention, the specific number and value of the optional values under each nonlinear parameter are not limited, but it can be determined that by selecting the current nonlinear parameter, the total number of optional values under the nonlinear parameter and the specific value of the optional value can be obtained. All the values of the optional values are traversed, and the currently selected optional value replaces the initial nonlinear parameter value to obtain the optional nonlinear parameter value, and the optional nonlinear model corresponding to each optional nonlinear parameter value is also the second type of typesetting template.
[0069] For each optional nonlinear template, after being input into the large model, there is a corresponding recognition result, that is, the current recognition error rate. In the embodiment of the present invention, unlike the linear parameter, since all optional values of the nonlinear parameter are traversed, the corresponding recognition error rate is obtained for all optional values of the current nonlinear parameter. All recognition error rates are obtained, and the adjusted nonlinear parameter value of the current nonlinear parameter is determined based on a certain calculation rule.
[0070] For nonlinear parameters, by replacing the initial values with all their optional values to generate different optional nonlinear models and templates, and obtaining the corresponding recognition error rates, it is possible to fully explore the value space of nonlinear parameters, fully tap the impact of different values on model performance, avoid missing possible optimal solutions, and provide a wider search range for finding the optimal nonlinear parameter values. The adjusted nonlinear parameter values are calculated based on all current recognition error rates, so that the adjustment of nonlinear parameters is based on a comprehensive evaluation of multiple value situations, and the parameter values suitable for the model can be determined more accurately, thereby optimizing the model's ability to handle nonlinear features in documents and improving the model's performance in processing complex document structures and nonlinear relationships.
[0071] Optionally, calculating the adjusted nonlinear parameter value of the current optional nonlinear parameter value according to all current recognition error rates of the current nonlinear parameter may include:
[0072] Obtaining current recognition error rates of all currently optional nonlinear parameter values of the current nonlinear parameter, and sorting all recognition error rates of the current nonlinear parameter in descending order;
[0073] According to the pre-configured hyperparameters, the target optional nonlinear parameter values corresponding to the top K recognition error rates are used to form a nonlinear error rate sequence;
[0074] Performing normalization calculation on the nonlinear error rate sequence to determine the proportion corresponding to each target optional nonlinear parameter value in the nonlinear error rate sequence;
[0075] Each target optional nonlinear parameter value is adjusted according to the proportion to obtain an adjusted nonlinear parameter value of the current nonlinear parameter.
[0076] For a specific nonlinear parameter, it is necessary to evaluate the performance of all possible values (optional nonlinear parameter values) on the model, that is, calculate the template recognition error rate under these values, and sort all recognition rates from high to low according to the error rate, in order to find those training templates with large deviations in recognition results and high complexity. The pre-configured hyperparameter can be a threshold k pre-set by the user or algorithm, which is used to determine how many optional values with the worst performance to consider, and select the values of the above k recognition error rates to form a "nonlinear error rate sequence".
[0077] Normalization is a data processing technology used to convert data of different ranges to the same scale. In an embodiment of the present invention, it is used to measure the relative importance or proportion of each target optional nonlinear parameter in the overall error. The calculated proportion can be regarded as the probability of each target optional nonlinear parameter value appearing in the entire training template in the next adjustment. For example, if the normalized proportion is 40%, 20%, 30% and 10%, it means that the hyperparameter is 4, the number of target optional nonlinear parameters is 4, and the probability of the corresponding optional value appearing in the next adjustment is 40%, 20%, 30% and 10%. The adjusted nonlinear parameter value of the current nonlinear parameter is regenerated according to the above proportion.
[0078] By sorting all the current recognition error rates of the current nonlinear parameters and selecting the target optional nonlinear parameter values corresponding to the top K error rates according to the hyperparameters to form a sequence, we can focus on the key parameter values that have a greater impact on the model performance, reduce the amount of subsequent calculations, and highlight the role of important parameters, thereby improving the calculation efficiency and optimization effect. The nonlinear error rate sequence is normalized to determine the weight, and then the target optional nonlinear parameter value is adjusted according to the weight. This method can make reasonable adjustments based on the relative importance of the impact of different parameter values on the error rate, so that the adjusted nonlinear parameter value is more in line with the model optimization requirements, further improving the accuracy and adaptability of the model when processing nonlinear features.
[0079] S140: Optimize the general large model using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
[0080] The embodiment of the present invention describes a method for optimizing a large document recognition model, which obtains a plurality of training parameters to be optimized (including nonlinear and linear parameters), and generates corresponding typesetting templates based on linear and nonlinear parameter adjustment rules, and then determines the training template to optimize the general large model, so that the model can be adjusted according to the typesetting characteristics of different types of documents, thereby improving the adaptability of the model to various document structures and formats. Generating training templates based on specific adjustment rules and optimizing the model using these templates can help the model better learn the features and patterns in the document, so that the document content can be more accurately identified and parsed in the document recognition task, thereby reducing the recognition error rate.
[0081] Embodiment 2
[0082] Figure 2 This is a schematic diagram of the structure of a device for optimizing a large document recognition model provided in the second embodiment of the present invention. Figure 2 As shown, the device comprises:
[0083] A training parameter acquisition module 210 is used to acquire a plurality of training parameters to be optimized for optimizing the general large model, wherein the training parameters include nonlinear parameters and linear parameters;
[0084] The first type of training template determination module 220 is used to iteratively generate the first type of typesetting template corresponding to each linear parameter by adjusting the linear parameter rule, and iteratively determine the first type of training template corresponding to each linear parameter by inputting the first type of typesetting template into the general large model for document recognition;
[0085] The second type training template determination module 230 is used to generate all the second type layout templates corresponding to each non-linear parameter by adjusting the non-linear parameter rules, and determine the second type training templates corresponding to each non-linear parameter by inputting all the second type layout templates corresponding to each non-linear parameter into the general large model for document recognition;
[0086] The document recognition large model optimization module 240 is used to optimize the general large model using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
[0087] The embodiment of the present invention describes a method for optimizing a large document recognition model, which obtains a plurality of training parameters to be optimized (including nonlinear and linear parameters), and generates corresponding typesetting templates based on linear and nonlinear parameter adjustment rules, and then determines the training template to optimize the general large model, so that the model can be adjusted according to the typesetting characteristics of different types of documents, thereby improving the adaptability of the model to various document structures and formats. Generating training templates based on specific adjustment rules and optimizing the model using these templates can help the model better learn the features and patterns in the document, so that the document content can be more accurately identified and parsed in the document recognition task, thereby reducing the recognition error rate.
[0088] Optionally, based on the above embodiments, the training parameter acquisition module 210 may include:
[0089] An initial parameter value acquisition unit, used to acquire a predefined initial typesetting model, identify a plurality of training parameters to be optimized in the initial training template, and acquire an initial parameter value of each of the training parameters to be optimized;
[0090] A parameter classification unit, used to classify each of the training parameters to be optimized into nonlinear parameters or linear parameters according to a preset classification standard;
[0091] The recognition error rate calculation module is used to generate an initial training template corresponding to the initial typesetting model, and input the initial training template into the general large model to obtain the general recognition error rate of the general large model for the initial training template.
[0092] Optionally, based on the above embodiments, the first type training template determination module 220 may include:
[0093] An initial linear parameter value acquisition unit, used to acquire a current linear parameter from the training parameters to be optimized, and to acquire a current initial linear parameter value of the current linear parameter;
[0094] A linear parameter adjustment unit, configured to adjust the current initial linear parameter value according to a preset initial parameter adjustment method to obtain an adjusted linear parameter value of the current initial linear parameter value;
[0095] A first type of typesetting template generating unit, configured to update the initial typesetting model according to the adjusted linear parameter value of the current initial linear parameter value, obtain a current updated linear model, and generate a current updated linear template corresponding to the current updated linear model;
[0096] A linear recognition error rate calculation unit, used for inputting the current updated linear template into the general large model to obtain the current recognition error rate of the general large model for the current updated linear model;
[0097] a recognition error rate updating unit, configured to calculate a difference between a current recognition error rate and a target recognition error rate, and to calculate a new adjusted linear parameter value of the current initial linear parameter value according to the difference value, and then update the current recognition error rate to a new target recognition error rate, wherein the target recognition error rate is initialized to a universal recognition error rate;
[0098] The loop iteration unit is used to return to execute the operation of updating the initial typesetting model according to the adjustment parameter value of the current initial parameter value until the end of the iteration condition is met.
[0099] Optionally, based on the above embodiments, the linear template generating unit can be used to generate a linear template according to the formula Calculate the current initial parameter value adjustment Δp i ;
[0100] Where k is the adjustment coefficient, dE is the difference value of the recognition error rate, which is obtained by calculating the difference between the current recognition error rate and the target recognition error rate, dp is the change of the current initial parameter value, which is obtained by calculating the difference between the current initial parameter value and the adjusted parameter value of the current initial parameter value, and p i is the adjusted parameter value of the current initial parameter value;
[0101] The adjusted parameter value of the current initial parameter value is adjusted according to the current initial parameter value adjustment amount to obtain a new adjusted parameter value of the current initial parameter value.
[0102] Optionally, based on the above embodiments, the second type training template determination module 230 may include:
[0103] An optional value traversal unit, used to obtain the current nonlinear parameter in the training parameters to be optimized, and obtain the current initial nonlinear parameter value and all optional values of the current nonlinear parameter;
[0104] The second type of typesetting template generating unit is used to replace the current initial nonlinear parameter value with an optional value to obtain the current optional nonlinear parameter value, and update the initial typesetting model according to the current optional nonlinear parameter value to obtain the current optional nonlinear model, and generate the current optional nonlinear template corresponding to the current optional nonlinear model;
[0105] A linear recognition error rate calculation module, used for inputting the current optional nonlinear template into the general large model to obtain the current recognition error rate of the general large model for multiple current optional nonlinear models;
[0106] A linear recognition error rate collection unit, used to return and execute replacing the current initial nonlinear parameter value with the optional value until the current recognition error rate is generated for all optional values of the current nonlinear parameter;
[0107] A nonlinear parameter adjustment unit, used for calculating an adjusted nonlinear parameter value of a current optional nonlinear parameter value according to all current recognition error rates of the current nonlinear parameter;
[0108] The second type of training template generating unit is used to update the initial typesetting model using the adjusted nonlinear parameter value of the current initial nonlinear parameter value to obtain the current updated nonlinear model, and generate the current updated nonlinear template corresponding to the current updated nonlinear model.
[0109] Optionally, based on the above embodiments, the nonlinear parameter adjustment unit may be used to obtain current recognition error rates of all currently optional nonlinear parameter values of the current nonlinear parameter, and sort all recognition error rates of the current nonlinear parameter in descending order;
[0110] According to the pre-configured hyperparameters, the target optional nonlinear parameter values corresponding to the top K recognition error rates are used to form a nonlinear error rate sequence;
[0111] Performing normalization calculation on the nonlinear error rate sequence to determine the proportion corresponding to each target optional nonlinear parameter value in the nonlinear error rate sequence;
[0112] Each target optional nonlinear parameter value is adjusted according to the proportion to obtain an adjusted nonlinear parameter value of the current nonlinear parameter.
[0113] An optimization device for a large document recognition model provided by an embodiment of the present invention can execute an optimization method for a large document recognition model provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0114] Embodiment 3
[0115] Figure 3 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0116] like Figure 3 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0117] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0118] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an optimization method for a large document recognition model.
[0119] That is: obtaining a plurality of training parameters to be optimized for optimizing the general large model, wherein the training parameters include nonlinear parameters and linear parameters;
[0120] Iteratively generate first-class typesetting templates corresponding to each linear parameter through linear parameter adjustment rules, and iteratively determine first-class training templates corresponding to each linear parameter by inputting the first-class typesetting templates into the general large model for document recognition;
[0121] By means of nonlinear parameter adjustment rules, all second-class typesetting templates corresponding to each nonlinear parameter are generated, and by inputting all second-class typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition, each second-class training template corresponding to each nonlinear parameter is determined;
[0122] The general large model is optimized using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
[0123] In some embodiments, a method for optimizing a large model for document recognition may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for optimizing a large model for document recognition described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute a method for optimizing a large model for document recognition in any other appropriate manner (e.g., by means of firmware).
[0124] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0126] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0127] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0128] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0129] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0130] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0131] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for optimizing a large document recognition model, characterized in that: include: Obtaining a plurality of training parameters to be optimized for optimizing a general large model, wherein the training parameters include nonlinear parameters and linear parameters; Iteratively generate first-class typesetting templates corresponding to each linear parameter through linear parameter adjustment rules, and iteratively determine first-class training templates corresponding to each linear parameter by inputting the first-class typesetting templates into the general large model for document recognition; By means of nonlinear parameter adjustment rules, all second-class typesetting templates corresponding to each nonlinear parameter are generated, and by inputting all second-class typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition, each second-class training template corresponding to each nonlinear parameter is determined; The general large model is optimized using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
2. The method according to claim 1, characterized in that Get multiple training parameters to be optimized contained in the training template for optimizing the general large model, including: Obtaining a predefined initial typesetting model, identifying a plurality of training parameters to be optimized in the initial training template, and obtaining an initial parameter value of each of the training parameters to be optimized; According to a preset classification standard, each of the training parameters to be optimized is classified into a nonlinear parameter or a linear parameter; An initial training template corresponding to the initial typesetting model is generated, and the initial training template is input into the general large model to obtain a general recognition error rate of the general large model for the initial training template.
3. The method according to claim 2, characterized in that By adjusting the linear parameter rules, iteratively generating the first type layout templates corresponding to each linear parameter, and by inputting the first type layout templates into the general large model for document recognition, iteratively determining the first type training templates corresponding to each linear parameter, including: Obtaining a current linear parameter from the training parameters to be optimized, and obtaining a current initial linear parameter value of the current linear parameter; According to a preset initial parameter adjustment method, the current initial linear parameter value is adjusted to obtain an adjusted linear parameter value of the current initial linear parameter value; According to the adjusted linear parameter value of the current initial linear parameter value, the initial typesetting model is updated to obtain a current updated linear model, and a current updated linear template corresponding to the current updated linear model is generated; Inputting the current updated linear template into the general large model to obtain the current recognition error rate of the general large model for the current updated linear model; Calculating the difference between the current recognition error rate and the target recognition error rate, and calculating a new adjusted linear parameter value of the current initial linear parameter value according to the difference, and then updating the current recognition error rate to the new target recognition error rate, wherein the target recognition error rate is initialized to the general recognition error rate; Return to execute the operation of updating the initial typesetting model according to the adjustment parameter value of the current initial parameter value until the end iteration condition is met.
4. The method according to claim 3, characterized in that Calculate the new adjustment parameter value of the current initial parameter value based on the difference value, including: According to the formula Calculate the current parameter adjustment Δp i ; Where k is the adjustment coefficient, dE is the difference value of the recognition error rate, which is obtained by calculating the difference between the current recognition error rate and the target recognition error rate, dp is the change of the current parameter value, which is obtained by calculating the difference between the current parameter value and the adjusted linear parameter value of the current parameter value, and p i is the adjusted linear parameter value of the current parameter value; The adjusted parameter value of the current initial parameter value is adjusted according to the current initial parameter value adjustment amount to obtain a new adjusted parameter value of the current initial parameter value.
5. The method according to claim 2, characterized in that: By means of nonlinear parameter adjustment rules, all second-class typesetting templates corresponding to each nonlinear parameter are generated, and by inputting all second-class typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition, the second-class training templates corresponding to each nonlinear parameter are determined, including: Obtain the current nonlinear parameter from the training parameters to be optimized, and obtain the current initial nonlinear parameter value and all optional values of the current nonlinear parameter; Replacing the current initial nonlinear parameter value with an optional value to obtain a current optional nonlinear parameter value, and updating the initial typesetting model according to the current optional nonlinear parameter value to obtain a current optional nonlinear model, and generating a current optional nonlinear template corresponding to the current optional nonlinear model; Inputting the current optional nonlinear template into the general large model to obtain the current recognition error rate of the general large model for multiple current optional nonlinear models; Return execution to replace the current initial nonlinear parameter value with the optional value until the current recognition error rate is generated for all optional values of the current nonlinear parameter; Calculating an adjusted nonlinear parameter value for a current optional nonlinear parameter value based on all current recognition error rates for the current nonlinear parameter; The initial typesetting model is updated by using the adjusted nonlinear parameter value of the current initial nonlinear parameter value to obtain a current updated nonlinear model, and a current updated nonlinear template corresponding to the current updated nonlinear model is generated.
6. The method according to claim 5, characterized in that Calculating an adjusted nonlinear parameter value of a current optional nonlinear parameter value based on all current recognition error rates of the current nonlinear parameter, including: Obtaining current recognition error rates of all currently optional nonlinear parameter values of the current nonlinear parameter, and sorting all recognition error rates of the current nonlinear parameter in descending order; According to the pre-configured hyperparameters, the target optional nonlinear parameter values corresponding to the top K recognition error rates are used to form a nonlinear error rate sequence; Performing normalization calculation on the nonlinear error rate sequence to determine the proportion corresponding to each target optional nonlinear parameter value in the nonlinear error rate sequence; Each target optional nonlinear parameter value is adjusted according to the proportion to obtain an adjusted nonlinear parameter value of the current nonlinear parameter.
7. An optimization device for a large document recognition model, characterized in that: include: A training parameter acquisition module, used to acquire a plurality of training parameters to be optimized for optimizing the general large model, wherein the training parameters include nonlinear parameters and linear parameters; A first-class training template determination module, for iteratively generating a first-class typesetting template corresponding to each linear parameter by adjusting a linear parameter rule, and iteratively determining each first-class training template corresponding to each linear parameter by inputting the first-class typesetting template into a universal large model for document recognition; The first type of training template determination module is used to generate all the second type of typesetting templates corresponding to each nonlinear parameter by adjusting the nonlinear parameter rules, and determine the second type of training templates corresponding to each nonlinear parameter by inputting all the second type of typesetting templates corresponding to each nonlinear parameter into the general large model for document recognition; The document recognition large model optimization module is used to optimize the general large model using each first-category training template and each second-category training template to obtain an optimized document recognition large model.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a method for optimizing a large document recognition model as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a method for optimizing a large document recognition model according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements a method for optimizing a large document recognition model according to any one of claims 1-6.