A method and system for training a field recognition model for a document

By annotating fields and training dynamic masks on various types of historical bills in the bill scenario, and combining reinforcement learning mechanisms, the problem of traditional models being highly dependent on local features was solved, and the stability and adaptability of the field recognition model in multiple feature dimensions were improved.

CN122116384AInactive Publication Date: 2026-05-29YOUDINGTE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YOUDINGTE TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-05-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional field recognition model training methods lack the ability to dynamically adjust the participation of different feature information during the training process, resulting in strong dependence of the model on local features, insufficient recognition stability, and difficulty in forming reliable and consistent field recognition results.

Method used

By annotating fields of various types of historical tickets, text training sets and image training sets are formed. Dynamic mask training is performed on different feature dimensions of the ticket scenario. Combined with the reinforcement learning mechanism of the ticket scenario, the parameters of the field recognition model are iteratively adjusted until the consistency of the model on different feature dimensions meets the preset conditions.

Benefits of technology

It enhances the accuracy and robustness of the model, reduces the dependence on a single feature dimension, significantly improves the accuracy, stability and adaptability of field recognition, and enhances the model's performance in complex invoice scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116384A_ABST
    Figure CN122116384A_ABST
Patent Text Reader

Abstract

The application relates to a field recognition model training method and system for bills. The method comprises the following steps: field labeling is performed on multi-class historical bills to obtain a text training set and an image training set; the text training set and the image training set are input into a field recognition model to be trained, dynamic mask training is performed on the text training set and the image training set in different feature dimensions of a bill scene, and the parameters of the field recognition model are iteratively adjusted in combination with a reinforcement learning mechanism corresponding to the bill scene; until the consistency of the field recognition model in different feature dimensions meets a preset condition, the field recognition model is taken as a trained target field recognition model for extracting and recognizing the fields of any type of bill. By combining dynamic mask training with a reinforcement learning mechanism, the model can effectively learn and autonomously adjust in multiple feature dimensions, so that the precision, stability and adaptability of field recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of invoice recognition technology, and in particular to a method and system for training a field recognition model for invoices. Background Technology

[0002] In the field of invoice recognition technology, the extraction and recognition of invoice fields are achieved through field recognition models. However, traditional field recognition model training methods only learn the static expression relationship of invoice fields through single or multiple feature information, lacking the ability to dynamically adjust the participation mode of different feature information during training. This results in the model being highly dependent on local features, and lacking recognition stability when features are missing or inconsistent, making it difficult to form reliable and consistent field recognition results. Summary of the Invention

[0003] Therefore, it is necessary to provide a method, system, computer device, and computer-readable storage medium for training a field recognition model for invoices, addressing the aforementioned technical problems.

[0004] Firstly, this application provides a method for training a field recognition model for invoices, including: Fields are labeled on multiple types of historical tickets to obtain text training sets and image training sets; The text training set and the image training set are input into the field recognition model to be trained. Dynamic mask training is performed on the text training set and the image training set on different feature dimensions of the bill scene. Combined with the reinforcement learning mechanism corresponding to the bill scene, the parameters of the field recognition model are iteratively adjusted. Once the consistency of the field recognition model across different feature dimensions in the invoice scenario meets the preset conditions, the field recognition model will be used as the trained target field recognition model for extracting and recognizing fields for any type of invoice.

[0005] Secondly, this application also provides a training system for a field recognition model for invoices, comprising: The annotation module is used to annotate fields of multiple preset types of historical tickets to obtain text training sets and image training sets. The training module is used to input the text training set and the image training set into the field recognition model to be trained, perform dynamic mask training on the text training set and the image training set on different feature dimensions of the invoice scenario, and iteratively adjust the parameters of the field recognition model in combination with the reinforcement learning mechanism corresponding to the invoice scenario. The generation module is used to select the field recognition model as a trained target field recognition model when the consistency of the field recognition model across different feature dimensions in the invoice scenario meets a preset condition, so as to extract and recognize fields for any type of invoice.

[0006] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the above steps.

[0007] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the above steps.

[0008] The aforementioned training method, system, computer equipment, and computer-readable storage medium for the field recognition model of invoices firstly, by annotating the fields of multiple types of historical invoices, text and image training sets are obtained. This ensures that the training data covers different invoice types and presentation formats, guaranteeing the diversity and broad applicability of the training model. Secondly, by inputting the text and image training sets into the field recognition model and performing dynamic mask training, combined with reinforcement learning mechanisms corresponding to the invoice scenario, the model can effectively learn the field correlations across feature dimensions during training and adaptively optimize the learning strategy. Finally, by evaluating the consistency of the field recognition model across different feature dimensions in the invoice scenario, the model can stably recognize fields across different feature dimensions, thereby enhancing the model's accuracy and robustness and reducing dependence on a single feature dimension. Based on this, the combination of dynamic mask training and reinforcement learning mechanisms in the entire technical solution enables the model to effectively learn and autonomously adjust across multiple feature dimensions, significantly improving the accuracy, stability, and adaptability of field recognition and enhancing the model's performance in complex invoice scenarios. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating a method for training a field recognition model for invoices in one embodiment. Figure 2 This is a block diagram of a field recognition model training system for invoices in one embodiment. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0012] In one embodiment, such as Figure 1 As shown, a method for training a field recognition model for invoices is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, or to a system including a terminal and a server, and can be implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S101 to S103.

[0013] Step S101: Field annotation is performed on the preset multiple types of historical tickets to obtain text training set and image training set.

[0014] Historical documents refer to invoices, receipts, or reimbursement vouchers that were created and archived in different historical periods.

[0015] For example, the fields in each historical ticket are manually or semi-automatically annotated to ensure that key information in the historical ticket has a recognizable correspondence at both the text and image levels. Specifically, at the text level, the positions of each field in the historical ticket are marked to create a clear start and end range for the field content in the text sequence, thereby constructing a text training set to represent the semantic structure of the fields. At the image level, based on the original images of the historical tickets, the positions of the fields in the ticket layout are annotated accordingly, so that the field content has a clear regional pointing relationship in the image space, thereby forming an image training set that corresponds one-to-one with the text training set.

[0016] By employing the above methods, a stable correlation is established between the textual and image representations of the same field, providing fundamental data support for subsequent models to process both textual and image information simultaneously. Furthermore, due to the differences in invoice types, this step covers multiple types of historical invoices, ensuring that the resulting training set possesses sufficient diversity in field distribution, layout structure, and content combination, thus providing representative input for subsequent model training.

[0017] Step S102: Input the text training set and the image training set into the field recognition model to be trained, perform dynamic mask training on the text training set and the image training set on different feature dimensions of the invoice scenario, and combine the reinforcement learning mechanism corresponding to the invoice scenario to iteratively adjust the parameters of the field recognition model.

[0018] Among them, the field recognition model refers to the computational model used in the training phase to recognize the field content in historical tickets, in order to learn the field structure, semantic information and the representation of fields at the text and image levels.

[0019] Among them, the different feature dimensions of the bill scenario represent different information expression levels used to describe the content of the bill, such as the text feature dimension reflected in the text content of the bill and the image feature dimension reflected in the layout of the bill.

[0020] Among them, the reinforcement learning mechanism corresponding to the invoice scenario represents a training feedback and adjustment method set for the invoice field recognition task. For example, it is a training mechanism that generates feedback signals based on factors such as the accuracy of field recognition and the consistency of field correspondence under different feature dimensions, and adjusts the direction of model parameter update accordingly.

[0021] For example, the text training set and the image training set are simultaneously input into the field recognition model to be trained. Inside the model, dynamic masking training is performed on the text training set and the image training set to address the differences in the expression of field content in different feature dimensions in the ticket scenario. That is, during the training process, some feature information is selectively masked according to the current training state, so that the model still needs to model the field relationship even when local information is missing, thereby prompting the model to gradually establish the ability to associate across feature dimensions.

[0022] At the same time, a reinforcement learning mechanism that matches the bill scenario is introduced, so that the model can obtain feedback signals in each round of training based on the comprehensive performance of the training results of different feature dimensions in the bill scenario, and adjust the model parameters according to the feedback signals, thereby guiding the training process to converge in a direction that is more in line with the characteristics of bill fields.

[0023] Based on this, by combining dynamic mask training with reinforcement learning mechanisms, the model continuously corrects its internal expression of the relationship between field structure, field position and field semantics during repeated iterations, thereby driving the model parameters to converge toward a state that better matches the actual distribution of characteristics of the invoice.

[0024] Step S103: Until the consistency of the field recognition model across different feature dimensions in the invoice scenario meets the preset conditions, the field recognition model is used as the trained target field recognition model for extracting and recognizing fields for any type of invoice.

[0025] For example, after multiple rounds of iterative adjustments to the field recognition model, the performance of the field recognition results output by the model in the current training state across different feature dimensions in a bill scenario is compared. This ensures that the recognition results of the same field in different feature dimensions can form comparable expressions, and determines whether a stable correspondence is maintained between each feature dimension in terms of field boundaries, field affiliation, and field content expression. Based on this, when the consistency of the field recognition results across different feature dimensions meets preset conditions—for example, the field boundaries, field affiliation, and field content expression are consistent and without significant deviations across feature dimensions—it indicates that the model has been able to form a relatively stable field recognition logical structure across different feature dimensions, thus enabling it to complete the bill field recognition task without relying on information from a single feature dimension.

[0026] Therefore, the field recognition model in the current training state is identified as the trained target field recognition model, and it is used as the core model in the subsequent invoice processing stage to support the extraction and recognition of fields for any type of invoice, thus completing the transition from the training stage to the practical application stage.

[0027] In the above-mentioned training method for the field recognition model of invoices, in step S101, text training sets and image training sets are obtained by annotating fields of multiple types of historical invoices, thereby ensuring that the training data can cover different invoice types and presentation forms, and ensuring the diversity and wide applicability of the training model. In step S102, the text training sets and image training sets are input into the field recognition model and subjected to dynamic mask training. Combined with the reinforcement learning mechanism corresponding to the invoice scenario, the model can effectively learn the field correlation across feature dimensions during training and adaptively optimize the learning strategy. In step S103, the consistency of the field recognition model on different feature dimensions of the invoice scenario is evaluated to ensure that the model can stably recognize fields across different feature dimensions, thereby enhancing the accuracy and robustness of the model and reducing the dependence on a single feature dimension. Based on this, in the entire technical solution, the combination of dynamic mask training and reinforcement learning mechanism enables the model to effectively learn and autonomously adjust in multiple feature dimensions, thereby significantly improving the accuracy, stability and adaptability of field recognition and enhancing the model's performance in complex invoice scenarios.

[0028] In an exemplary embodiment, dynamic mask training is performed on the text training set and the image training set on different feature dimensions of the invoice scenario, and the parameters of the field recognition model are iteratively adjusted in combination with the reinforcement learning mechanism corresponding to the invoice scenario, including steps S201 to S203.

[0029] Step S201: Perform text mask training on the text training set in the text feature dimension of the ticket scene, perform image mask training on the image training set in the image feature dimension of the ticket scene, and perform composite mask training on the text training set and image training set in the combined dimension of text features and image features of the ticket scene, to obtain the mask training results for each feature dimension.

[0030] For example, based on the differences in field content across different feature dimensions in a bill scenario, targeted processing is applied to the text training set and the image training set respectively, enabling the model to be exposed to information changes under each feature dimension during the training phase.

[0031] Specifically, in the text feature dimension, by performing text masking on some field content or contextual information in the text training set, the model still needs to complete field recognition based on text sequence relationships even when the text information is incomplete, thus prompting the model to focus on the semantic relationships between fields. In the image feature dimension, by performing image masking on the regions related to fields in the image training set, the model still needs to complete field recognition based on spatial layout relationships even when page information is limited, thus prompting the model to focus on the visual relationships between fields. On this basis, in the dimension of combining text and image features, composite masking is performed on both the text training set and the image training set, so that even when the information expression of a specified category is limited, the model still needs to identify the information expression of another category based on the field association relationship between feature dimensions, thus prompting the model to focus on the cross-dimensional relationships between fields.

[0032] Based on this, the above method generates mask training results corresponding to each feature dimension, enabling the model to fully experience training states under different combinations of missing features during training, providing multiple types of basic data for subsequent evaluation.

[0033] Step S202: Based on the reinforcement learning mechanism corresponding to the ticket scenario, perform reinforcement evaluation on the mask training results of each feature dimension based on field association features to obtain scenario reinforcement evaluation results used to characterize the effectiveness of each mask training process.

[0034] For example, the mask training results for each feature dimension are organized so that the field recognition outputs corresponding to different feature dimensions in the same training round can form a comparable dataset, thus providing a consistent data foundation for subsequent analysis. Based on this, a reinforcement learning mechanism matching the invoice scenario is introduced. This allows the training process to move beyond simply following a fixed training objective, instead combining the existing field association features between invoice fields to evaluate the field structure representation reflected in the mask training results. For instance, by analyzing the field correspondence, field arrangement order, and field positional associations under different masking methods, the applicability of various masking training methods to the invoice field structure representation in the current training round is reflected.

[0035] In other words, the purpose of this evaluation process is to transform the mask training results into an intermediate representation that reflects the effectiveness of the training process, enabling the training process to provide feedback based on the characteristics of the ticket fields. Based on this, through the above processing, a scenario adaptability evaluation result is obtained to represent the training state of the current training round, so that the subsequent training process can guide the adjustment direction of the model parameters based on the evaluation result, thereby ensuring consistency and coherence between the model training process and the structural expression of the ticket fields.

[0036] Step S203: Based on the scene adaptability evaluation results of each training round, the parameters of the field recognition model are iteratively adjusted.

[0037] For example, by comparing the scene reinforcement evaluation results under different training rounds, the model can identify training patterns that are more effective for field recognition in the current invoice scenario, and iteratively adjust the update direction of the model's internal parameters accordingly. This iterative adjustment process enables the model to gradually strengthen its response to effective feature representation methods in subsequent training, while reducing its dependence on inefficient training states, thereby driving the overall training process of the model to converge in a direction that is more in line with the characteristics of the invoice scenario.

[0038] Based on this, by repeatedly performing this parameter adjustment process in each round of training, the field recognition model gradually develops the ability to coordinate and process information of multiple feature dimensions through continuous iteration, and finally stabilizes the model parameters to a training state that can support subsequent invoice field recognition tasks.

[0039] For example, in consecutive training epochs, when the model maintains high stability in the attribution, corresponding position, and field boundary consistency of the same field in both text and image features during a particular epoch, the scene reinforcement evaluation result for that epoch is relatively better. Conversely, if the correspondence of fields shifts across different feature dimensions in another epoch, its scene reinforcement evaluation result is lower. By comparing these evaluation results, the model tends to strengthen the training state of the aforementioned stable field correspondences when updating parameters, thereby gradually weakening the impact of unfavorable training states on the model.

[0040] In this embodiment, in step S201, mask training is performed on the text feature dimension, image feature dimension, and their combined dimension of the bill scene, thereby realizing a multi-dimensional training process of the model under different information expression conditions. In step S202, by introducing a reinforcement learning mechanism that matches the bill scene, the mask training results of each feature dimension are evaluated based on field association features, so that the adaptability of different mask training methods to the bill field structure expression can be effectively determined. In step S203, the model parameters are iteratively adjusted based on the scene reinforcement evaluation results formed under different training rounds, thereby guiding the model to gradually strengthen its response to better field structure expression methods. Based on this, the entire technical solution realizes the field recognition model training process that combines multi-feature dimension mask training and reinforcement learning mechanism, enabling the model to stably form a field recognition logical structure that conforms to the characteristics of bill fields.

[0041] In an exemplary embodiment, text mask training is performed on the text training set in the text feature dimension of the bill scene, and image mask training is performed on the image training set in the image feature dimension of the bill scene. Composite mask training is performed on the text training set and the image training set in the combined dimension of text features and image features of the bill scene to obtain the mask training results for each feature dimension, including steps S301 to S303.

[0042] Step S301: On the text feature dimension of the bill scenario, mask any field text in the text training set, and predict the text content of the masked field text based on the semantic distribution of the unmasked field text in the bill structure, so as to obtain the mask training result of the text feature dimension.

[0043] For example, the text corresponding to any field in the text training set is selected and masked, so that the text of this field no longer participates in the model calculation as explicit text content in the current training state. After the text of this field is masked, the remaining unmasked text of the fields is retained as context information and input into the model, so that the model can only predict the text content of the masked text field based on the distribution relationship, relative position relationship and overall semantic structure of other fields in the text sequence of the document.

[0044] For example, in a bill, if the text corresponding to the "Invoice Date" field is obscured, the remaining fields such as "Invoice Code," "Invoice Number," "Buyer's Name," and "Amount" still serve as contextual information. The model can determine that the obscured field should be located before the Amount field and after the Bill Number field based on the order, relative position, and common semantic structure relationships of these fields in the text sequence. Combining this with the overall semantic structure of the bill, the model can predict that the text content of the obscured field is date format information, thereby restoring the text of the field.

[0045] By employing the above processing methods, the model gradually establishes semantic relationships between fields during training, rather than relying solely on the explicit textual features of the fields themselves. This encourages the model to focus on the inherent connections between fields within the document text structure. Consequently, the model's predicted output is compared with the original text content corresponding to the masked fields, and a masked training result is generated based on this comparison. This result reflects the model's ability to understand and reconstruct the document text structure even when local field text information is missing.

[0046] Step S302: On the image feature dimension of the ticket scene, any visual region in the image training set is masked, and the image content of the masked visual region is predicted according to the visual layout of the unmasked visual region in the ticket layout, so as to obtain the mask training result of the image feature dimension.

[0047] For example, any visual region corresponding to the ticket field is selected from the image training set, and this visual region is occluded so that it no longer provides explicit image content information to the model in the current training round. After the visual region is occluded, the remaining unoccluded visual regions are used as an overall reference for the ticket layout and input into the model, so that the model can only predict the image content of the occluded region based on the spatial layout, relative positional relationship, and overall visual structure of other visual regions in the ticket.

[0048] For example, in a bill image, if the numbering area in the upper right corner is obscured, the remaining areas, such as the title area, amount area, and bottom information area, still serve as an overall reference for the bill layout. The model can determine the fixed layout position of the obscured area in the upper right corner of the bill based on the relative position and layout structure of these areas in the bill. Combining this with the common functions of this position in the bill structure, the model can predict that the obscured area should be presented as numbered image content, thus completing the prediction of the visual area.

[0049] Through the above processing method, the model gradually learns the structural relationships between different visual regions in the ticket layout during training, rather than simply recording local image features. Therefore, the model's predicted output is compared with the original image content corresponding to the obscured visual regions, and the masked training results in the image feature dimension are obtained accordingly. This is used to characterize the model's ability to understand and reconstruct the visual structure of the ticket when layout information is limited.

[0050] Step S303: In the dimension of combining text features and image features in the ticket scene, mask any field text in the text training set, and predict the cross-modal alignment state of the masked field text and the corresponding visual region at the same position by judging whether the visual region at the same position in the image training set is masked, so as to obtain the mask training result of the combined dimension.

[0051] For example, the text corresponding to any field in the text training set is masked so that the text cannot be directly obtained in the current training state. At the same time, the corresponding masking state of the visual region in the image training set that is in the same position as the text is masked is obtained and used as part of the model input. Under this condition, the model no longer relies solely on the single information of text or image, but predicts whether the masked text and the corresponding visual region maintain a cross-modal alignment state based on the spatial correspondence between the text and the visual region in the document.

[0052] For example, in a document, when the text corresponding to the "Total Amount" field is obscured, the model simultaneously receives the obscuration status of the visual area corresponding to the field's location. Other unobscured areas, such as the title area and date area, are still used as input. Based on the fact that the amount field is typically located in the numerical area of ​​a document and is usually close to the currency symbol, the model can determine whether the obscured text and the corresponding visual area maintain a consistent obscuration status, thereby predicting whether cross-modal alignment is valid when they are on the same page.

[0053] Through the above processing methods, the model gradually establishes a cognitive association between the text field and its position on the page during training, rather than processing text or image information in isolation. As a result, the model's cross-modal prediction results are compared with the actual alignment, and a mask training result combining text and image features is generated to characterize the model's ability to understand the structure of document fields in the context of multi-dimensional information association.

[0054] Optionally, in addition to masking the field text and predicting its cross-modal alignment with the visual region, the model can also mask the visual region corresponding to a certain field text in the image training set, while retaining the field text as input information. In this case, the model determines whether the masked visual region should be on the same page as the field text based on the position of the field text in the document and its structural relationship with other fields, and predicts the cross-modal alignment between the visual region and the field text accordingly.

[0055] In this embodiment, in step S301, the model's ability to understand and reconstruct the document text structure is improved by masking the field text in the text feature dimension and predicting based on the semantic distribution of the remaining field text. In step S302, the model's ability to understand and reconstruct the document visual structure is improved by masking the visual region in the image feature dimension and predicting based on the visual layout of the remaining visual region. In step S303, the model's ability to model the correspondence between field text and layout position is strengthened by masking the field text in the dimension combining text and image features and predicting the cross-modal alignment state of the masked field text and the corresponding visual region at the same position. Based on this, in the entire technical solution, by implementing masking training under different feature dimensions, the model can simultaneously learn the document text structure, the document visual structure, and the correspondence between the two, thereby forming a stable and consistent document field recognition logical structure.

[0056] In an exemplary embodiment, based on the reinforcement learning mechanism corresponding to the ticket scenario, the mask training results of each feature dimension are subjected to reinforcement evaluation based on field association features to obtain scenario reinforcement evaluation results that characterize the effectiveness of each mask training process, including steps S401 to S402.

[0057] Step S401: Based on the degree to which the mask training results of each feature dimension preserve the field-related features, determine the deviation of each mask training result in the representation of the ticket scenario, and generate reward feedback information to indicate the effectiveness of the corresponding mask training process based on the corresponding deviation.

[0058] For example, the mask training results for each feature dimension are organized so that the field recognition relationships output by different feature dimensions in the same training round can be compared within the same analytical framework. Based on this, according to the original association structure between fields in the document, the field correspondence, field arrangement order, and field position association reflected in each mask training result are checked to determine the degree of preservation of field association features under different mask conditions.

[0059] For example, in the same document, fields typically have a fixed arrangement order and positional relationship, such as the number field being below the title, the amount field being in the middle right of the page, and the date field near the bottom of the document. In training results with different masks, the field correspondences output by the model are compared with this existing structure. If the fields still maintain their original relative order and positional correspondence, it indicates a high degree of preservation of field association features under this mask condition; conversely, if the field order or positional relationship shifts significantly, it indicates a low degree of preservation of field association features under this mask condition.

[0060] Through the above verification process, the deviation of fields from the overall structure of the document in the mask training results can be identified, and this deviation is expressed in a quantitative form, forming a deviation quantity that reflects the degree of difference between the mask training results and the document scene representation. Based on this, reward feedback information is generated to indicate the effectiveness of the corresponding mask training process according to the determined deviation quantity, so that the training results are no longer just in the form of predicted output, but are further transformed into feedback basis that can be used for subsequent evaluation and adjustment.

[0061] For example, within the same training epoch, the mask training results are compared with the original field structure of the ticket for the text feature dimension, the image feature dimension, and the combined dimension of the two. Under the text feature dimension, if the field order is basically consistent with the semantic correspondence, a small text deviation is formed. Under the image feature dimension, if the field position is offset from the layout, a large image deviation is formed. Under the combined dimension, if the correspondence between the text and the visual region is partially mismatched, a moderate deviation is formed. Based on these multi-level deviations, corresponding reward feedback information is generated to comprehensively reflect the effectiveness of the mask training process under different feature dimensions.

[0062] Step S402: Based on the reward feedback information, perform a reinforcement evaluation on the learning benefits of the mask training strategy for each feature dimension in the ticket scenario to obtain the scenario reinforcement evaluation result.

[0063] For example, the reward feedback information corresponding to mask training of different feature dimensions in the current training round is associated with the corresponding mask training strategy, so that the learning effect of various mask training strategies in the ticket scenario can be independently represented. On this basis, according to the preservation of field association features reflected by the reward feedback information, the degree of matching of each mask training strategy with the ticket field structure expression in the current training round is evaluated, thereby forming a judgment on its learning benefit level.

[0064] Through the above evaluation process, the reward feedback information is transformed into a scenario reinforcement evaluation result to characterize the effectiveness of mask training strategies for different feature dimensions. Based on this scenario reinforcement evaluation result, the adjustment direction of model parameters in subsequent training stages is determined, so that the parameter update process is consistent with the field structure expression requirements in the invoice scenario.

[0065] For example, in a training epoch, mask training is performed for text feature dimensions, image feature dimensions, and the combination of the two. Corresponding reward feedback information is generated and associated with each mask training strategy. If the reward feedback information for a certain feature dimension indicates that the field order, position, or correspondence is consistent with the existing structure of the ticket, the mask training strategy is determined to have a high learning gain in the current epoch, and this is used to strengthen the update weights or adjust the priority of the corresponding parameters in subsequent training. Conversely, if the reward feedback information reflects a large deviation in the field structure, its learning gain is determined to be low, and this is used to weaken the update weights or adjust the priority of the corresponding parameters in subsequent training. Based on the above judgments, a scene reinforcement evaluation result is formed, which guides the direction of parameter adjustment in subsequent training, making model parameter updates more conducive to maintaining the stable expression of the ticket field structure.

[0066] In this embodiment, in step S401, the deviation in the representation of the bill scene is determined based on the degree to which the mask training results of each feature dimension preserve the field-related features, thereby enabling the effectiveness of each mask training process in the bill scene to be expressed on a uniform scale. In step S402, the learning gains of the mask training strategies for each feature dimension are comprehensively evaluated based on the reward feedback information corresponding to the deviation, thereby obtaining a scene reinforcement evaluation result reflecting the adaptability of different mask training strategies in the bill scene. Based on this, in the entire technical solution, by transforming the mask training results into comparable and feedback-enabled scene evaluation information, the model training process can be guided around the stable representation of the bill field structure, thereby improving the overall coordination and effectiveness of multi-feature dimension training.

[0067] In an exemplary embodiment, the field recognition model is used as the trained target field recognition model until the consistency of the field recognition model across different feature dimensions in the invoice scenario meets a preset condition, including steps S501 to S502.

[0068] Step S501: Obtain the field recognition results of the field recognition model in the current training round, and based on the field expression of each field in the text feature dimension, image feature dimension and combination dimension in the field recognition results, perform structured processing on the field expression relationship of each feature dimension to obtain the cross-dimensional consistency index.

[0069] For example, after obtaining the field recognition results of the field recognition model in the current training round, the expression of each field across different feature dimensions is organized based on the field recognition results, rather than directly judging the correctness of the recognition. Specifically, for the same field, the text expression positional relationship corresponding to it in the text feature dimension, the layout area positional relationship corresponding to it in the image feature dimension, and the correspondence between the text and the visual area reflected in the combination dimension are extracted respectively, and the above information is uniformly mapped to the same field level for alignment processing; thereby, the field expression originally scattered across different feature dimensions can be summarized in a structured form, thus eliminating the incomparability caused by differences in expression forms.

[0070] After structuring, the field representations of each field under different feature dimensions are merged to comprehensively characterize the positional correspondence, belonging consistency, and structural coherence of fields across different feature dimensions. Based on this, a cross-dimensional consistency index is formed to characterize the stability of the model across feature dimensions in the current training round. This cross-dimensional consistency index is used to centrally reflect the overall coordination state of field recognition results across multiple feature dimensions, providing a clear and unified quantitative basis for subsequent judgment on whether the model training has met the preset conditions.

[0071] For example, in the same document, a certain field corresponds to a fixed text position in the text feature dimension, a specific visual area in the layout in the image feature dimension, and a stable correspondence between the text position and the visual area in the combined dimension. By merging the field expressions under the above different feature dimensions, it can be determined whether the field points to the same position in different feature dimensions, whether it belongs to the same field category, and whether the overall structure remains consistent, and a field-level consistency index is formed to characterize the field. Furthermore, by calculating the average, weighted average, or other statistical methods of each field-level consistency index, the field-level consistency indices corresponding to each field are summarized to obtain a cross-dimensional consistency index that reflects the overall cross-dimensional stability of the model in the current training round.

[0072] Step S502: If the cross-dimensional consistency index meets the preset threshold condition, then the field recognition model of the current training round is used as the trained target field recognition model.

[0073] For example, comparing the cross-dimensional consistency index with a pre-set threshold condition allows for a unified measurement of whether the model's field representations across text feature dimensions, image feature dimensions, and the combined dimension of the two have reached a stable level. Essentially, this threshold condition limits the minimum consistency the model should maintain for the same field recognition results across different feature dimensions, thus preventing situations where the model performs stably in one feature dimension but still deviates significantly in other dimensions.

[0074] When the cross-dimensional consistency index reaches or exceeds the threshold indicated by the threshold condition, it indicates that the field recognition model has been able to form a stable correspondence between field position, field affiliation and structural relationship under multiple feature dimensions in the current training round. At this time, it is no longer necessary to correct the expression differences between different feature dimensions through further training. Therefore, the field recognition model under the current training round is determined as the trained target field recognition model.

[0075] Optionally, in the early stages of model training, statistical analysis is performed on the consistency of field representation across different feature dimensions based on historical invoices. The distribution range of field-level consistency indices for each field during the stable training phase is calculated, and the lower limit or mean of this distribution is used as a threshold reference for the overall cross-dimensional consistency index. During actual training, when the overall field-level consistency index of each field reaches or exceeds this threshold reference in the current training round, the model is considered to have formed a stable representation across multiple feature dimensions; otherwise, training continues. This method of determining the threshold based on statistical features ensures that the threshold conditions are derived from the actual data distribution and reflect the model's stable recognition level in invoice scenarios.

[0076] In this embodiment, in step S501, based on the field recognition results of the field recognition model in the current training round, the field expression relationships of each field in the text feature dimension, image feature dimension, and combination dimension are structurally merged, so that the expression state of the same field in different feature dimensions is comparable, and a cross-dimensional consistency index is further formed to characterize the consistency of the cross-feature dimension expression of each field. In step S502, by comparing the cross-dimensional consistency index with a preset threshold condition, it can be determined whether the field expression of the field recognition model in the current training round has reached a stable and consistent state in multiple feature dimensions, and if the threshold condition is met, the current model is determined as the trained target field recognition model. Based on this, in the entire technical solution, by using the cross-feature dimension consistency of the field as the training termination criterion, the model training process aims at multi-dimensional stable expression, ensuring that the finally obtained field recognition model has a consistent and stable field recognition capability across feature dimensions.

[0077] In an exemplary embodiment, obtaining the field recognition result of the field recognition model in the current training round includes steps S601 to S602; the method further includes step S603.

[0078] Step S601: Obtain the initial field recognition results of the field recognition model in the current training round.

[0079] The initial field recognition results are derived from the field outputs generated by the model when processing historical invoices during training. They reflect the model's ability to recognize various fields, including fields such as date, amount, and invoice number. Their recognition status represents the model's performance in the current training round.

[0080] Step S602: In the preset verification structure, the initial field recognition result is jointly verified according to the business structure of a single field and the business logic of multiple fields in the invoice scenario, and the initial field recognition result that passes the verification is used as the field recognition result of the field recognition model in the current training round.

[0081] The verification structure represents a framework used to verify whether the initial field recognition results conform to the business structure and business logic in the invoice, ensuring the semantic and structural rationality of the fields.

[0082] Among them, the business structure of a single field represents the format and expression rules of the field in the invoice, which is used to ensure that the field content meets the expected format requirements. For example, the "amount" field usually needs to contain numbers and decimal points, while the "date" field must follow a specific date format (such as "YYYY-MM-DD", which means "four-digit year-two-digit month-two-digit date").

[0083] Among them, the business logic of multiple fields represents the interrelationships and constraints between multiple fields, which is used to ensure the logical consistency of field content. For example, the value of the "Amount" field should be consistent with the calculation result of the "Total" field, or the "Amount Payable" field should be equal to the sum of the "Goods Amount" and "Taxes" fields.

[0084] For example, based on the preset verification structure, the initial field recognition results are jointly verified for the business structure of a single field and the business logic of multiple fields in the invoice scenario. After this process, only the field recognition results that pass the verification will be considered as valid recognition outputs and will be used as the final field recognition results of the model in this training round. If the field recognition results of some fields fail the verification, they will be considered invalid, and the model will not use these field recognition results for subsequent processing. Thus, erroneous field recognition results are filtered out to improve the overall performance and stability of the model.

[0085] Step S603: Based on the joint verification process of the initial field recognition results, optimize the parameters of the verification structure to obtain an optimized verification structure, which is used to verify the fields identified by the field recognition model during the training and deployment processes.

[0086] For example, the feedback during the joint validation process is analyzed to identify which fields failed the validation or which field identifications deviated from the expected business structure or logic. Based on this information, the relevant parameters in the validation structure are optimized to make the validation structure more adaptable to the characteristics of the current training data and improve its ability to validate field identification results.

[0087] Through the optimization process described above, the verification structure can more effectively capture and correct recognition biases during model training, thereby improving the accuracy and stability of the model in subsequent training and practical applications. Based on this, the optimized verification structure will be used to verify the field recognition results of the field recognition model during training and deployment, ensuring that the model's output conforms to the preset business structure and logic.

[0088] For example, during a joint validation process, the "Amount" field failed validation because its format on some documents was inconsistent with expectations, or there was a deviation in the logical relationship between the field and related fields. In this case, by analyzing the fields that failed validation and the specific deviations, the validation rules in the validation structure can be adjusted and optimized. For instance, the decimal place limit for the amount field can be relaxed, or the calculation relationship between the amount and tax fields can be modified, so that the validation structure can better adapt to the characteristics of the current training data.

[0089] Furthermore, the optimization of the verification structure is not limited to the model training phase, but is continuously carried out during both the training and deployment processes. During iterative training, the verification structure continuously adjusts its parameters based on the field deviations exposed during joint verification, gradually aligning it with the business structure and logic in the training data. After model deployment, the verification structure continues to verify the actual field recognition results and updates its parameters based on verification feedback accumulated over long-term operation. This enables the verification structure to adapt to changes in field representation in real-world invoice scenarios, continuously ensuring the stability and consistency of field recognition results at the business level.

[0090] In this embodiment, in steps S601 and S602, based on a preset verification structure, the initial field recognition results of the current round are jointly verified in the context of the single-field business structure and multi-field business logic in the invoice scenario. This ensures the semantic and structural consistency of the fields and determines the field recognition results that pass the verification. In step S603, based on the feedback from the joint verification process, the parameters of the verification structure are optimized to improve the effectiveness and adaptability of the verification structure in verifying field recognition results during subsequent training and deployment. Therefore, in the entire technical solution, the application and optimization of the verification structure ensures the consistency of field recognition results in business structure and business logic, thereby improving the stability and accuracy of the model.

[0091] In an exemplary embodiment, after using the field recognition model of the current training round as the trained target field recognition model, the method further includes steps S701 to S703.

[0092] Step S701: Using the target field recognition model as a general field recognition model applicable to all types of bills, and based on the bill structure of different bill types, perform type association on the cross-dimensional consistency index of the general field recognition model in the last training round to obtain the type association index corresponding to each bill type.

[0093] For example, after determining the trained target field recognition model, the target field recognition model is used as a general field recognition model applicable to multiple types of bill scenarios. On this basis, based on the differences in bill structure of different bill types, the cross-dimensional consistency index formed by the general field recognition model in the last training round is processed by type association. That is, according to the structural characteristics of different bill types in terms of the number of fields, the order of field arrangement and the combination of fields, the cross-dimensional consistency index is decomposed and reorganized so that it can reflect the stability of the field expression of the model under various bill structures.

[0094] Through the above type association process, the cross-dimensional consistency index, which was originally used to judge the convergence of the overall training, is mapped to the structural context corresponding to different invoice types, thereby obtaining the type association index corresponding to each invoice type. This type association index is used to represent the adaptability of the general field recognition model when facing different invoice types, and provides a basis for distinguishing the differences in model-level requirements of different invoice types.

[0095] For example, for invoices, which have a large number of fields and a fixed field order, the part related to the stability of field order and position correspondence is extracted from the cross-dimensional consistency index. For receipts, whose field combinations are relatively simple and some fields can be omitted, the focus is on extracting the part related to the existence of fields and the stability of their combination relationships. By decomposing and reorganizing the consistency index according to the structural characteristics of different document types, each type of document obtains a type-related index that reflects the stability of field expression in the model under that document structure.

[0096] Step S702: Based on the type association indicators corresponding to each invoice type, select the fine-tuning part of the general field recognition model to obtain the model fine-tuning method corresponding to each invoice type.

[0097] For example, by analyzing the stability of field expressions reflected by the type association indicators corresponding to each type of invoice, the differences in the degree of matching of the field structure of each part of the general field recognition model when facing different invoice types can be identified. For the part of the model that shows high consistency under a certain invoice type, its parameter state is kept unchanged, so that it continues to be retained as a general capability; while for the part of the model that shows insufficient consistency under that invoice type, it is identified as the fine-tuning part corresponding to that invoice type.

[0098] Because different types of invoices differ in the number of fields, the arrangement of fields, and the rules for combining fields, the fine-tuning parts corresponding to different invoice types are not the same in this analysis process. Instead, they point to the parts of the model that are most relevant to their structural differences.

[0099] For example, for invoices, which have a large number of fields and a relatively fixed order, the fine-tuning mainly focuses on the model parts that are strongly related to the order and position of the fields. For receipts, which have a relatively small number of fields and some fields can be omitted, the corresponding fine-tuning focuses more on the model parts that are related to the existence and combination of fields.

[0100] Based on this, through the above-mentioned differentiation and selection process, each type of bill forms a set of model fine-tuning parts that match its own structural characteristics, thereby avoiding the use of a uniform fine-tuning strategy to cover all bill types. After completing the above processing, corresponding model fine-tuning methods are obtained for each bill type, thus providing a clear basis for implementing differentiated fine-tuning according to bill type.

[0101] Step S703: Based on the model fine-tuning method corresponding to each type of invoice, the general field recognition model is fine-tuned separately to obtain the special field recognition models applicable to the corresponding invoice types.

[0102] For example, for each type of invoice, based on its corresponding model fine-tuning method, only the parameters of the selected fine-tuning part in the general field recognition model are adjusted, while the rest of the model parts that maintain general capabilities remain unchanged, thereby ensuring that the model can adapt to a specific invoice type without destroying the established overall recognition capability.

[0103] During the fine-tuning process, training data matching the specific invoice type is introduced, allowing the selected fine-tuning portion to gradually align with the structural characteristics of this type of invoice in terms of the number of fields, their order, and their combinations through repeated iterations. As the fine-tuning progresses, the model's field representation for the corresponding invoice type gradually stabilizes, forming a field recognition pattern consistent with that type of invoice. After completing the fine-tuning process, the adjusted model is saved as a dedicated field recognition model suitable for this invoice type.

[0104] Based on this, by performing the above-mentioned special fine-tuning process on different types of invoices, the same general field recognition model is transformed into multiple special field recognition models. Each special field recognition model is structurally adapted to its corresponding invoice type, so that in practical applications, it can provide field recognition capabilities that are more in line with the structural characteristics of different invoice types.

[0105] In this embodiment, in step S701, the cross-dimensional consistency index formed by the general field recognition model in the last training round is associated with the bill structure of different bill types to obtain various type association indices, thereby distinguishing the stability of field expression under different bill types. In step S702, based on the various type association indices, the fine-tuning parts of the general field recognition model that need to be adapted to the bill type are selected, so that different bill types correspond to different model fine-tuning methods. In step S703, based on the model fine-tuning methods corresponding to each bill type, targeted special fine-tuning is performed on the general field recognition model to obtain special field recognition models applicable to different bill types. Based on this, in the entire technical solution, by introducing a fine-tuning mechanism that distinguishes by bill type while maintaining the model's generality, the field recognition model can simultaneously take into account both generality and type adaptability, improving the overall stability and accuracy of field recognition in multiple bill scenarios.

[0106] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0107] Based on the same inventive concept, this application also provides a field recognition model training system for invoices to implement the above-described method for training a field recognition model for invoices. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the field recognition model training system for invoices provided below can be found in the limitations of the field recognition model training method for invoices described above, and will not be repeated here.

[0108] In one exemplary embodiment, such as Figure 2 As shown, a field recognition model training system for invoices is provided, including: an annotation module 201, a training module 202, and a generation module 203, wherein: The annotation module 201 is used to annotate fields of preset historical tickets to obtain text training sets and image training sets. The training module 202 is used to input the text training set and the image training set into the field recognition model to be trained, to perform dynamic mask training on the text training set and the image training set on different feature dimensions of the bill scene, and to iteratively adjust the parameters of the field recognition model in combination with the reinforcement learning mechanism corresponding to the bill scene. The generation module 203 is used to take the field recognition model as the trained target field recognition model when the consistency of the field recognition model in different feature dimensions of the invoice scenario meets the preset conditions, so as to extract and recognize fields for any type of invoice.

[0109] The modules in the aforementioned training system for the field recognition model of invoices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0110] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above embodiments.

[0111] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above embodiments.

[0112] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0113] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0114] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for training a field recognition model for invoices, characterized in that, The method includes: Fields are labeled on multiple types of historical tickets to obtain text training sets and image training sets; The text training set and the image training set are input into the field recognition model to be trained. Dynamic mask training is performed on the text training set and the image training set on different feature dimensions of the bill scene. Combined with the reinforcement learning mechanism corresponding to the bill scene, the parameters of the field recognition model are iteratively adjusted. Once the consistency of the field recognition model across different feature dimensions in the invoice scenario meets the preset conditions, the field recognition model will be used as the trained target field recognition model for extracting and recognizing fields for any type of invoice.

2. The method according to claim 1, characterized in that, The step of dynamically masking the text training set and the image training set across different feature dimensions of the invoice scenario, and iteratively adjusting the parameters of the field recognition model by combining a reinforcement learning mechanism corresponding to the invoice scenario, includes: Text mask training is performed on the text training set in the text feature dimension of the ticket scene, and image mask training is performed on the image training set in the image feature dimension of the ticket scene. Composite mask training is performed on the text training set and the image training set in the combined dimension of text features and image features of the ticket scene, so as to obtain the mask training results of each feature dimension. Based on the reinforcement learning mechanism corresponding to the bill scenario, the reinforcement evaluation based on field association features is performed on the mask training results of each feature dimension to obtain the scenario reinforcement evaluation results used to characterize the effectiveness of each mask training process. Based on the scene adaptability evaluation results of each training round, the parameters of the field recognition model are iteratively adjusted.

3. The method according to claim 2, characterized in that, The process involves training text masks on the text training set based on the text feature dimension of the ticket scene, training image masks on the image training set based on the image feature dimension of the ticket scene, and training composite masks on the text training set and the image training set based on the combined dimension of text and image features of the ticket scene, to obtain mask training results for each feature dimension, including: In the text feature dimension of the bill scenario, any field text in the text training set is masked, and the text content of the masked field text is predicted based on the semantic distribution of the unmasked field text in the bill structure, so as to obtain the mask training result of the text feature dimension. In the image feature dimension of the ticket scene, any visual region in the image training set is masked, and the image content of the masked visual region is predicted based on the visual layout of the unmasked visual region in the ticket layout, so as to obtain the mask training result of the image feature dimension. In the dimension of combining text features and image features in the ticket scenario, any field of text in the text training set is masked, and by judging whether the visual region in the same position of the image training set is masked, the cross-modal alignment state of the masked field text and the corresponding visual region in the same position is predicted, so as to obtain the mask training result of the combined dimension.

4. The method according to claim 2, characterized in that, The step involves performing reinforcement evaluation based on field-related features on the mask training results of each feature dimension according to the reinforcement learning mechanism corresponding to the ticket scenario, to obtain scenario reinforcement evaluation results that characterize the effectiveness of each mask training process, including: Based on the degree to which the mask training results of each feature dimension preserve the field-related features, the deviation of each mask training result in the representation of the bill scenario is determined, and reward feedback information is generated based on the corresponding deviation to indicate the effectiveness of the corresponding mask training process. Based on the reward feedback information, the learning benefits of the mask training strategy for each feature dimension in the ticket scenario are evaluated to obtain the scenario enhancement evaluation results.

5. The method according to claim 1, characterized in that, The step of using the field recognition model as the trained target field recognition model until the consistency of the field recognition model across different feature dimensions in the invoice scenario meets a preset condition includes: Obtain the field recognition results of the field recognition model in the current training round, and based on the field expression of each field in the text feature dimension, image feature dimension and combination dimension in the field recognition results, perform structured processing on the field expression relationship of each feature dimension to obtain the cross-dimensional consistency index. If the cross-dimensional consistency index meets the preset threshold condition, then the field recognition model of the current training round will be used as the trained target field recognition model.

6. The method according to claim 5, characterized in that, The step of obtaining the field recognition result of the field recognition model in the current training round includes: Obtain the initial field recognition results of the field recognition model in the current training round; In the preset verification structure, the initial field recognition result is jointly verified according to the business structure of a single field and the business logic of multiple fields in the bill scenario, and the initial field recognition result that passes the verification is used as the field recognition result of the field recognition model in the current training round. The method further includes: Based on the joint verification process of the initial field recognition results, the parameters of the verification structure are optimized to obtain an optimized verification structure, which is used to verify the fields identified by the field recognition model during the training and deployment processes.

7. The method according to claim 5, characterized in that, After using the field recognition model of the current training round as the trained target field recognition model, the method further includes: The target field recognition model is used as a general field recognition model applicable to all types of bills. Based on the bill structure of different bill types, the cross-dimensional consistency index of the general field recognition model in the last training round is type-associated to obtain the type association index corresponding to each bill type. Based on the type association indicators corresponding to each invoice type, the fine-tuning part of the general field recognition model is selected to obtain the model fine-tuning method corresponding to each invoice type; Based on the model fine-tuning method corresponding to each type of invoice, the general field recognition model is fine-tuned separately to obtain special field recognition models applicable to the corresponding invoice types.

8. A training system for a field recognition model for invoices, characterized in that, The system includes: The annotation module is used to annotate fields of multiple preset types of historical tickets to obtain text training sets and image training sets. The training module is used to input the text training set and the image training set into the field recognition model to be trained, perform dynamic mask training on the text training set and the image training set on different feature dimensions of the invoice scenario, and iteratively adjust the parameters of the field recognition model in combination with the reinforcement learning mechanism corresponding to the invoice scenario. The generation module is used to select the field recognition model as a trained target field recognition model when the consistency of the field recognition model across different feature dimensions in the invoice scenario meets a preset condition, so as to extract and recognize fields for any type of invoice.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.