Information extraction model training method and device, storage medium and electronic equipment
By generating virtual documents and training in a multimodal large model, the problem of low accuracy of information extraction models due to difficulty in obtaining data in the prior art is solved, and efficient and accurate information extraction model training is achieved.
Patent Information
- Application Number
- CN202510166762.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The existing training methods of information extraction models rely on a large amount of real picture-text pairing data, which is restricted by privacy and regulations, and are difficult to obtain on a large scale, resulting in low model accuracy.
By creating the original structural information and script code of the virtual certificate, the virtual certificate is generated and image enhancement processing is performed, input into the multimodal large model for training, constructing the target loss function and fine-tuning the model parameters to obtain a high-precision information extraction model.
The high-precision information extraction model can be trained without real document data, which solves the problem of data acquisition difficulties and improves the model generation efficiency.
Smart Images

Figure CN120107989A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the technical field of image data processing, and in particular, to a training method for an information extraction model, a training device for an information extraction model, a computer-readable storage medium, and an electronic device. Background Art
[0002] The existing training methods for information extraction models require a large amount of image-text pairing data; however, real document data is difficult to obtain on a large scale due to privacy and regulatory restrictions, which results in low accuracy of the resulting information extraction model.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0004] The purpose of the present disclosure is to provide a training method for an information extraction model, a training device for an information extraction model, a computer-readable storage medium and an electronic device, thereby at least to a certain extent overcoming the problem of low accuracy of the information extraction model caused by the limitations and defects of the relevant technology.
[0005] According to one aspect of the present disclosure, a method for training an information extraction model is provided, comprising:
[0006] Creating original certificate structure information required for generating a first original virtual certificate, and creating script code required for generating the first original virtual certificate according to the original certificate structure information;
[0007] generating the first original virtual certificate based on the script code, and performing image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate;
[0008] Inputting the first standard virtual certificate into the multimodal large model to obtain a first certificate information prediction result, and constructing a target loss function according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate;
[0009] Based on the objective loss function, the low-rank matrix parameters of the multimodal large model in the document information extraction scenario are determined, and the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain an information extraction model.
[0010] In an exemplary embodiment of the present disclosure, creating original certificate structure information required for generating a first original virtual certificate includes:
[0011] Create the name of the field required for generating the first original virtual certificate, the position of the field in the first original virtual certificate, the style of the field in the first original virtual certificate, and the color information, texture information, and image information required for generating the first original virtual certificate;
[0012] The original certificate structure information required for generating the first original virtual certificate is created according to the name, position, style, color information, texture information and image information of the fields.
[0013] In an exemplary embodiment of the present disclosure, the script code required for generating the first original virtual certificate is created according to the original certificate structure information, including:
[0014] Determine, according to the position of the field, the name of the field and a field insertion position code of a field value corresponding to the name of the field in the first original virtual certificate;
[0015] Determine the name of the field and the font code, font size code and text color code of the field value in the first original virtual certificate according to the style of the field;
[0016] Determine a background color code of the first original virtual certificate according to the texture information, and determine an image insertion position code of the image information in the first original virtual certificate;
[0017] The script code required for generating the first original virtual certificate is created according to the field insertion position code, font code, font size code, text color code, background color code and image insertion position code.
[0018] In an exemplary embodiment of the present disclosure, generating the first original virtual certificate based on the script code includes:
[0019] Determining a type of the field based on the name of the field, and generating a random field value corresponding to the type of the field;
[0020] Verifying the legitimacy of the random field value according to the type of the field, and after the legitimacy verification passes, performing diversified processing on the name of the field and the language type of the random field value to obtain names and field values in multiple different language types;
[0021] The names and field values in a plurality of different language types and the script code are compiled to obtain the first original virtual certificate.
[0022] In an exemplary embodiment of the present disclosure, performing image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate includes:
[0023] Performing noise processing and / or blur processing and / or compression distortion processing on the first original virtual certificate to obtain a virtual certificate with different noise and / or different definition and / or different pixels; and / or
[0024] The positions of the fields and / or the font styles of the fields included in the first original virtual certificate are randomly adjusted to obtain a virtual certificate with a different position and / or a different font style.
[0025] In an exemplary embodiment of the present disclosure, the multimodal large model includes a modality encoder, an input projector, a large language model, an output projector, and a modality generator;
[0026] The first standard virtual certificate is input into the multimodal large model to obtain the first certificate information prediction result, including:
[0027] Encoding the first standard virtual certificate based on the modality encoder to obtain a first image feature;
[0028] Performing feature mapping processing on the first image feature based on the input projector to obtain a first feature mapping result;
[0029] Performing feature extraction on the first feature mapping result based on the large language model to obtain a first feature extraction result;
[0030] Performing feature conversion processing on the first feature extraction result based on the output projector to obtain a first feature conversion result;
[0031] The first feature conversion result is modally converted based on the modality generator to obtain the first certificate information prediction result.
[0032] In an exemplary embodiment of the present disclosure, the large language model includes an embedding mapping layer, a first encoding layer, and a hybrid expert model layer;
[0033] The step of performing feature extraction on the first feature mapping result based on the large language model to obtain a first feature extraction result includes:
[0034] Acquire preset information classification prompt information and preset modality fusion prompt information corresponding to the first standard virtual certificate, and generate context information to be predicted according to the preset information classification prompt information and the preset modality fusion prompt information;
[0035] Performing embedding mapping processing on the first feature mapping result based on the embedding mapping layer to obtain a second image feature, and performing embedding mapping processing on the context information to be predicted based on the embedding mapping layer to obtain a first context marker sequence;
[0036] Based on the first coding layer, the second image feature and the first context marker sequence are encoded to obtain a first context overall representation, and based on the hybrid expert model layer, the first context marker sequence and the first context overall representation are predicted to obtain the first feature extraction result.
[0037] In an exemplary embodiment of the present disclosure, the hybrid expert model layer includes a first gating network model, a second gating network model, and a plurality of expert neural network models;
[0038] The first feature extraction result is obtained by predicting the first context marker sequence and the first context overall representation based on the hybrid expert model layer, including:
[0039] Determine, based on the first gating network model and the first context flag sequence, a first model weight of the expert neural network model in the image feature extraction dimension, and determine, from the expert neural network according to the first model weight, a first target neural network model required to perform the information prediction task in the image feature extraction dimension;
[0040] Determine a second model weight of the expert neural network model on the text feature extraction dimension based on the second gating network model according to the first contextual marker sequence, and determine a second target neural network model required to perform the information prediction task on the text feature extraction dimension from the expert neural network according to the second model weight;
[0041] The first context marker sequence and the first context overall representation are input into the first target neural network model to obtain image feature prediction results in the image feature extraction dimension, and the first context marker sequence and the first context overall representation are input into the second target neural network model to obtain text feature prediction results in the text feature extraction dimension, so as to determine the first feature extraction result based on the image feature prediction results and the text feature prediction results.
[0042] In an exemplary embodiment of the present disclosure, the multimodal large model is fine-tuned based on low-rank matrix parameters to obtain an information extraction model, including:
[0043] Fine-tune the multimodal large model based on the low-rank matrix parameters to obtain an adjusted large model, and extract information from the second original virtual certificate based on the adjusted large model to obtain a second certificate information prediction result;
[0044] Calculating a character edit distance between the second certificate information prediction result and the second actual certificate information corresponding to the second original virtual certificate, and determining a first character accuracy rate of the second certificate information prediction result according to the character edit distance;
[0045] Determine a field accuracy rate in the second certificate information prediction result according to the character edit distance, and determine a second character accuracy rate of the second certificate information prediction result according to the field accuracy rate;
[0046] The model accuracy of the adjusted large model is determined according to the first character accuracy and the second character accuracy, and when it is determined that the model accuracy is greater than or equal to a preset threshold, the adjusted large model is used as the information extraction model.
[0047] According to one aspect of the present disclosure, there is provided a training device for an information extraction model, comprising:
[0048] A script code creation module, used to create original certificate structure information required to generate the first original virtual certificate, and to create the script code required to generate the first original virtual certificate according to the original certificate structure information;
[0049] An image enhancement processing module, used to generate the first original virtual certificate based on the script code, and perform image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate;
[0050] A loss function construction module, used for inputting the first standard virtual certificate into the multimodal large model to obtain a first certificate information prediction result, and constructing a target loss function according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate;
[0051] A large model fine-tuning module is used to determine the low-rank matrix parameters of the multimodal large model in the document information extraction scenario based on the target loss function, and fine-tune the multimodal large model based on the low-rank matrix parameters to obtain an information extraction model.
[0052] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the information extraction model described in any one of the above is implemented.
[0053] According to one aspect of the present disclosure, there is provided an electronic device, including:
[0054] Processor; and
[0055] A memory, configured to store executable instructions of the processor;
[0056] Wherein, the processor is configured to execute any one of the above-mentioned information extraction model training methods by executing the executable instructions.
[0057] The disclosed embodiment provides a training method for an information extraction model. On the one hand, the original certificate structure information required for generating a first original virtual certificate is created, and the script code required for generating the first original virtual certificate is created based on the original certificate structure information; then the first original virtual certificate is generated based on the script code, and the image enhancement processing is performed on the first original virtual certificate to obtain a first standard virtual certificate; the first standard virtual certificate is then input into a multimodal large model to obtain a first certificate information prediction result, and a target loss function is constructed based on the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate; finally, the multimodal large model is determined based on the target loss function in the certificate. The low-rank matrix parameters in the document information extraction scenario are obtained, and the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain the information extraction model; since virtual documents can be constructed for model training, virtual documents of corresponding types and numbers can be constructed according to actual needs, and the model training can be realized without obtaining real document data, thereby solving the problem of low accuracy of the information extraction model obtained due to the inability to obtain sufficient document data in the prior art, and improving the accuracy of the obtained information extraction model; on the other hand, since the information extraction model can be obtained by fine-tuning the multimodal large model, there is no need to train from scratch, thereby improving the generation efficiency of the information extraction model.
[0058] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0060] Figure 1 A flowchart of a method for training an information extraction model according to an example embodiment of the present disclosure is schematically shown.
[0061] Figure 2 An example diagram of a hierarchical structure of a multimodal large model according to an example embodiment of the present disclosure is schematically shown.
[0062] Figure 3 A diagram schematically shows an example structure of a large language model in a multimodal large model according to an example embodiment of the present disclosure.
[0063] Figure 4 The following schematically shows an example structure diagram of a hybrid expert model layer in a large language model according to an example embodiment of the present disclosure.
[0064] Figure 5 A diagram schematically shows an example scenario of a first original virtual certificate obtained according to an example embodiment of the present disclosure.
[0065] Figure 6 A diagram schematically shows an example structure of an input projector in a multimodal large model according to an example embodiment of the present disclosure.
[0066] Figure 7 A diagram schematically shows an example structure of a training device for an information extraction model according to an example embodiment of the present disclosure.
[0067] Figure 8 An electronic device for implementing a training method for an information extraction model according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0068] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, the known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0069] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0070] The extraction of certificate information refers to extracting the key information of the certificate holder from mixed text and image data in various certificate formats; the certificates in different certificate formats recorded here may include but are not limited to identity cards, driver's licenses, passports, and business licenses; the key information recorded here may include but is not limited to name, gender, certificate number, and validity period, etc. In the actual process of extracting certificate information, there are the following challenges: on the one hand, template style diversity: that is, different types of certificates vary greatly in overall layout, fonts, colors, etc., so it is impossible to extract information from different certificates based on the same extraction method; on the other hand, rule design complexity; that is, traditional methods need to design corresponding parsing rules for each certificate separately, and it is difficult to uniformly handle all styles, which makes the extraction efficiency of certificate information low; on the other hand, multi-language support challenges; that is, certificate information usually involves multiple languages, and certificate information in different regions not only has layout differences, but also involves different languages and character sets (such as Chinese, English, Arabic, etc.); especially in international scenarios, it is difficult to support small languages.
[0071] Traditional methods for extracting certificate information usually rely on template matching, rule engines or deep learning models to achieve this. In the specific process of extracting certificate information, optical character recognition (OCR) technology and deep learning models are usually combined to achieve certificate information extraction through the following steps: First, use OCR to recognize the text content in the certificate; Second, use a rule engine or a specific deep learning model to post-process the recognition results to achieve the purpose of extracting key information; Although information extraction based on this method can be applied to common templates and the effect is relatively stable, it can achieve high accuracy with the support of large-scale samples; However, it also has the following disadvantages: On the one hand, it has strong template dependence; that is, it takes a lot of resources to build a template library, and each new certificate template requires the design of rules or fine-tuning of models, which has poor scalability; On the other hand, it lacks multi-language support; that is, different OCR models need to be used for different languages, which increases deployment and maintenance costs; On the other hand, it has a large demand for data; that is, training high-performance models requires a large amount of labeled data, and the cost of data collection and labeling is extremely high, especially in minority language scenarios.
[0072] Under this premise, the relevant scheme proposes a method for extracting certificate information based on a large multimodal model. Specifically, with the development of deep learning, especially large multimodal models (LMM), its performance in natural language processing, image recognition and other fields is getting better and better; in the actual information extraction process, large multimodal models combine the understanding of vision and language, and can process image and text information at the same time. Its advantages in the field of information extraction include but are not limited to the following aspects: on the one hand, it has end-to-end information extraction capabilities; that is, large multimodal models can achieve the extraction of key information from images without the help of additional OCR modules, and thus can achieve efficient processing; on the other hand, it has strong cross-language capabilities; that is, multimodal models can achieve the purpose of supporting multilingual scenarios through large-scale pre-training; on the other hand, it uniformly processes various templates; that is, information extraction based on large multimodal models does not require separate design of rules for each certificate style, and the model has strong versatility.
[0073] However, although large multimodal models have strong versatility and cross-language capabilities, they still face the following challenges in practical applications: on the one hand, data acquisition is difficult; that is, since model training requires a large amount of image-text pairing data, real document data is difficult to obtain on a large scale due to privacy and regulatory restrictions, so there is a problem of difficulty in data acquisition; on the other hand, support for small languages is insufficient; that is, there is a lack of corresponding labeled data in small language scenarios, and model performance is difficult to guarantee; on the other hand, the training cost is high; that is, the training of multimodal models requires a lot of computing resources, which further increases the optimization cost.
[0074] Based on this, this exemplary embodiment first provides a training method for an information extraction model, which can be run on a terminal device, a server, a server cluster or a cloud server, etc. Of course, those skilled in the art can also run the method disclosed in this disclosure on other platforms as required, and this exemplary embodiment does not specifically limit this. Figure 1 As shown, the training method of the information extraction model may include the following steps:
[0075] Step S110. Create original certificate structure information required for generating the first original virtual certificate, and create script code required for generating the first original virtual certificate according to the original certificate structure information;
[0076] Step S120: generating the first original virtual certificate based on the script code, and performing image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate;
[0077] Step S130: Input the first standard virtual certificate into the multimodal large model to obtain a first certificate information prediction result, and construct a target loss function according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate;
[0078] Step S140. Determine the low-rank matrix parameters of the multimodal large model in the document information extraction scenario based on the target loss function, and fine-tune the multimodal large model based on the low-rank matrix parameters to obtain an information extraction model.
[0079] In the training method of the information extraction model recorded above, the training method of the information extraction model, on the one hand, creates the original certificate structure information required to generate the first original virtual certificate, and creates the script code required to generate the first original virtual certificate based on the original certificate structure information; then generates the first original virtual certificate based on the script code, and performs image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate; then inputs the first standard virtual certificate into the multimodal large model to obtain the first certificate information prediction result, and constructs the target loss function based on the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate; finally, determines the multimodal large model based on the target loss function. The low-rank matrix parameters of the model in the scenario of document information extraction are obtained, and the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain the information extraction model; since virtual documents can be constructed for model training, virtual documents of corresponding types and numbers can be constructed according to actual needs, and the model training can be realized without obtaining real document data, thereby solving the problem of low accuracy of the information extraction model obtained in the prior art due to the inability to obtain sufficient document data, and improving the accuracy of the obtained information extraction model; on the other hand, since the information extraction model can be obtained by fine-tuning the multimodal large model, there is no need to train from scratch, thereby improving the generation efficiency of the information extraction model.
[0080] Hereinafter, the training method of the information extraction model described in the exemplary embodiment of the present disclosure will be explained and illustrated in detail with reference to the accompanying drawings.
[0081] First, the technical implementation principle of the example embodiment of the present disclosure is explained and illustrated. Specifically, the training method of the information extraction model recorded in the example embodiment of the present disclosure aims to train a multimodal large model by synthesizing graphic samples of virtual certificates, and apply the multimodal large model in the field of certificate information extraction to improve the efficiency and accuracy of certificate information extraction. In the specific application process, by synthesizing virtual certificates and combining the capabilities of the multimodal large model, efficient extraction of multi-template and multi-language certificate information is achieved; in the actual model training process, first, a variety of certificate samples are automatically generated based on the corresponding certificate generation script, including virtual certificates in different languages and layouts; secondly, based on the generated virtual data, the multimodal large model is fine-tuned to improve the accuracy and robustness of certificate information extraction; further, by generating multilingual virtual certificates, the problem of data scarcity in small language scenarios can also be solved. Based on this solution, not only can we reduce the cost of data annotation by synthesizing a large number of diverse virtual certificates, but we can also achieve the goal of uniformly processing certificates of various different formats based on multimodal large models and diverse virtual data. Moreover, we can reduce the dependence on real data in minority languages by generating multilingual certificate data. On this basis, we can also quickly expand the method of generating virtual samples to new templates and new language scenarios, thereby improving the generalization ability of the information extraction model.
[0082] Next, the multimodal macro model described in the exemplary embodiment of the present disclosure is explained and illustrated. Figure 2 As shown, the multimodal large model may include a first input layer 210, a modality encoder 220, an input projector 230, a large language model 240, an output projector 250, a modality generator 260, and a first output layer 270; further, refer to Figure 3 As shown, the large language model described herein may include an embedding mapping layer 301, a first encoding layer 302, and a hybrid expert model layer 303; further, referring to Figure 4 As shown, the hybrid expert model layer recorded here may include a first gated network model 401, a second gated network model 402 and multiple expert neural network models 403; wherein, the specific functions performed by each functional layer included in the multimodal large model in the information extraction process will be described in detail later, and no further elaboration will be given here.
[0083] The following will be combined Figure 2-Figure 4 right Figure 1 The training method of the information extraction model shown in is further explained and illustrated. Specifically:
[0084] In step S110, original certificate structure information required for generating a first original virtual certificate is created, and script code required for generating the first original virtual certificate is created according to the original certificate structure information.
[0085] In this example embodiment, first, the original document structure information required to generate the first original virtual document is created; specifically, it can be achieved in the following ways: create the name of the field required to generate the first original virtual document, the position of the field in the first original virtual document, the style of the field in the first original virtual document, and the color information, texture information and image information required to generate the first original virtual document; according to the name, position, style, color information, texture information and image information of the field, the original document structure information required to generate the first original virtual document is created. Specifically, the first original virtual document recorded here may include but is not limited to common documents such as identity cards, driver's licenses, passports, business licenses, passes, marriage certificates, employee cards, student cards, etc. In the process of actual application, they can be selected according to actual needs, and this example does not impose special restrictions on this.
[0086] In a possible example embodiment, in the process of creating the original certificate structure information, first, it is necessary to define the structure information of different types of certificates through a configuration file or a database; specifically, the certificate structure information recorded here may include but is not limited to field names, field positions, field styles, and certificate backgrounds, etc.; wherein, the field names recorded here may include but are not limited to information such as name, certificate number, and address. In the actual application process, different fields can be configured according to the type of certificate, and this example does not impose special restrictions on this; the field position recorded here refers to the placement position of the field and the field value corresponding to the field in the certificate; the field style recorded here refers to the font, font size, and font color, etc. of the field and the field value corresponding to the field in the certificate; the certificate background recorded here refers to the background color, background texture, and pattern of the certificate (for example, a photo in an ID card or driver's license, etc.).
[0087] Secondly, the script code required for generating the first original virtual certificate is created based on the original certificate structure information; specifically, it can be achieved in the following ways: determine the name of the field and the field insertion position code of the field value corresponding to the name of the field in the first original virtual certificate based on the position of the field; determine the font code, font size code and text color code of the name of the field and the field value in the first original virtual certificate based on the style of the field; determine the background color code of the first original virtual certificate based on the texture information, and determine the image insertion position code of the image information in the first original virtual certificate; create the script code required for generating the first original virtual certificate based on the field insertion position code, font code, font size code, text color code, background color code and image insertion position code. Specifically, in the process of script code generation, first, the field content is automatically inserted according to the field position; secondly, the font, font size, color, etc. are set according to the style definition; finally, the pattern or background color, etc. are set according to the background definition; when the field content, font, font size, color, pattern or background color are all set, the corresponding script code can be obtained; wherein the obtained script code can be a JSON file or an XML file, and this example does not impose any special restrictions on this.
[0088] In step S120, the first original virtual certificate is generated based on the script code, and image enhancement processing is performed on the first original virtual certificate to obtain a first standard virtual certificate.
[0089] In this example embodiment, first, a first original virtual certificate is generated based on the script code; specifically, it can be implemented in the following manner: based on the name of the field, the type of the field is determined, and a random field value corresponding to the type of the field is generated; the legitimacy of the random field value is verified according to the type of the field, and after the legitimacy verification is passed, the name of the field and the language type of the random field value are diversified to obtain names and field values with multiple different language types; the names and field values with multiple different language types and the script code are compiled to obtain the first original virtual certificate; wherein the obtained first original virtual certificate can refer to Figure 5As shown. Specifically, in the process of generating the first original virtual certificate, in order to improve the accuracy of the obtained first original virtual certificate, it is necessary to first verify the legitimacy of the field value; specifically, in the process of verifying the legitimacy, it can be verified according to the type of the field; for example, date type, certificate number type, etc.; when verifying the legitimacy, it can be based on the number of data bits, specific data format, and specific value range of the field value of the corresponding type; if the legitimacy check passes, then perform subsequent multi-language processing or multi-format processing; if the legitimacy check fails, then regenerate the corresponding field value until the legitimacy check passes; finally, perform the corresponding compilation processing to obtain the first original virtual certificate. It should also be noted here that in the process of generating the first original virtual certificate, the field name, field value, and field position also need to be marked to ensure that the obtained first original virtual certificate is consistent with the real certificate.
[0090] In a possible example embodiment, in the process of generating random field values, it can be implemented based on the corresponding random function; for example, a random function for generating an ID number; a random function for generating an address; or a random function for generating a name, etc. The specific random function used can be defined according to actual needs; further, in the actual application process, the corresponding field can be selected according to the ID type, and the corresponding random function can be selected according to the field type. This example does not impose any special restrictions on this.
[0091] Secondly, the first original virtual certificate is subjected to image enhancement processing to obtain a first standard virtual certificate; specifically, this can be achieved in the following ways: the first original virtual certificate is subjected to noise processing and / or blur processing and / or compression distortion processing to obtain a virtual certificate with different noise and / or different clarity and / or different pixels; and / or the position of the fields and / or the font style of the fields included in the first original virtual certificate are randomly adjusted to obtain a virtual certificate with different positions and / or different font styles. Specifically, the data enhancement processing described herein may include but is not limited to image processing such as noise, blur, compression distortion, and random adjustment of field position and font style, etc.; further, by performing image enhancement processing on the first original virtual certificate, the generalization ability of the model can be further improved, thereby achieving the purpose of improving the accuracy of the extracted certificate information.
[0092] In an example embodiment, after obtaining the first standard virtual certificate, in order to adapt to the extraction of local information in the certificate and thus achieve the purpose of improving the model's recognition robustness of key information, it is also necessary to resample the first standard virtual certificate. In the specific resampling process, the first standard virtual certificate can be rotated, cropped, and noise added in a variety of ways to achieve the purpose of improving the generalization ability of the model; at the same time, it is also necessary to convert the graphic data into a format suitable for the training of the multimodal large model, so that the multimodal large model can accurately extract the certificate information and obtain the first certificate information prediction result.
[0093] In step S130, the first standard virtual certificate is input into the multimodal large model to obtain a first certificate information prediction result, and a target loss function is constructed according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate.
[0094] In this example embodiment, first, a first standard virtual certificate is input into a multimodal large model to obtain a first certificate information prediction result; specifically, this can be achieved in the following manner: encoding the first standard virtual certificate based on the modal encoder to obtain a first image feature; feature mapping the first image feature based on the input projector to obtain a first feature mapping result; feature extraction is performed on the first feature mapping result based on the large language model to obtain a first feature extraction result; feature conversion is performed on the first feature extraction result based on the output projector to obtain a first feature conversion result; modality conversion is performed on the first feature conversion result based on the modal generator to obtain the first certificate information prediction result. Specifically, the modality encoder (Modality Encoder) recorded here can be used to encode the first standard virtual document into a representation (i.e., a feature vector) that can be understood by the model. The modality encoder can be a Vision Transformer, and of course it can also be other encoders. This example does not impose any special restrictions on this. The input projector (Input Projector) recorded here can be used to project the first image features into a specific feature space so that these features can be understood by the large language model. Among them, the input projector recorded here can be a multi-layer perceptron or a cross-attention unit, etc. This example does not impose any special restrictions on this. Further, taking the input projector as a multi-layer perceptron as an example, the specific model structure can be referred to. Figure 6As shown; the output projector recorded here can be used to convert the output signal of a large language model (LLM) into a feature representation suitable for use by different modal generators; the output projector recorded here can be a Tiny Transformer or a multi-layer perceptron, and this example does not impose any special restrictions on this; the modal generator recorded here can be used to generate outputs of different modalities, such as images, videos or audio, etc.; wherein, the modal generator recorded here can be a Stable Diffusion or other generators, and this example does not impose any special restrictions on this.
[0095] In an exemplary embodiment, feature extraction is performed on the first feature mapping result based on the large language model to obtain a first feature extraction result, which can be achieved in the following manner: obtaining preset information classification prompt information and preset modal fusion prompt information corresponding to the first standard virtual certificate, and generating context information to be predicted based on the preset information classification prompt information and the preset modal fusion prompt information; embedding mapping is performed on the first feature mapping result based on the embedding mapping layer to obtain a second image feature, and embedding mapping is performed on the context information to be predicted based on the embedding mapping layer to obtain a first context marker sequence; encoding is performed on the second image feature and the first context marker sequence based on the first encoding layer to obtain a first context overall representation, and predicting the first context marker sequence and the first context overall representation based on the hybrid expert model layer to obtain the first feature extraction result. Specifically, the embedding mapping layer recorded here may include an Embedding embedding mapping layer and a Bert embedding mapping layer. In actual application, the first feature mapping result can be embedded and mapped based on the Embedding embedding mapping layer to obtain the second image feature, and the context information to be predicted can be embedded and mapped based on the Bert embedding mapping layer to obtain the first context marker sequence; the first encoding layer recorded here can be a bidirectional multi-layer Transformer, or other encoding layers can be selected, and this example does not impose any special restrictions on this.
[0096] In an example embodiment, for the task of extracting certificate information, the design of prompts is an important link; in the training of a large multimodal model, prompts will help the model understand the task objectives and strengthen the extraction of specific information. Based on this, the example embodiment of the present disclosure designs information classification prompt information and modal fusion prompt information; wherein, the information classification prompt information can be used to design specific prompts according to the task objectives to guide the model to extract specific certificate information; the information classification prompt information recorded here may include but is not limited to: extracting ID card information: "extract the name, date of birth and certificate number on the ID card"; extracting passport information: "extract the name, nationality, certificate number and validity period on the passport"; extracting driver's license information: "extract the name, gender, certificate number, and type of vehicle allowed to be driven on the driver's license". Furthermore, the modal fusion prompt information recorded here can be used to process the combination of image and text information, and design multimodal fusion prompt words to guide the model to extract the corresponding type of certificate information; the modal fusion prompt information recorded here may include but is not limited to: "Extract the certificate number and name from this picture"; "Extract the date of birth and address information based on the text content in the picture"; "This picture contains name, ID number, date of birth and other information, which are extracted and marked one by one."
[0097] In an exemplary embodiment, the first context marker sequence and the first context overall representation are predicted based on the hybrid expert model layer to obtain the first feature extraction result, which can be achieved in the following manner: based on the first gating network model, according to the first context marker sequence, the first model weight of the expert neural network model in the image feature extraction dimension is determined, and according to the first model weight, the first target neural network model required to perform the information prediction task in the image feature extraction dimension is determined from the expert neural network; based on the second gating network model, according to the first context marker sequence, the second model weight of the expert neural network model in the text feature extraction dimension is determined, and according to the second model weight, the second target neural network model required to perform the information prediction task in the text feature extraction dimension is determined from the expert neural network; the first context marker sequence and the first context overall representation are input into the first target neural network model to obtain the image feature prediction result in the image feature extraction dimension, and the first context marker sequence and the first context overall representation are input into the second target neural network model to obtain the text feature prediction result in the text feature extraction dimension, so as to determine the first feature extraction result according to the image feature prediction result and the text feature prediction result.
[0098] In an example embodiment, in the process of determining the first feature extraction result based on the image feature prediction result and the text feature prediction result, the image feature prediction result and the text feature prediction result can be directly spliced to obtain the first feature extraction result, or the image feature prediction result and the text feature prediction result can be configured with corresponding weight values and then obtained by weighted summation. This example does not impose any special restrictions on this.
[0099] In an example embodiment, after obtaining the first certificate information prediction result, a target loss function can be constructed based on the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate; wherein the target loss function used here may include a cross-entropy loss function, and of course other loss functions may be used according to actual needs, and this example does not impose any special restrictions on this.
[0100] In step S140, the low-rank matrix parameters of the multimodal large model in the document information extraction scenario are determined based on the target loss function, and the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain an information extraction model.
[0101] In this example embodiment, first, the low-rank matrix parameters of the multimodal large model in the document information extraction scenario are determined based on the target loss function; wherein the low-rank matrix parameters recorded here may include but are not limited to the matrix parameters of the modal encoder in the multimodal large model, the matrix parameters of the input projector, the matrix parameters of the large language model, the matrix parameters of the output projector, and the matrix parameters of the modal generator, etc.; secondly, the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain an information extraction model; specifically, it can be achieved in the following way: the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain the adjusted large model, and the second original virtual document is fine-tuned based on the adjusted large model Information extraction, obtaining a second certificate information prediction result; calculating the character edit distance between the second certificate information prediction result and the second actual certificate information corresponding to the second original virtual certificate, and determining the first character accuracy of the second certificate information prediction result based on the character edit distance; determining the field accuracy in the second certificate information prediction result based on the character edit distance, and determining the second character accuracy of the second certificate information prediction result based on the field accuracy; determining the model accuracy of the adjusted large model based on the first character accuracy and the second character accuracy, and when it is determined that the model accuracy is greater than or equal to a preset threshold, using the adjusted large model as the information extraction model.
[0102] In an example embodiment, verification of the fine-tuned multimodal large model can be achieved in the following manner: first, determine the average character accuracy (i.e., the first character accuracy); wherein, the first character accuracy can be used to measure the accuracy of text extraction by the model; in actual application, first, calculate the character edit distance (NED, Normalized Edit Distance metric) between the predicted result x and the true result y of each evaluation field, calculate the average NED of all samples of the field, and obtain the character error rate, and finally calculate the single word accuracy (i.e., the first character accuracy) based on the character error rate; wherein, the specific calculation formula of the character edit distance can be referred to as shown in the following formula (1); the specific calculation formula of the first character accuracy can be referred to as shown in formula (2).
[0103]
[0104]
[0105] Where NED is the character edit distance, LD(x,y) is the amount of characters that need to be edited between the field prediction result x and the true result y, |x| is the number of characters included in the field prediction result x, and y| is the number of characters included in the true result y; ACC char is the first character accuracy, is the average NED of all samples.
[0106] Secondly, determine the average field accuracy (i.e., the second character accuracy); the second character accuracy can be used to measure the accuracy of the model in extracting the document field; in the calculation process of the second character accuracy, assuming that the result x predicted by the evaluation field is exactly the same as the true result y, it means that the model predicts a hit f(x, y) = 1, otherwise it means a miss f(x, y) = 0; under this premise, the specific calculation formula for the second character accuracy can be shown in the following formula (3).
[0107]
[0108] Among them, ACC line is the second character accuracy, f(x i ,y i ) is the indicator function. When the result x predicted by the evaluation field is exactly the same as the true result y, f(x i ,y i ) takes the value of 1; when the result x predicted by the evaluation field is not exactly the same as the actual result y, the value is 0.
[0109] So far, the training method of the information extraction model recorded in the exemplary embodiment of the present disclosure has been fully realized. Based on the above-mentioned content, it can be known that in the process of actual application, since the fine-tuning training of the model is an effective means to improve the application effect of the multimodal large model in the document information extraction scenario, the present disclosure combines data generation, prompt word design, SFT / LoRA fine-tuning and strict verification steps to effectively improve the performance of the multimodal information extraction model based on virtual documents; among them, in the process of model training, every detail in the data preprocessing, model training and verification process is very critical, thus ensuring the accuracy and robustness of the final model.
[0110] Furthermore, after obtaining the information extraction model, the information extraction task can also be performed based on the obtained information extraction model; specifically, when the information extraction model performs the document information extraction task, first, the virtual document data required for model training is synthesized using an automated script method; then, the synthesized virtual document samples are used to train a multimodal large model to improve the document information extraction capability of the model; further, in the process of document information extraction, the following definition is required: given an document image I and a set of target fields F = {f 1 ,f 2 ,...,f n}, the trained model needs to extract each target field f from the document image I i The corresponding field value v i , to form a structured output O = {f 1 :v 1 ,f 2 :v 2 ,...,f n :v n}; Furthermore, when extracting the fields and field values included in the certificate information based on the certificate information extraction model, the specific input is a certificate image I and a prompt word (Prompt); wherein the prompt word is used to specify the target field and its semantics; the output obtained is the structured information O of the certificate. Based on this method, not only can the extraction of certificate information of different styles be realized, but also the extraction of certificate information of different languages can be realized, so that the extraction of certificate information of multiple languages and types can be realized on the basis of improving the accuracy of the extracted certificate information.
[0111] The following are embodiments of the device disclosed herein, which can be used to execute the method embodiments disclosed herein. For details not disclosed in the device embodiments disclosed herein, please refer to the method embodiments disclosed herein.
[0112] The exemplary embodiment of the present disclosure also provides a training device for an information extraction model. Figure 7As shown, the training device of the information extraction model may include a script code creation module 710, an image enhancement processing module 720, a loss function construction module 730 and a large model fine-tuning module 740. Among them:
[0113] The script code creation module 710 may be used to create original certificate structure information required for generating the first original virtual certificate, and to create the script code required for generating the first original virtual certificate according to the original certificate structure information;
[0114] The image enhancement processing module 720 may be used to generate the first original virtual certificate based on the script code, and perform image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate;
[0115] The loss function construction module 730 may be used to input the first standard virtual certificate into the multimodal large model to obtain a first certificate information prediction result, and to construct a target loss function according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate;
[0116] The large model fine-tuning module 740 can be used to determine the low-rank matrix parameters of the multimodal large model in the document information extraction scenario based on the target loss function, and fine-tune the multimodal large model based on the low-rank matrix parameters to obtain an information extraction model.
[0117] In an exemplary embodiment of the present disclosure, original certificate structure information required for generating a first original virtual certificate is created, including: creating the name of the field required for generating the first original virtual certificate, the position of the field in the first original virtual certificate, the style of the field in the first original virtual certificate, and the color information, texture information and image information required for generating the first original virtual certificate; creating the original certificate structure information required for generating the first original virtual certificate according to the name, position, style, color information, texture information and image information of the field.
[0118] In an exemplary embodiment of the present disclosure, a script code required for generating a first original virtual certificate is created based on the original certificate structure information, including: determining the name of the field and the field insertion position code of the field value corresponding to the name of the field in the first original virtual certificate based on the position of the field; determining the font code, font size code and text color code of the name of the field and the field value in the first original virtual certificate based on the style of the field; determining the background color code of the first original virtual certificate based on the texture information, and determining the image insertion position code of the image information in the first original virtual certificate; creating the script code required for generating the first original virtual certificate based on the field insertion position code, font code, font size code, text color code, background color code and image insertion position code.
[0119] In an exemplary embodiment of the present disclosure, the first original virtual certificate is generated based on the script code, including: determining the type of the field based on the name of the field, and generating a random field value corresponding to the type of the field; verifying the legitimacy of the random field value according to the type of the field, and after the legitimacy check passes, diversifying the language type of the name of the field and the random field value to obtain names and field values of multiple different language types; compiling the names and field values of multiple different language types and the script code to obtain the first original virtual certificate.
[0120] In an exemplary embodiment of the present disclosure, image enhancement processing is performed on the first original virtual certificate to obtain a first standard virtual certificate, including: performing noise processing and / or blur processing and / or compression distortion processing on the first original virtual certificate to obtain a virtual certificate with different noise and / or different clarity and / or different pixels; and / or randomly adjusting the position of the fields and / or the font style of the fields included in the first original virtual certificate to obtain a virtual certificate with different positions and / or different font styles.
[0121] In an exemplary embodiment of the present disclosure, the multimodal large model includes a modal encoder, an input projector, a large language model, an output projector and a modal generator; wherein, the first standard virtual certificate is input into the multimodal large model to obtain a first certificate information prediction result, including: encoding the first standard virtual certificate based on the modal encoder to obtain a first image feature; performing feature mapping processing on the first image feature based on the input projector to obtain a first feature mapping result; performing feature extraction on the first feature mapping result based on the large language model to obtain a first feature extraction result; performing feature conversion processing on the first feature extraction result based on the output projector to obtain a first feature conversion result; performing modality conversion on the first feature conversion result based on the modal generator to obtain the first certificate information prediction result.
[0122] In an exemplary embodiment of the present disclosure, the large language model includes an embedding mapping layer, a first encoding layer and a hybrid expert model layer; wherein, feature extraction is performed on the first feature mapping result based on the large language model to obtain a first feature extraction result, including: obtaining preset information classification prompt information and preset modal fusion prompt information corresponding to the first standard virtual certificate, and generating context information to be predicted based on the preset information classification prompt information and preset modal fusion prompt information; embedding mapping is performed on the first feature mapping result based on the embedding mapping layer to obtain a second image feature, and embedding mapping is performed on the context information to be predicted based on the embedding mapping layer to obtain a first context marker sequence; encoding is performed on the second image feature and the first context marker sequence based on the first encoding layer to obtain a first context overall representation, and the first context marker sequence and the first context overall representation are predicted based on the hybrid expert model layer to obtain the first feature extraction result.
[0123] In an exemplary embodiment of the present disclosure, the hybrid expert model layer includes a first gating network model, a second gating network model and a plurality of expert neural network models; wherein, the first context marker sequence and the first context overall representation are predicted based on the hybrid expert model layer to obtain the first feature extraction result, including: determining the first model weight of the expert neural network model in the image feature extraction dimension based on the first gating network model according to the first context marker sequence, and determining the first target neural network model required to perform the information prediction task in the image feature extraction dimension from the expert neural network according to the first model weight; determining the second model weight of the expert neural network model in the text feature extraction dimension based on the second gating network model according to the first context marker sequence, and determining the second target neural network model required to perform the information prediction task in the text feature extraction dimension from the expert neural network according to the second model weight; inputting the first context marker sequence and the first context overall representation into the first target neural network model to obtain the image feature prediction result in the image feature extraction dimension, and inputting the first context marker sequence and the first context overall representation into the second target neural network model to obtain the text feature prediction result in the text feature extraction dimension, so as to determine the first feature extraction result according to the image feature prediction result and the text feature prediction result.
[0124] In an exemplary embodiment of the present disclosure, the multimodal large model is fine-tuned based on low-rank matrix parameters to obtain an information extraction model, including: fine-tuning the multimodal large model based on low-rank matrix parameters to obtain an adjusted large model, and extracting information from the second original virtual certificate based on the adjusted large model to obtain a second certificate information prediction result; calculating the character editing distance between the second certificate information prediction result and the second actual certificate information corresponding to the second original virtual certificate, and determining the first character accuracy of the second certificate information prediction result according to the character editing distance; determining the field accuracy in the second certificate information prediction result according to the character editing distance, and determining the second character accuracy of the second certificate information prediction result according to the field accuracy; determining the model accuracy of the adjusted large model based on the first character accuracy and the second character accuracy, and when it is determined that the model accuracy is greater than or equal to a preset threshold, using the adjusted large model as the information extraction model.
[0125] The specific details of each module in the training device of the above-mentioned information extraction model have been described in detail in the corresponding training method of the information extraction model, so they will not be repeated here.
[0126] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0127] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0128] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. A person skilled in the art will appreciate that various aspects of the present disclosure may be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which may be collectively referred to herein as a circuit, module, or system.
[0129] Refer to the following Figure 8 The electronic device 800 according to this embodiment of the present disclosure is described. Figure 8 The electronic device 800 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0130] like Figure 8 As shown, the electronic device 800 is in the form of a general computing device. The components of the electronic device 800 may include, but are not limited to: the at least one processing unit 810, the at least one storage unit 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0131] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 810 can perform the following steps: Figure 1Step S110 shown in: creating original document structure information required to generate a first original virtual document, and creating script code required to generate the first original virtual document based on the original document structure information; step S120: generating the first original virtual document based on the script code, and performing image enhancement processing on the first original virtual document to obtain a first standard virtual document; step S130: inputting the first standard virtual document into the multimodal large model to obtain a first document information prediction result, and constructing a target loss function based on the first document information prediction result and the first actual document information of the first standard virtual document; step S140: determining the low-rank matrix parameters of the multimodal large model in the document information extraction scenario based on the target loss function, and fine-tuning the multimodal large model based on the low-rank matrix parameters to obtain an information extraction model.
[0132] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .
[0133] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0134] Bus 830 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0135] The electronic device 800 may also communicate with one or more external devices 900 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 850. Furthermore, the electronic device 800 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 860. As shown, the network adapter 860 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0136] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0137] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary implementations of the present disclosure described in the above "Exemplary Method" section of the present specification.
[0138] According to the program product for implementing the above method in the embodiment of the present disclosure, it can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0139] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0140] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0141] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0142] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0143] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0144] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the inventions invented herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not invented by the present disclosure. The specification and examples are to be considered merely exemplary, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A training method for an information extraction model, characterized in that: include: Creating original certificate structure information required for generating a first original virtual certificate, and creating script code required for generating the first original virtual certificate according to the original certificate structure information; generating the first original virtual certificate based on the script code, and performing image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate; Inputting the first standard virtual certificate into the multimodal large model to obtain a first certificate information prediction result, and constructing a target loss function according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate; Based on the objective loss function, the low-rank matrix parameters of the multimodal large model in the document information extraction scenario are determined, and the multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain an information extraction model.
2. The training method of the information extraction model according to claim 1, characterized in that: The original certificate structure information required to generate the first original virtual certificate is created, including: Create the name of the field required for generating the first original virtual certificate, the position of the field in the first original virtual certificate, the style of the field in the first original virtual certificate, and the color information, texture information, and image information required for generating the first original virtual certificate; The original certificate structure information required for generating the first original virtual certificate is created according to the name, position, style, color information, texture information and image information of the fields.
3. The training method of the information extraction model according to claim 2, characterized in that: The script code required to generate the first original virtual certificate is created according to the original certificate structure information, including: Determine, according to the position of the field, the name of the field and a field insertion position code of a field value corresponding to the name of the field in the first original virtual certificate; Determine the name of the field and the font code, font size code and text color code of the field value in the first original virtual certificate according to the style of the field; Determine a background color code of the first original virtual certificate according to the texture information, and determine an image insertion position code of the image information in the first original virtual certificate; The script code required for generating the first original virtual certificate is created according to the field insertion position code, font code, font size code, text color code, background color code and image insertion position code.
4. The training method of the information extraction model according to claim 2, characterized in that: Generating the first original virtual certificate based on the script code includes: Determining a type of the field based on the name of the field, and generating a random field value corresponding to the type of the field; Verifying the legitimacy of the random field value according to the type of the field, and after the legitimacy verification passes, performing diversified processing on the name of the field and the language type of the random field value to obtain names and field values in multiple different language types; The names and field values in a plurality of different language types and the script code are compiled to obtain the first original virtual certificate.
5. The training method of the information extraction model according to claim 1, characterized in that: Performing image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate includes: Performing noise processing and / or blur processing and / or compression distortion processing on the first original virtual certificate to obtain a virtual certificate with different noise and / or different definition and / or different pixels; and / or The positions of the fields and / or the font styles of the fields included in the first original virtual certificate are randomly adjusted to obtain a virtual certificate with a different position and / or a different font style.
6. The training method of the information extraction model according to claim 1, characterized in that: The multimodal large model includes a modality encoder, an input projector, a large language model, an output projector, and a modality generator; The first standard virtual certificate is input into the multimodal large model to obtain the first certificate information prediction result, including: Encoding the first standard virtual certificate based on the modality encoder to obtain a first image feature; Performing feature mapping processing on the first image feature based on the input projector to obtain a first feature mapping result; Performing feature extraction on the first feature mapping result based on the large language model to obtain a first feature extraction result; Performing feature conversion processing on the first feature extraction result based on the output projector to obtain a first feature conversion result; The first feature conversion result is modally converted based on the modality generator to obtain the first certificate information prediction result.
7. The training method of the information extraction model according to claim 6, characterized in that: The large language model includes an embedding mapping layer, a first encoding layer, and a hybrid expert model layer; The step of performing feature extraction on the first feature mapping result based on the large language model to obtain a first feature extraction result includes: Acquire preset information classification prompt information and preset modality fusion prompt information corresponding to the first standard virtual certificate, and generate context information to be predicted according to the preset information classification prompt information and the preset modality fusion prompt information; Performing embedding mapping processing on the first feature mapping result based on the embedding mapping layer to obtain a second image feature, and performing embedding mapping processing on the context information to be predicted based on the embedding mapping layer to obtain a first context marker sequence; Based on the first coding layer, the second image feature and the first context marker sequence are encoded to obtain a first context overall representation, and based on the hybrid expert model layer, the first context marker sequence and the first context overall representation are predicted to obtain the first feature extraction result.
8. The training method of the information extraction model according to claim 7, characterized in that: The hybrid expert model layer includes a first gating network model, a second gating network model and a plurality of expert neural network models; The first feature extraction result is obtained by predicting the first context marker sequence and the first context overall representation based on the hybrid expert model layer, including: Determine, based on the first gating network model and the first context flag sequence, a first model weight of the expert neural network model in the image feature extraction dimension, and determine, from the expert neural network according to the first model weight, a first target neural network model required to perform the information prediction task in the image feature extraction dimension; Determine a second model weight of the expert neural network model on the text feature extraction dimension based on the second gating network model according to the first contextual marker sequence, and determine a second target neural network model required to perform the information prediction task on the text feature extraction dimension from the expert neural network according to the second model weight; The first context marker sequence and the first context overall representation are input into the first target neural network model to obtain image feature prediction results in the image feature extraction dimension, and the first context marker sequence and the first context overall representation are input into the second target neural network model to obtain text feature prediction results in the text feature extraction dimension, so as to determine the first feature extraction result based on the image feature prediction results and the text feature prediction results.
9. The training method of the information extraction model according to claim 1, characterized in that: The multimodal large model is fine-tuned based on the low-rank matrix parameters to obtain an information extraction model, including: Fine-tune the multimodal large model based on the low-rank matrix parameters to obtain an adjusted large model, and extract information from the second original virtual certificate based on the adjusted large model to obtain a second certificate information prediction result; Calculating a character edit distance between the second certificate information prediction result and the second actual certificate information corresponding to the second original virtual certificate, and determining a first character accuracy rate of the second certificate information prediction result according to the character edit distance; Determine a field accuracy rate in the second certificate information prediction result according to the character edit distance, and determine a second character accuracy rate of the second certificate information prediction result according to the field accuracy rate; The model accuracy of the adjusted large model is determined according to the first character accuracy and the second character accuracy, and when it is determined that the model accuracy is greater than or equal to a preset threshold, the adjusted large model is used as the information extraction model.
10. A training device for an information extraction model, characterized in that: include: A script code creation module, used to create original certificate structure information required to generate the first original virtual certificate, and to create the script code required to generate the first original virtual certificate according to the original certificate structure information; An image enhancement processing module, used to generate the first original virtual certificate based on the script code, and perform image enhancement processing on the first original virtual certificate to obtain a first standard virtual certificate; A loss function construction module, used for inputting the first standard virtual certificate into the multimodal large model to obtain a first certificate information prediction result, and constructing a target loss function according to the first certificate information prediction result and the first actual certificate information of the first standard virtual certificate; A large model fine-tuning module is used to determine the low-rank matrix parameters of the multimodal large model in the document information extraction scenario based on the target loss function, and fine-tune the multimodal large model based on the low-rank matrix parameters to obtain an information extraction model.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method of the information extraction model described in any one of claims 1 to 9 is implemented.
12. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the training method of the information extraction model described in any one of claims 1-9 by executing the executable instructions.