License information extraction method and system, training method of information extraction model, computing device, computer readable storage medium and computer program product

By fusing text semantics and layout position features into the information extraction model and combining it with information extraction prompts, the problems of poor scalability and low efficiency in certificate information extraction are solved, and high-precision information extraction of various certificate images is achieved.

CN120726656APending Publication Date: 2025-09-30HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410384124.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies for extracting certificate information suffer from poor scalability and low efficiency, making it difficult to effectively identify and extract information from various types of certificate images.

Method used

By adopting the information extraction model, integrating the text semantics and typesetting position features and combining the information extraction prompt information, high-precision information extraction of certificate images can be achieved.

Benefits of technology

The accuracy and efficiency of certificate information extraction are improved, and it can be applied to various types of certificate images, including untrained certificate types, improving the generalization and efficiency of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726656A_ABST
    Figure CN120726656A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a license information extraction method and system, an information extraction model training method, computing equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining a license image to be subjected to information extraction; identifying a text in the license image and typesetting position information of the text in the license image; analyzing the information extraction prompt information, the text and the typesetting position information through an information extraction model to obtain certificate information which needs to be extracted and is extracted by the information extraction model; wherein the information extraction prompt information is used for describing the information type of to-be-extracted license information in the license image, and the information extraction model extracts the license information corresponding to the information type based on text semantic typesetting fusion features. The text semantic typesetting fusion feature is obtained by fusing a semantic feature corresponding to the text and a typesetting feature corresponding to the typesetting position information. The method can improve the extraction effect of the license information in the license image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a method for extracting certificate information, a system for extracting certificate information, a method for training an information extraction model, a device for extracting certificate information, a device for training an information extraction model, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Currently, in many scenarios, it is necessary to extract information from images in order to process the extracted information accordingly.

[0003] For example, in the government sector, it is necessary to extract certificate information. Certificates refer to documents and licenses. Certificate information extraction can be performed on images of certificates. There are many types of certificates, such as identity cards, birth certificates, and business licenses. The information types and layouts of different certificates vary, resulting in the need for improved certificate information extraction. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a method for extracting license information, which can improve the extraction of license information from license images. One or more embodiments of this specification also relate to a license information extraction system, a method for training an information extraction model, a license information extraction device, a training device for an information extraction model, a computing device, a computer-readable storage medium, and a computer program product.

[0005] According to a first aspect of an embodiment of this specification, a method for extracting license information is provided, comprising:

[0006] Obtaining a certificate image for information extraction;

[0007] Identifying text in the certificate image and typeset position information of the text in the certificate image;

[0008] The information extraction prompt information, the text and the typesetting position information are analyzed by an information extraction model to obtain the certificate information to be extracted by the information extraction model; wherein, the information extraction prompt information is used to describe the information type of the certificate information to be extracted in the certificate image, and the information extraction model extracts the certificate information corresponding to the information type based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic feature corresponding to the text and the typesetting feature corresponding to the typesetting position information.

[0009] According to a second aspect of the embodiments of this specification, a method for training an information extraction model is provided, comprising:

[0010] Obtaining sample prompt information and a sample certificate image marked with a certificate information tag, wherein the sample prompt information is used to describe the information type of the certificate information to be extracted from the sample certificate image;

[0011] Identifying sample text in the sample ID image and sample layout position information of the sample text in the sample ID image;

[0012] The sample prompt information, the sample text, and the sample typeset position information are analyzed by the information extraction model to be trained to obtain the predicted license information to be extracted by the information extraction model; wherein the information extraction model extracts the predicted license information based on text semantic typeset fusion features, and the text semantic typeset fusion features are obtained by fusing semantic features corresponding to the sample text and typeset features corresponding to the sample typeset position information;

[0013] The predicted license information is compared with the license information label, and the parameters of the information extraction model are adjusted based on the comparison result to obtain a trained information extraction model.

[0014] According to a third aspect of the embodiments of this specification, a method for extracting license information is provided, which is applied to a server and includes:

[0015] In response to a certificate information extraction request for a certificate image sent by a client, obtaining information extraction prompt information; wherein the information extraction prompt information is used to describe the information type of the certificate information to be extracted;

[0016] Identifying text in the certificate image and typeset position information of the text in the certificate image;

[0017] The information extraction prompt information, the text, and the typesetting position information are analyzed by an information extraction model to obtain the license information to be extracted by the information extraction model; wherein the information extraction model extracts the license information corresponding to the information type based on a text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information;

[0018] The certificate information to be extracted, which is extracted by the information extraction model, is sent to the client, so that the client can display the certificate information.

[0019] According to a fourth aspect of the embodiments of this specification, a device for extracting certificate information is provided, comprising:

[0020] A first acquisition module is used to acquire a certificate image to be extracted;

[0021] A first recognition module, configured to recognize text in the certificate image and typeset position information of the text in the certificate image;

[0022] The first information extraction module is used to analyze the information extraction prompt information, the text and the typesetting position information through an information extraction model to obtain the certificate information to be extracted by the information extraction model; wherein, the information extraction prompt information is used to describe the information type of the certificate information to be extracted in the certificate image, and the information extraction model extracts the certificate information corresponding to the information type based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic feature corresponding to the text and the typesetting feature corresponding to the typesetting position information.

[0023] According to a fifth aspect of the embodiments of this specification, a training device for an information extraction model is provided, comprising:

[0024] an acquisition module, configured to acquire sample prompt information and a sample license image labeled with a license information tag, wherein the sample prompt information is used to describe the type of license information to be extracted from the sample license image;

[0025] an identification module, configured to identify sample text in the sample certificate image and sample layout position information of the sample text in the sample certificate image;

[0026] An information extraction module is configured to analyze the sample prompt information, the sample text, and the sample typeset position information using an information extraction model to be trained, to obtain the predicted license information to be extracted by the information extraction model; wherein the information extraction model extracts the predicted license information based on a text semantic typeset fusion feature, wherein the text semantic typeset fusion feature is obtained by fusing semantic features corresponding to the sample text with typeset features corresponding to the sample typeset position information;

[0027] A parameter adjustment module is used to compare the predicted license information with the license information label, and adjust the parameters of the information extraction model based on the comparison result to obtain a trained information extraction model.

[0028] According to a sixth aspect of the embodiments of this specification, a device for extracting license information is provided, which is applied to a server and includes:

[0029] an acquisition module, configured to obtain information extraction prompt information in response to a certificate information extraction request for a certificate image sent by a client; wherein the information extraction prompt information is used to describe the information type of the certificate information to be extracted;

[0030] A recognition module, configured to recognize text in the certificate image and typeset position information of the text in the certificate image;

[0031] An information extraction module is configured to analyze the information extraction prompt information, the text, and the typesetting position information using an information extraction model to obtain the license information to be extracted by the information extraction model; wherein the information extraction model extracts the license information corresponding to the information type based on a text semantic typesetting fusion feature, wherein the text semantic typesetting fusion feature is obtained by fusing semantic features corresponding to the text with typesetting features corresponding to the typesetting position information;

[0032] The sending module is used to send the certificate information to be extracted by the information extraction model to the client, so that the client can display the certificate information.

[0033] According to a seventh aspect of the embodiments of this specification, there is provided a system for extracting certificate information, including: a client and a server;

[0034] The client is used to: send a certificate information extraction request to the server;

[0035] The server is configured to: execute the above method based on the certificate information extraction request, extract the certificate information to be extracted, and send the certificate information to be extracted to the client;

[0036] The client is also used to display the received certificate information.

[0037] According to an eighth aspect of the embodiments of this specification, there is provided a computing device, including: a memory and a processor;

[0038] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.

[0039] According to a ninth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the steps of the above method are implemented when the computer program / instruction is executed by a processor.

[0040] According to a tenth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above method when the computer program / instruction in the computer program product is executed in a processor.

[0041] In an embodiment of the present specification, the text and its typeset position information in the certificate image to be subjected to information extraction can be identified. The information extraction model fuses the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information to obtain a text semantic typesetting fusion feature, and then extracts the certificate information corresponding to the information type described by the information extraction prompt information based on the text semantic typesetting fusion feature and the information extraction prompt information. In this certificate information extraction method, the information extraction model can analyze the relationship between the text semantics and its typesetting position, and adds information extraction prompt information to describe the certificate information to be extracted. This certificate information extraction method can be applied to various types of certificate images. Even for certificate types that are not involved in the training of the information extraction model, the certificate information can be extracted with high accuracy by analyzing the relationship between the semantics of the text in the certificate image and its typesetting position, combined with the information extraction prompt information. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a schematic diagram of the structure of a certificate information extraction system provided in an embodiment of this specification;

[0043] Figure 2 This is a flow chart of a method for extracting license information provided in one embodiment of this specification;

[0044] Figure 3 This is a schematic diagram of the structure of an information extraction model provided in an embodiment of this specification;

[0045] Figure 4 This is a flowchart of another method for extracting license information provided in an embodiment of this specification;

[0046] Figure 5 This is a flowchart of a training method for an information extraction model provided in one embodiment of this specification;

[0047] Figure 6 is a flowchart of another information extraction model training method provided in one embodiment of this specification;

[0048] Figure 7 This is a schematic diagram of the structure of a certificate information extraction device provided in one embodiment of this specification;

[0049] Figure 8 This is a schematic diagram of the structure of another device for extracting license information provided in an embodiment of this specification;

[0050] Figure 9 This is a schematic diagram of the structure of a training device for an information extraction model provided in one embodiment of this specification;

[0051] Figure 10 This is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION

[0052] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0053] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items. The term "at least one" in one or more embodiments of this specification refers to "one or more" and "a plurality" refers to "two or more". The term "including" is an open description and should be understood as "including but not limited to", and may include other content on the basis of what has been described.

[0054] It should be understood that although the terms "first," "second," and the like may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, without departing from the scope of one or more embodiments of this specification, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first." Depending on the context, the word "if" as used herein may be interpreted as "at the time of," "when," or "in response to determining."

[0055] In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant standards and requirements, and corresponding operation entrances must be provided for users to choose to authorize or refuse.

[0056] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. By pre-training a large model with large-scale unlabeled corpus, a pre-trained model with more than 100 million parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.

[0057] When the large model is actually used, only a small amount of samples are needed to fine-tune the pre-training model and it can be applied to different tasks. The large model can be widely used in natural language processing (NLP), computer vision and other fields, and can be specifically applied to computer vision tasks such as visual question answering (VQA), image description (IC), image generation, and text-based sentiment classification, text summary generation, machine translation and other natural language processing tasks. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. The models involved in the embodiments of this specification include but are not limited to the above-mentioned large models, and can be any type of machine learning model. For example, it can be an end-to-end seq2seq model, a convolutional neural network (CNN) model, a Transformer model, etc.

[0058] In the government sector, it is often necessary to identify certificate images and extract certificate information from them. There are many types of certificates, such as identity cards, driver's licenses, birth certificates, and business licenses. The field information in each certificate is different, and the format and text layout of different types of certificates are also different. These factors increase the difficulty of certificate information extraction. Currently, there are some solutions that use models to extract certificate information, but these solutions have poor scalability and can usually only achieve good recognition and information extraction effects for the certificate categories to which the training samples used during model training belong. For certificates not involved in model training, the training dataset needs to be re-labeled for iterative model training. There are also some solutions that can only extract information about one keyword in the certificate image each time certificate information is extracted, resulting in low certificate information extraction efficiency.

[0059] This specification provides a method for extracting license information. This method utilizes a highly generalizable information extraction model to extract license information with high accuracy for a wide range of license image types, while also improving information extraction efficiency. This specification also relates to a license information extraction system, a method for training an information extraction model, a license information extraction device, a training device for a license extraction model, a computing device, a computer-readable storage medium, and a computer program product, each of which is described in detail in the following embodiments.

[0060] Figure 1 This is a schematic diagram of the structure of a certificate information extraction system provided in an embodiment of this specification. Figure 1 As shown, the certificate information extraction system may include a server 101 and a client 102, and the client 102 may establish a communication connection with the server 101. For example, the server 101 may be a cloud server or a server cluster, and the client 102 may be a smartphone, a desktop computer, a laptop computer, a tablet computer, or a wearable device.

[0061] The certificate information extraction method provided in this specification can be applied to the server 101. The server 101 can provide a certificate information extraction service. The user can operate the client 102 to send a certificate information extraction request to the server 101, and the server 101 can provide a certificate information extraction service to the client 101 in response to the certificate information extraction request. The certificate information extraction request is a request for a certificate image. In one way, the certificate information extraction request can directly carry the certificate image. In another way, the certificate information extraction request can specify the certificate image. If the certificate image and the certificate information extraction request are sent based on different data packets, the certificate information extraction request carries the identifier of the certificate image, and the server 101 obtains the corresponding certificate image based on the identifier. For example, if the certificate image is stored in other devices, the certificate information extraction request carries the storage location of the certificate image, and the server 101 obtains the certificate image from other devices based on the storage location. In another embodiment, the certificate image may be stored in the server 101 before the certificate information extraction request is sent. The server 101 determines the certificate image targeted by the certificate information extraction request from the stored images based on the certificate information extraction request.

[0062] For example, client 102 sends a certificate image to be extracted, along with the information type corresponding to the certificate information to be extracted from the certificate image, to server 101, so that server 101 extracts the certificate information corresponding to the information type from the certificate image and sends the certificate information to client 102. After receiving the certificate information, client 102 may display the certificate information. Optionally, client 102 may send one or more certificate images to server 101 at a time, and may send information of one or multiple information types for each certificate image, instructing server 101 to extract certificate information corresponding to one or multiple information types from the certificate image.

[0063] The server 101 can use the information extraction model to extract the certificate information to be extracted. The information extraction model can be a large model. For example, the server 101 can use the information of the information type sent by the client 102 to generate information extraction prompt information (i.e., prompt information) for guiding the information extraction model. The information extraction prompt information and the certificate image can then be analyzed and processed by the information extraction model to obtain the certificate information output by the information extraction model. The information extraction prompt information is used to describe the information type of the certificate information to be extracted in the certificate image. The information extraction prompt information can also be used to make some settings for the execution and output method of the information extraction model. The specific process of extracting certificate information using the information extraction model will be introduced in detail later and will not be expanded here.

[0064] For example, client 102 may have an application installed that corresponds to the license information extraction service. By running this application, client 102 can connect to server 101 and obtain the license information extraction service provided by server 101. Client 102 may also access server 101 through a webpage to obtain the license information extraction service. Optionally, this application or webpage may be used solely to provide the license information extraction service, or may also be used to provide other services (such as information query, text analysis, document generation, etc.).

[0065] In one implementation, the client 102 page may display various functional areas, such as an area for uploading a certificate image, an area for inputting the type of certificate information to be extracted, an information output area, and various control buttons (such as a start recognition button and a stop recognition button). The user can operate on each functional area on the page to send the certificate image and information type to the server 101, and request the server 101 to extract the certificate information. After the server 101 completes the certificate information extraction, the extracted certificate information is displayed in the information output area.

[0066] In another implementation, the client 102 may display a dialog page where the user can send a photo of their ID and enter their desired information through a dialog. This allows the server 101 to receive the ID information and information type, allowing it to extract the ID information accordingly. For example, a user may send a photo of their ID and enter text such as "Please identify ID information of type A in this image." After the server 101 identifies the ID information of that type, it can provide the ID information through the dialog page.

[0067] In the embodiments of this specification, the client 102 may also be disconnected from the server 101, and the client 102 may independently execute the certificate information extraction method to extract information from the certificate image. In this method, the client 102 may be deployed with an information extraction model to extract certificate information using the information extraction model.

[0068] Figure 2 This is a flowchart of a method for extracting license information provided in an embodiment of this specification. The method can be applied to a license information extraction device. The license information extraction device can be Figure 1 The server 101 or client 102 in the certificate information extraction system shown. Figure 2 As shown, the method may include:

[0069] Step 202: Obtain a certificate image for information extraction.

[0070] For example, if the certificate information extraction method is applied to server 101, server 101 may receive a certificate image sent by client 102 to obtain the certificate image from which information extraction is to be performed. If the certificate information extraction method is applied to client 102, client 102 may, based on a user's selection instruction, determine a selected image from the stored images as the certificate image from which information extraction is to be performed. Alternatively, the client may also determine an image sent by another device as the certificate image from which information extraction is to be performed.

[0071] Step 204: Identify the text in the certificate image and the layout position information of the text in the certificate image.

[0072] Typesetting position information refers to the position of the area occupied by the corresponding text in the certificate image. In some embodiments, if the text in the certificate image is carried by a specific page element, the typesetting position information of the text may refer to the position of the page element carrying the text. The page element can be any type of component, such as a text box component.

[0073] The certificate information extraction device can perform optical character recognition (OCR) on the acquired certificate image set to obtain the text in the certificate image and the layout position information of the text in the certificate image. Performing OCR on the image can identify the text area, which can be represented by a rectangular text detection box. Each text area can include one or more characters, such as a text area can include a word or a sentence. Performing OCR on the image can also identify the individual characters in the text area. The text recognized on the certificate image can include the individual characters in each text area, and the layout position information of the text can include information about the text area (also called box information).

[0074] Step 206: Analyze the information extraction prompt information and the text and typesetting position information identified for the certificate image through the information extraction model to obtain the certificate information to be extracted by the information extraction model; wherein, the information extraction prompt information is used to describe the information type of the certificate information to be extracted in the certificate image, and the information extraction model extracts the certificate information corresponding to the information type based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information.

[0075] The certificate information extraction device can obtain information extraction prompt information for guiding the information extraction model. The information extraction prompt information can be used to describe the information type of the certificate information to be extracted in the certificate image, such as describing only one information type or multiple information types. The information extraction prompt information guides the information extraction model to extract the certificate information corresponding to the information type. For example, the certificate image can be an ID card image, and the information type described by the information extraction prompt information may include name, gender, ethnicity, place of origin, ID card number and other information types. The user can specify the information type of the certificate information to be extracted, and the certificate information extraction device can generate information extraction prompt information based on the information type; alternatively, the information extraction prompt information can also be directly input by the user.

[0076] In the embodiments of this specification, the information extraction model has the ability to understand the semantics of the text and can obtain the semantic features corresponding to the text based on the input text. The information extraction model also analyzes the corresponding typesetting features based on the input typesetting position information. The information extraction model fuses the semantic features and typesetting features to obtain a text semantic typesetting fusion feature. Based on the text semantic typesetting fusion feature, the correspondence between each information type and the license information, as well as the location of each information type and the license information, can be determined. Furthermore, the license information corresponding to the information type described in the information extraction prompt information can be determined.

[0077] In the embodiments of this specification, the information extraction prompt model has the ability to understand the semantics of the text, and the ability to understand the relationship between semantics and typeset position, and the information extraction prompt information can describe the specific information type in the certificate image. The information extraction model can analyze the relationship between the text semantics and its typeset position, and adds information extraction prompt information to describe the certificate information to be extracted. This certificate information extraction method can be applied to various types of certificate images. Even for certificate types that are not involved in the training of the information extraction model, the certificate information of each information type can be known, so as to achieve more accurate extraction of the required certificate information in the certificate image of that type. In addition, the information extraction prompt information can be used to extract the required certificate information of multiple information types at one time, thereby improving the efficiency of information extraction.

[0078] In the embodiments of this specification, information extraction can be performed on certificate images, and the layout position information of the text in the image is also taken into consideration. The prompt mechanism is used to guide the model to extract the required information. All required certificate information can be directly extracted without extracting information one by one, ensuring high information extraction efficiency. The certificate information extraction scheme has good generalization and can not only have a good information extraction effect on known certificate images, but can also be well applied to unknown certificates. For example, the certificate information extraction method provided in the embodiments of this specification can be applied to scenarios such as one-stop government affairs, AI certificate review, or intelligent certificate recognition.

[0079] The information extraction model in the embodiments of this specification may be an end-to-end (seq2seq) text information extraction model. Figure 3 This is a schematic diagram of the structure of an information extraction model provided in an embodiment of this specification. Figure 3 As shown, the information extraction model may include: an encoder, an adaptation module (Adapt Layer) and a decoder. The encoder includes a text semantic analysis module (also known as the text flow module in the figure) and a layout analysis module (also known as the layout flow module in the figure).

[0080] In this specification, the information extraction prompt information, text and layout position information are analyzed by the information extraction model to obtain the license information to be extracted by the information extraction model, which may include the following steps:

[0081] Splicing the information extraction prompt information with the text to obtain text splicing features, and determining the position coding information corresponding to the text splicing features;

[0082] Based on the typesetting position information, determining the splicing position information corresponding to the text splicing feature, and determining the position coding information corresponding to the splicing position information;

[0083] Inputting the text splicing features and the position coding information corresponding to the text splicing features into the text semantic analysis module to obtain the semantic features corresponding to the text output by the text semantic analysis module;

[0084] Inputting the splicing position information and the position coding information corresponding to the splicing position information into the layout analysis module to obtain the layout features corresponding to the layout position information output by the layout analysis module;

[0085] The semantic features and typographic features are semantically fused through the adaptation module to obtain the text semantic typographic fusion features;

[0086] The text semantic typesetting fusion features are decoded through the decoder to obtain the certificate information to be extracted.

[0087] Specifically, the certificate information extraction device can encode each word (also called token) in the text obtained by the certificate image recognition to obtain corresponding word encoding information (also known as Token Embedding in the figure). T1, t2, t3, t4 and t5 in the figure refer to the word encoding information obtained by encoding different words. The certificate information extraction device also encodes each word in the information extraction prompt information to obtain corresponding word encoding information, such as p1 and p2 in the figure. In the embodiment of this specification, the text can be represented by the word encoding information of each word therein, and the information extraction prompt information can also be represented by the word encoding information of each word therein. The certificate information extraction device splices the information extraction prompt information with the text, which can be splicing the encoding information of each word of the information extraction prompt information and the encoding information of each word of the text, and uses the information obtained after splicing as a text splicing feature. Optionally, the certificate information extraction device can also first splice the information extraction prompt information with the text, and then encode each word in the information obtained after splicing to obtain a text splicing feature.

[0088] The certificate information extraction device can also determine the position coding information corresponding to the text splicing features, specifically including determining the position coding information corresponding to each word in the text and the position coding information corresponding to each word in the information extraction prompt information. The position coding information can represent the arrangement order of each word. For example, the position of each word in the text can be encoded, and the position of each word in the information extraction prompt information can also be encoded to obtain the corresponding position coding information (that is, the 1D position Embedding in the figure). Figure 3 It can be determined that the position coding information corresponding to each word in the information extraction prompt information includes 0 and 1, and the position coding information corresponding to each word in the text includes 1, 2, 3, 4 and 5.

[0089] The certificate information extraction device can input the text splicing features and the corresponding position coding information into the text semantic analysis module to obtain the semantic features corresponding to the text output by the text semantic analysis module.

[0090] The certificate information extraction device can encode the typesetting position information of each text obtained by the certificate image recognition to obtain the corresponding typesetting coding information (that is, the 2D position Embedding in the figure). b1, b2, b3, b4 and b5 in the figure refer to the typesetting coding information obtained by encoding the typesetting position information of different words. The certificate information extraction device also supplements the typesetting position information for each word in the information extraction prompt information. bp in the figure represents the typesetting coding information obtained by encoding the typesetting position information supplemented by the information extraction prompt information, and the content it contains can all be 0, so that the alignment of the text splicing features and the splicing position information can be guaranteed. In the embodiment of this specification, the typesetting position information can be represented by typesetting coding information, and the information composed of the typesetting position information of the text and the typesetting position information of the information extraction prompt information can be called splicing position information.

[0091] The license information extraction device can also obtain position coding information corresponding to the splicing position information (such as the 1D position embedding in the figure). For example, the position coding information corresponding to the above-mentioned text splicing feature can be directly determined as the position coding information corresponding to the splicing position information. The license information extraction device can input the splicing position information and its corresponding position coding information into the layout analysis module to obtain the layout features corresponding to the layout position information output by the layout analysis module.

[0092] The above semantic features and typesetting features are semantically fused to obtain text semantic typesetting fusion features, and the text semantic typesetting fusion features are decoded by a decoder to obtain the certificate information to be extracted indicated by the information extraction prompt information. Figure 3 For example, the information types of the certificate information to be extracted include tasks and places of employment, and the extracted certificate information includes person X and company Y.

[0093] In one embodiment, the text semantic analysis module and the layout analysis module each include multiple transformer layers; the method provided in the embodiment of this specification may further include:

[0094] Information is exchanged through the transformer layer in the text semantic analysis module and the transformer layer in the layout analysis module, so that both the semantic features and the typesetting features contain the relationship information between the semantics of the text and the typesetting position.

[0095] Please continue to refer to Figure 3, the text semantic analysis module and the layout analysis module can both include multiple transformer layers, and the decoder can also include multiple transformer decoding layers, Figure 3 Take the example of each module including 4 transformer layers. The data in each transformer layer in the encoder can be MatMul, Scale, MaskOut, SoftMax and MatMul in sequence, where MatMul represents the multiplication of two vector matrices, Scale represents scaling, MaskOut represents mask output, and Softmax represents classification. The certificate information extraction device can exchange information through the transformer layer in the text semantic analysis module and the transformer layer in the layout analysis module, so that the obtained semantic features and typesetting features contain the relationship information between the semantics of the text and the typesetting position. The number of transformer layers in the text semantic analysis module and the layout analysis module is the same, and the data in every two transformer layers can exchange information after scaling.

[0096] In the embodiments of this specification, the encoder may map the result of the identification of the certificate image (text and layout position) and the information extraction prompt information into a semantically rich feature, and the decoder may decode the feature into a structured representation of the certificate information.

[0097] In the embodiment of this specification, before step 206, the method for extracting license information may further include:

[0098] Prompt information is obtained from the certificate information extraction request sent by the client, and structured information extraction prompt information is generated based on the prompt information according to a preset prompt information structure.

[0099] The certificate information extraction device can receive a certificate information extraction request sent by a client and, in response to the request, execute the above-described certificate information extraction method. The certificate information extraction request can include prompt information indicating the type of certificate information to be extracted from the certificate image. The certificate information extraction device processes the prompt information included in the certificate information extraction request, such as converting the prompt information into structured information extraction prompt information according to a preset prompt information structure. The information extraction model utilizes this structured information extraction prompt information for processing, thereby improving processing efficiency and accuracy.

[0100] For example, it is necessary to extract name and gender information from an ID card image. Assuming that the text obtained after the ID card image is recognized is: "Name A Gender Male Nationality Han Birthplace M Citizen ID Number 10010", the structured information extraction prompt information can be: <spot>Name <spot>Gender". The input of the text analysis module can be: <spot>Name <spot>gender<extra_id2> Name A Gender Male Nationality Han Place of Birth M Citizen ID Number 10010"<extra_id2> Indicates the separator between the prompt information and the main text.

[0101] The decoder can express the extracted license information into a structured language for output. Continuing with the above example, the license information output by the information extraction model can be "<extra_id0><extra_id0> Name<extra_id5> A<extra_id1><extra_id0> gender<extra_id5> male<extra_id1><extra_id1> ",in,<extra_id0> Indicates the start of extracting information.<extra_id5> Indicates the start of the extracted information content,<extra_id1> Indicates the end of extracting information.<extra_id0> and<extra_id1> Nesting is supported.

[0102] Figure 4 This is a flowchart of another method for extracting license information provided in an embodiment of this specification. This method can be applied to Figure 1 The server 101 in the certificate information extraction system shown in FIG. Figure 4 As shown, the method may include:

[0103] Step 402: In response to the certificate information extraction request for the certificate image sent by the client, obtain information extraction prompt information; wherein the information extraction prompt information is used to describe the information type of the certificate information to be extracted.

[0104] In the embodiment of this specification, step 402 can refer to Figure 1 The relevant content in the introduction of , will not be repeated here for the content introduced earlier.

[0105] Step 404: Identify the text in the certificate image and the layout position information of the text in the certificate image.

[0106] In the embodiment of this specification, the method of obtaining the sample prompt information in step 404 can refer to the introduction in step 204, and the content introduced above will not be repeated here.

[0107] Step 406: Analyze the information extraction prompt information, the text, and the typesetting position information through the information extraction model to obtain the certificate information to be extracted by the information extraction model; wherein, the information extraction model extracts the certificate information corresponding to the information type based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information.

[0108] In the embodiment of this specification, the method of obtaining the sample prompt information in step 406 can refer to the introduction in step 206, and the content introduced above will not be repeated here.

[0109] Step 408: Send the certificate information to be extracted by the information extraction model to the client, so that the client can display the certificate information.

[0110] In the embodiment of this specification, step 408 can refer to Figure 1 The relevant content in the introduction of , will not be repeated here for the content introduced earlier.

[0111] In summary, in the certificate information extraction method provided in the embodiments of this specification, the text and its typeset position information in the certificate image to be extracted can be identified for the certificate image to be extracted. The information extraction model fuses the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information to obtain the text semantic typesetting fusion features, and then extracts the certificate information corresponding to the attribute described by the information extraction prompt information based on the text semantic typesetting fusion features and the information extraction prompt information. In this certificate information extraction method, the information extraction model can analyze the relationship between the text semantics and its typeset position, and adds information extraction prompt information to describe the certificate information to be extracted. This certificate information extraction method can be applied to various types of certificate images. Even for certificate types that are not involved in the training of the information extraction model, the certificate information can be extracted with high accuracy by analyzing the relationship between the semantics of the text in the certificate image and its typeset position, combined with the information extraction prompt information.

[0112] In the embodiments of this specification, before analyzing the information extraction prompt information, text and typesetting position information through the information extraction model, the information extraction model may be trained in advance. The following introduces the training method of the information extraction model. Figure 5 This is a flowchart of a method for training an information extraction model provided in one embodiment of this specification. This method can be applied to a training device for an information extraction model, which is hereinafter referred to as a training device. For example, the training device can be Figure 1 The server 101 or client 102 in the certificate information extraction system shown can also be other devices other than the server 101 and client 102. Figure 5 As shown, the method may include:

[0113] Step 502: Obtain sample prompt information and a sample certificate image marked with a certificate information tag, wherein the sample prompt information is used to describe the information type of the certificate information to be extracted from the sample certificate image.

[0114] In the embodiment of this specification, the method of obtaining the sample prompt information in step 502 can refer to the introduction of obtaining information extraction prompt information in step 206, and the content introduced above will not be repeated here.

[0115] For sample ID images, all or part of the ID information corresponding to the information type in the ID image can be manually annotated. The annotated ID information can be used as the ID information label, and the annotated ID image can be used as the sample ID image. This sample ID image is used to train the information extraction model.

[0116] Step 504: Identify the sample text in the sample certificate image and the sample layout position information of the sample text in the sample certificate image.

[0117] In the embodiments of this specification, the method for identifying the sample text and sample typesetting position information in the sample certificate image is similar to the method for identifying the text and typesetting position information in the certificate image to be extracted. Step 504 can refer to the relevant introduction of step 204, and the content introduced above will not be repeated here.

[0118] Step 506: Analyze the sample prompt information, sample text, and sample typesetting position information through the information extraction model to be trained to obtain the predicted license information to be extracted by the information extraction model; wherein, the information extraction model extracts the predicted license information based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the sample text and the typesetting features corresponding to the sample typesetting position information.

[0119] In the embodiment of this specification, the information extraction model to be trained and the trained information extraction model process the input information in a similar manner. Step 506 can refer to the relevant introduction of step 206, and the content introduced above will not be repeated here.

[0120] A sample prompt, along with sample text and sample layout position information recognized for a license image, can constitute a training sample. The information extraction model training device can construct multiple training samples, using these multiple training samples to train the information extraction model multiple times. Optionally, each training sample can be input into the information extraction model multiple times to train the model multiple times. For each training sample, the information extraction model can output corresponding predicted license information.

[0121] Step 508: Compare the predicted license information with the license information label, and adjust the parameters of the information extraction model based on the comparison result to obtain a trained information extraction model.

[0122] The training device of the information extraction model can compare the predicted license information and the license information label corresponding to each training sample. For example, in a certain training sample, the sample prompt information instructs the information extraction model to extract the license information of information type A in the sample license image. After the information extraction model outputs the predicted license information corresponding to the training sample, it can compare whether the predicted license information is the same as the labeled license information of information type A (that is, the license information label). In the case that the predicted license information and the license information label are different, the parameters of the information extraction model can be adjusted, and the steps of obtaining the predicted license information for the training sample and comparing the predicted license information and the license information label can be repeated until the training stop condition is reached, thereby obtaining a trained information extraction model.

[0123] For example, if a certain percentage of the training samples participating in model training correspond to the same predicted license information and license information labels, the training device for the information extraction model may determine that a training termination condition has been met. Alternatively, the training termination condition may be determined to have been met when the number of training times reaches a certain threshold.

[0124] In some cases, the number of sample license images labeled with license information is small. In this case, data enhancement can be performed based on the sample license images, and the images obtained by data enhancement are used as new sample license images (hereinafter referred to as new sample images). For example, the training method of the information extraction model provided in this specification also includes the following steps:

[0125] Adjusting the content and / or position of the sample text in the sample certificate image to obtain a new sample image;

[0126] Determine the sample prompt information and the certificate information label of the newly added sample image based on the sample prompt information and the certificate information label of the sample certificate image;

[0127] The information extraction model to be trained is used to analyze the sample prompt information, sample text, and sample layout position information of the newly added sample image to obtain the predicted license information to be extracted by the information extraction model;

[0128] The predicted license information of the newly added sample image is compared with the license information label of the newly added license image, and the parameters of the information extraction model are adjusted based on the comparison results to obtain a trained information extraction model.

[0129] In the embodiment of this specification, the same steps as those for the original sample certificate images can be performed for the newly added sample images, such as the relevant introduction in the above steps 506 and 508, which will not be repeated here.

[0130] The newly added sample images generated by the training device may include two types of images, one focusing on semantic understanding and the other focusing on typesetting understanding. For the newly added sample images focusing on semantic understanding, the training device may adjust the content of the sample text in the sample certificate image that is labeled with the certificate information label. For the newly added sample images focusing on typesetting understanding, the training device may adjust the typesetting position of the text in the sample certificate image. The sample text in the sample certificate image may include multiple information types and corresponding certificate information, and the labeled certificate information label includes the certificate information corresponding to the target information type. Adjustments can be made to the information type and the corresponding certificate information to obtain the newly added sample image.

[0131] Optionally, the sample text in the sample certificate image includes multiple information types and corresponding certificate information, and the annotated certificate information tag includes the certificate information corresponding to the target information type; adjusting the content and / or position of the sample text in the sample certificate image to obtain a new sample image includes at least one of the following four steps:

[0132] For any information type and corresponding certificate information in the sample certificate image, the certificate information is replaced with alternative information of the information type to obtain a first newly added sample image;

[0133] Replacing any information type and corresponding certificate information in the sample certificate image with the alternative information type and corresponding certificate information to obtain a second newly added sample image;

[0134] Merging at least two adjacent text regions in the sample certificate image to obtain a third newly added sample image; wherein the text region includes the location of at least part of the text of any information type or certificate information;

[0135] The text area in the sample certificate image is shifted to obtain a fourth newly added sample image.

[0136] The above four steps correspond to four ways of obtaining new sample images.

[0137] In the first approach, for all or part of the information types in any sample certificate image, multiple certificate information of that information type can be pre-set. A sample certificate image may include only one certificate information of that information type, with the remaining certificate information serving as alternative information of that information type. For any information type and corresponding certificate information in a sample certificate image, the training device can replace the current certificate information with the alternative information of that information type to obtain a first newly added sample image. This first newly added sample image can be a newly added sample image that focuses on semantic understanding.

[0138] In the second approach, for any sample ID image, the training device can select, from among the multiple acquired information types, information types that differ from the information type in the sample ID image as alternative information types. The training device can replace any information type and corresponding ID information in the sample ID image with the alternative information type and corresponding ID information, generating a second additional sample image. This second additional sample image can be a sample image that focuses on comprehension of typography.

[0139] In a third approach, the training device may merge at least two adjacent text regions in the sample ID image to generate a third additional sample image. This text region includes the location of at least a portion of text within any information type or ID information. For example, this text region may be a text detection frame obtained when recognizing the sample ID image. The training device may randomly merge adjacent text detection frames to simulate different text detection scenarios.

[0140] In a fourth approach, the training device can shift the text area in the sample certificate image in any direction, up, down, left, or right, to obtain a fourth additional sample image. The fourth additional sample image can be used to simulate the situation of median printing deviation in the certificate.

[0141] In the embodiments of this specification, the training device may generate new sample images using only the four aforementioned methods, or may combine any of the four aforementioned methods to generate new sample images. For example, the training device may combine the first and second methods with the third and fourth methods to generate more new sample images.

[0142] In the embodiments of this specification, the training device can generate a large number of new sample images based on limited annotated data (that is, sample certificate images), which can ensure that the information extraction model can be trained with sufficient training data, ensure that the information extraction model fully understands the certificate format and text semantics, and improves the accuracy of information extraction. By fine-tuning the pre-trained information extraction model with the sample certificate images and the newly added sample images, a model that can effectively extract global certificate information can be obtained. The global certificate refers to various types of certificates, such as identity cards, birth certificates, industrial and commercial business licenses, etc. In this way, the information extraction model can achieve a good certificate extraction effect using only a small number of annotated samples, and ensure good generalization for new certificates and a low delay in information extraction.

[0143] In one embodiment, before step 506 , the information extraction model training method provided in this specification further includes the following steps S4002 to S4010 .

[0144] S4002. Obtain corpus data from the information publishing platform.

[0145] The information publishing platform may include various websites on the Internet (such as news websites, popular science websites, and social networking websites, etc.), and may also include certain personal information publishing channels to which the training device has access rights. For example, the corpus data may include news corpus, commentary corpus, encyclopedia knowledge corpus, etc., and may also include a Chinese knowledge graph dataset. The corpus data can be collected directly from the information publishing platform by the training device, or it can be collected through other devices, which is not limited here. The training device can obtain multiple corpus data, such as the number of corpus data can reach millions.

[0146] S4004: Arrange the corpus data in a variety of layouts to obtain a plurality of sample images.

[0147] The multiple layout styles may be manually pre-configured, or may be automatically generated by a computing device analyzing the layout of images or pages containing information. For example, the layout style may include the layout style of information in a certificate, or the layout style of information in a table, document, slide, or web page.

[0148] The training device can obtain multiple typesetting methods and arrange the collected corpus data according to the typesetting methods. For example, each corpus can be arranged according to the multiple typesetting methods to obtain multiple sample images. Optionally, different corpus data can also be arranged according to different typesetting methods, which is not limited here. For example, the training device can print the corpus data onto a blank image according to various typesetting methods to obtain sample images. The typesetting method can also correspond to a document order, which can correspond to the user's reading order of the information under the typesetting method. The training device can print the corpus data according to the document order.

[0149] The corpus data collected by the training device may include attributes and attribute values ​​of multiple objects, and each object may have multiple attributes and corresponding attribute values. The training device may arrange the attributes and corresponding attribute values ​​of each object in a variety of layouts to obtain multiple sample images. For example, if the object is a certain person, the attributes of the object may include name, gender, ethnicity, place of origin, and age, etc. The training device may obtain a public knowledge graph dataset as corpus data, and based on the knowledge graph dataset, the attributes and corresponding attribute values ​​of multiple objects may be obtained. For example, a relatively large number of "attribute-value" pairs may be obtained, and each "attribute-value" pair includes an attribute of an object and a corresponding attribute value. The attribute can be analogous to the information type in the certificate image, and the attribute value can be analogous to the certificate information corresponding to the information type in the certificate image.

[0150] Optionally, the training device can also generate more sample images based on the obtained sample images. For example, the training device can adjust the relative positions of the attributes and their attribute values ​​in the sample images to obtain newly added sample images. The relative position adjustment can include size scaling and / or position offset for the area where the attribute is located or the area where the attribute value is located. For example, for at least one attribute in the sample image, the area where the attribute value corresponding to the attribute is located can be scaled, or the area where the attribute value is located can be offset in any direction relative to the area where the attribute is located. The size scaling can be proportional scaling in all directions, or it can be scaling only along a certain direction.

[0151] S4006: Identify the sample text in the sample image and the sample layout position information of the sample text in the sample image.

[0152] The training device can perform OCR on each sample image to obtain sample text and sample layout position information in the sample image. The sample layout position information can include information about the text detection frame. This step can refer to the relevant description of identifying text and layout position information in step 204, and the content previously described is not repeated here.

[0153] S4008. Based on the sample text and sample typesetting position information, the initial information extraction model is trained to obtain a pre-trained information extraction model.

[0154] In the embodiments of this specification, a training device can pre-train an initial information extraction model based on sample text and sample layout position information recognized for a sample image, so that the resulting pre-trained information extraction model can understand the semantics of the text and the relationship between the information to be extracted and the layout position. This pre-trained information extraction model can be further trained to be applicable to specific information extraction scenarios.

[0155] In a pre-training mode, the training device trains the initial information extraction model to enable it to have the ability to understand the semantics of the text. Step S4008 may include the following steps:

[0156] Perform mask processing on the sample text and sample layout position information to obtain pre-training data;

[0157] Input the pre-trained data into the initial information extraction model to obtain the prediction results of the complete sample text and the complete sample typesetting position information output by the initial information extraction model;

[0158] The prediction results are compared with the sample text and sample layout position information, and the parameters of the initial information extraction model are adjusted based on the comparison results to obtain a pre-trained information extraction model.

[0159] For example, the training device removes part of the sample text and part of the sample typesetting position information in the sample image, and uses the obtained incomplete sample text and sample typesetting position information as pre-training data. The complete sample text and sample typesetting position information before masking can be used as data labels for the pre-training data. The training device can input each pre-training data into the initial information extraction model to obtain the corresponding prediction results of the complete sample text and complete sample typesetting position information output by the initial information extraction model. For the process of adjusting the model parameters based on the comparison results here, please refer to the relevant introduction to the parameter adjustment of the information extraction model in step 508.

[0160] In another pre-training approach, a training device trains an initial information extraction model to enable it to understand the relationship between the information to be extracted and the layout position. The corpus data includes attributes and attribute values ​​of multiple objects. The training device can use sample images obtained based on these attributes and attribute values ​​to train the model.

[0161] In this pre-training method, step S4004 may include: arranging the attributes and attribute values ​​of each object according to a plurality of layouts to obtain a plurality of sample images;

[0162] Step S4008 may include:

[0163] Inputting the target attribute prompt information, sample text, and sample layout position information into the initial information extraction model to obtain the predicted attribute value output by the initial information extraction model; wherein the target attribute prompt information is used to describe the target attribute corresponding to the attribute value to be extracted in the sample image;

[0164] The predicted attribute value is compared with the attribute value of the target attribute in the sample image, and the parameters of the initial information extraction model are adjusted based on the comparison results to obtain a pre-trained information extraction model.

[0165] During the pre-training process, the training device can obtain target attribute prompt information, which is used to describe the target attribute corresponding to the attribute value to be extracted in the sample image. The target attribute prompt information is similar to the above-mentioned information extraction prompt information. Please refer to the above-mentioned introduction to the information extraction prompt information and will not be repeated here. The attribute value corresponding to the target attribute in the sample image can be used as a data label. The initial information extraction model can predict the attribute value corresponding to the target attribute and obtain the predicted attribute value output by the initial information extraction model. For the process of adjusting the model parameters based on the comparison results here, please refer to the relevant introduction to the parameter adjustment of the information extraction model in step 508.

[0166] In the embodiments of this specification, the training device may simultaneously use the above two pre-training methods to pre-train the initial information extraction model, or may use any one of the two pre-training methods to pre-train the initial information extraction model, which is not limited here.

[0167] S4010: Use the pre-trained information extraction model as the information extraction model to be trained.

[0168] The information extraction model to be trained is also the information extraction model to be trained in step 506 , and then step 506 can be executed to further train the model.

[0169] In summary, in the training method of the information extraction model provided in the embodiments of this specification, the sample text and its sample typesetting position information can be identified for the sample certificate image. The information extraction model is trained using the sample prompt information, sample text and sample typesetting position information, and the information extraction model can fuse the semantic features corresponding to the sample text and the typesetting features corresponding to the sample typesetting position information to obtain the text semantic typesetting fusion feature, and then extract the predicted certificate information based on the text semantic typesetting fusion feature and the sample prompt information. The information extraction model trained in this way can be applied to the certificate information extraction of various types of certificate images. Even for the certificate types that are not involved in the training of the information extraction model, the certificate information can be extracted with high accuracy by analyzing the relationship between the semantics of the text in the certificate image and its typesetting position, combined with the prompt information describing the information type of the certificate information.

[0170] The following is a detailed introduction to the training process of the information extraction model. Figure 6 This is a flow chart of another information extraction model training method provided in one embodiment of this specification. This method can be applied to a training device for an information extraction model, which is hereinafter referred to as a training device. Figure 6 As shown, the method may include:

[0171] Step 602: Acquire corpus data from the information publishing platform.

[0172] Step 604: Arrange the corpus data in a variety of layouts to obtain a plurality of sample images.

[0173] Step 606: Identify the sample text in the sample image and the sample layout position information of the sample text in the sample image.

[0174] Step 608: Input the sample text and sample typesetting position information into the initial information extraction model for model training to obtain a pre-trained information extraction model.

[0175] In the embodiment of this specification, steps 602 to 608 may refer to the relevant introduction in steps 402 to 408 and will not be repeated here.

[0176] After obtaining the pre-trained information extraction model, the training device can further fine-tune the information extraction model for a specific usage scenario (e.g., a license information recognition scenario) to obtain an information extraction model for license information recognition. For example, the training device can fine-tune the information extraction model using steps 610 through 616 below. Steps 610 through 616 can refer to the description of steps 502 through 508 above and will not be further detailed below.

[0177] Step 610: Obtain sample prompt information and a sample certificate image marked with a certificate information tag, wherein the sample prompt information is used to describe the information type of the certificate information to be extracted from the sample certificate image.

[0178] Step 612: Identify the sample text in the sample ID image and the sample layout position information of the sample text in the sample ID image.

[0179] Step 614: Analyze the sample prompt information, sample text, and sample typesetting position information through a pre-trained information extraction model to obtain the predicted license information extracted by the information extraction model; wherein, the information extraction model extracts the predicted license information based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the sample text and the typesetting features corresponding to the sample typesetting position information.

[0180] Step 616: Compare the predicted license information with the license information label, and adjust the parameters of the information extraction model based on the comparison result to obtain a trained information extraction model.

[0181] In summary, in the training method of the information extraction model provided in the embodiments of this specification, the sample text and its sample typesetting position information can be identified for the sample certificate image. The information extraction model is trained using the sample prompt information, sample text and sample typesetting position information, and the information extraction model can fuse the semantic features corresponding to the sample text and the typesetting features corresponding to the sample typesetting position information to obtain the text semantic typesetting fusion feature, and then extract the predicted certificate information based on the text semantic typesetting fusion feature and the sample prompt information. The information extraction model trained in this way can be applied to the certificate information extraction of various types of certificate images. Even for the certificate types that are not involved in the training of the information extraction model, the certificate information can be extracted with high accuracy by analyzing the relationship between the semantics of the text in the certificate image and its typesetting position, combined with the prompt information describing the information type of the certificate information.

[0182] Corresponding to the above method embodiment, this specification also provides an embodiment of a certificate information extraction device, Figure 7 This is a schematic diagram of the structure of a certificate information extraction device provided in one embodiment of this specification. Figure 7 As shown, the certificate information extraction device includes:

[0183] An acquisition module 701 is used to acquire a certificate image to be extracted;

[0184] Identification module 702, used to identify text in the certificate image and the layout position information of the text in the certificate image;

[0185] The information extraction module 703 is used to analyze the information extraction prompt information, text and typesetting position information through the information extraction model to obtain the certificate information to be extracted by the information extraction model; wherein, the information extraction prompt information is used to describe the information type of the certificate information to be extracted in the certificate image, and the information extraction model extracts the certificate information corresponding to the information type based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information.

[0186] Optionally, the information extraction model includes an encoder, an adaptation module, and a decoder, wherein the encoder includes a text semantic analysis module and a layout analysis module; the information extraction module 703 is used to:

[0187] Splicing the information extraction prompt information with the text to obtain text splicing features, and determining the position coding information corresponding to the text splicing features;

[0188] Based on the typesetting position information, determining the splicing position information corresponding to the text splicing feature, and determining the position coding information corresponding to the splicing position information;

[0189] Inputting the text splicing features and the position coding information corresponding to the text splicing features into the text semantic analysis module to obtain the semantic features corresponding to the text output by the text semantic analysis module;

[0190] Inputting the splicing position information and the position coding information corresponding to the splicing position information into the layout analysis module to obtain the layout features corresponding to the layout position information output by the layout analysis module;

[0191] The semantic features and typographic features are semantically fused through the adaptation module to obtain the text semantic typographic fusion features;

[0192] The text semantic typesetting fusion features are decoded through the decoder to obtain the certificate information to be extracted.

[0193] Optionally, the text semantic analysis module and the layout analysis module each include multiple transformer layers; and the certificate information extraction device further includes:

[0194] The information interaction module is used to perform information interaction through the transformer layer in the text semantic analysis module and the transformer layer in the layout analysis module, so that both the semantic features and the typesetting features contain the relationship information between the semantics of the text and the typesetting position.

[0195] Optionally, the certificate information extraction device further includes:

[0196] The information processing module is used to obtain prompt information from the certificate information extraction request sent by the client before analyzing the information extraction prompt information, text and typesetting position information through the information extraction model, and generate structured information extraction prompt information based on the prompt information according to the preset prompt information structure.

[0197] Corresponding to the above method embodiment, this specification also provides an embodiment of a certificate information extraction device, which is applied to a server. Figure 8 This is a schematic diagram of the structure of another device for extracting license information provided in one embodiment of this specification. Figure 8 As shown, the certificate information extraction device includes:

[0198] The acquisition module 801 is configured to obtain information extraction prompt information in response to a certificate information extraction request for a certificate image sent by a client; wherein the information extraction prompt information is used to describe the type of information to be extracted from the certificate information;

[0199] Identification module 802, used to identify text in the certificate image and the layout position information of the text in the certificate image;

[0200] The information extraction module 803 is configured to input the information extraction prompt information, text, and typesetting position information into the information extraction model to obtain the license information extracted by the information extraction model; wherein the information extraction model extracts the license information corresponding to the information type based on the text semantic typesetting fusion feature, which is obtained by fusing the semantic features corresponding to the text with the typesetting features corresponding to the typesetting position information;

[0201] The sending module 804 is used to send the certificate information extracted by the information extraction model to the client, so that the client can display the certificate information.

[0202] The above is a schematic diagram of a license information extraction device according to this embodiment. It should be noted that the technical solution of this license information extraction device and the technical solution of the license information extraction method described above are based on the same concept. For details not described in detail in the technical solution of the license information extraction device, please refer to the description of the technical solution of the license information extraction method described above.

[0203] Corresponding to the above method embodiment, this specification also provides an embodiment of a training device for an information extraction model. Figure 9 This is a structural diagram of a training device for an information extraction model provided in one embodiment of this specification. Figure 9 As shown, the training device of the information extraction model includes:

[0204] A first acquisition module 901 is configured to acquire sample prompt information and a sample license image labeled with a license information tag, wherein the sample prompt information is used to describe the type of license information to be extracted from the sample license image;

[0205] A first recognition module 902 is configured to recognize sample text in a sample ID image and sample layout position information of the sample text in the sample ID image;

[0206] A first information extraction module 903 is configured to analyze the sample prompt information, sample text, and sample typesetting position information using a to-be-trained information extraction model to obtain predicted license information extracted by the information extraction model. The information extraction model extracts the predicted license information to be extracted based on a text semantic typesetting fusion feature, which is obtained by fusing semantic features corresponding to the sample text with typesetting features corresponding to the sample typesetting position information.

[0207] The first parameter adjustment module 904 is used to compare the predicted license information with the license information label and adjust the parameters of the information extraction model based on the comparison result to obtain a trained information extraction model.

[0208] Optionally, the training device of the information extraction model further includes:

[0209] The second acquisition module is used to obtain corpus data from the information publishing platform before inputting the information extraction prompt information, text and typesetting position information into the information extraction model;

[0210] The data arrangement module is used to arrange the corpus data in a variety of layout methods to obtain multiple sample images;

[0211] A second recognition module is used to recognize sample text in the sample image and sample typesetting position information of the sample text in the sample image;

[0212] A pre-training module is used to train the initial information extraction model based on the sample text and sample typesetting position information to obtain a pre-trained information extraction model;

[0213] The first determining module is configured to use a pre-trained information extraction model as the information extraction model to be trained.

[0214] Optionally, pre-training modules are used to:

[0215] Perform mask processing on the sample text and sample layout position information to obtain pre-training data;

[0216] Input the pre-trained data into the initial information extraction model to obtain the prediction results of the complete sample text and the complete sample typesetting position information output by the initial information extraction model;

[0217] The prediction results are compared with the sample text and sample layout position information, and the parameters of the initial information extraction model are adjusted based on the comparison results to obtain a pre-trained information extraction model.

[0218] Optionally, the corpus data includes attributes and attribute values ​​of multiple objects; the data arrangement module is used to:

[0219] Arrange the attributes and attribute values ​​of each object in a variety of layouts to obtain multiple sample images;

[0220] Pre-training modules are used for:

[0221] Inputting the target attribute prompt information, sample text, and sample layout position information into the initial information extraction model to obtain the predicted attribute value output by the initial information extraction model; wherein the target attribute prompt information is used to describe the target attribute corresponding to the attribute value to be extracted in the sample image;

[0222] The predicted attribute value is compared with the attribute value of the target attribute in the sample image, and the parameters of the initial information extraction model are adjusted based on the comparison results to obtain a pre-trained information extraction model.

[0223] Optionally, the training device for the information extraction model further includes:

[0224] A third acquisition module is configured to acquire sample prompt information and a sample license image labeled with license information tags before inputting the information extraction prompt information, text, and layout position information into the information extraction model, wherein the sample prompt information is used to describe the type of license information to be extracted from the sample license image;

[0225] A third recognition module is used to recognize sample text in the sample certificate image and sample layout position information of the sample text in the sample certificate image;

[0226] A second information extraction module is configured to input the sample prompt information, sample text, and sample typesetting position information into the information extraction model to be trained, and obtain predicted license information extracted by the information extraction model; wherein the information extraction model extracts the predicted license information based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the sample text and the typesetting features corresponding to the sample typesetting position information;

[0227] The second parameter adjustment module is used to compare the predicted license information with the license information label, and adjust the parameters of the information extraction model based on the comparison results to obtain a trained information extraction model.

[0228] Optionally, the training device for the information extraction model further includes:

[0229] An image adjustment module, configured to adjust the content and / or position of the sample text in the sample certificate image to obtain a new sample image;

[0230] A second determining module is configured to determine the sample prompt information and the certificate information label of the newly added sample image based on the sample prompt information and the certificate information label of the sample certificate image;

[0231] The third information extraction module is used to analyze the sample prompt information, sample text and sample layout position information of the newly added sample image through the information extraction model to be trained, and obtain the predicted license information to be extracted by the information extraction model;

[0232] The third parameter adjustment module is used to compare the predicted license information of the newly added sample image with the license information label of the newly added license image, and adjust the parameters of the information extraction model based on the comparison result to obtain a trained information extraction model.

[0233] Optionally, the sample text in the sample certificate image includes multiple information types and corresponding certificate information, and the marked certificate information tag includes the certificate information corresponding to the target information type; the image adjustment module is configured to perform at least one of the following steps:

[0234] For any information type and corresponding certificate information in the sample certificate image, replace the certificate information with the alternative information of the information type to obtain a first newly added sample image;

[0235] Replacing any information type and corresponding certificate information in the sample certificate image with the alternative information type and corresponding certificate information to obtain a second newly added sample image;

[0236] Merging at least two adjacent text regions in the sample certificate image to obtain a third newly added sample image; wherein the text region includes the location of at least part of the text of any information type or certificate information;

[0237] The text area in the sample certificate image is shifted to obtain a fourth newly added sample image.

[0238] The above is a schematic diagram of a training device for an information extraction model according to this embodiment. It should be noted that the technical solution of the training device for this information extraction model is based on the same concept as the technical solution of the training method for the information extraction model described above. For details not described in detail in the technical solution for the training device for the information extraction model, please refer to the description of the technical solution for the training method for the information extraction model described above.

[0239] Figure 10 10 is a block diagram of a computing device according to an embodiment of the present disclosure. The computing device 1000 includes, but is not limited to, a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 via a bus 1030, and a database 1050 is used to store data.

[0240] The computing device 1000 also includes an access device 1040 that enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0241] In one embodiment of the present specification, the above components of the computing device 1000 and Figure 10 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 10 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0242] Computing device 1000 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1000 may also be a mobile or stationary server.

[0243] The processor 1020 is used to execute computer programs / instructions, which, when executed by the processor, implement the above Figure 2 or Figures 4 to 6 Any of the methods shown.

[0244] As for the computing device embodiment, since it is basically similar to the certificate information extraction method embodiment and the information extraction model training method embodiment, the description is relatively simple. For relevant details, please refer to the description of the method embodiment part.

[0245] One embodiment of this specification also provides a computer-readable storage medium, which stores a computer program / instruction, which implements the steps of the above-mentioned certificate information extraction method when executed by a processor. The computer program / instruction includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.

[0246] One embodiment of the present specification further provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when the computer program / instruction is executed in a processor.

[0247] As for the computer-readable storage medium embodiment and the computer program product embodiment, since they are basically similar to the certificate information extraction method embodiment and the information extraction model training method embodiment, the description is relatively simple. For relevant details, please refer to the description of the certificate information extraction method embodiment and the information extraction model training method embodiment.

[0248] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0249] It should be noted that the above description is of a specific embodiment of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0250] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0251] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.< / spot> < / spot> < / spot> < / spot>

Claims

1. A method for extracting certificate information, comprising: Obtaining a certificate image for information extraction; Identifying text in the certificate image and typeset position information of the text in the certificate image; The information extraction prompt information, the text and the typesetting position information are analyzed by an information extraction model to obtain the certificate information to be extracted by the information extraction model; wherein, the information extraction prompt information is used to describe the information type of the certificate information to be extracted in the certificate image, and the information extraction model extracts the certificate information corresponding to the information type based on the text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic feature corresponding to the text and the typesetting feature corresponding to the typesetting position information.

2. The method according to claim 1, wherein The information extraction model includes an encoder, an adaptation module, and a decoder. The encoder includes a text semantic analysis module and a layout analysis module. The information extraction model analyzes the information extraction prompt information, the text, and the typesetting position information to obtain the license information extracted by the information extraction model, including: Splicing the information extraction prompt information with the text to obtain text splicing features, and determining position coding information corresponding to the text splicing features; Based on the typesetting position information, determining the splicing position information corresponding to the text splicing feature, and determining the position coding information corresponding to the splicing position information; Inputting the text splicing features and the position coding information corresponding to the text splicing features into the text semantic analysis module to obtain the semantic features corresponding to the text output by the text semantic analysis module; Inputting the splicing position information and the position coding information corresponding to the splicing position information into the layout analysis module, and obtaining the typesetting features corresponding to the typesetting position information output by the layout analysis module; The adaptation module semantically fuses the semantic features and the typesetting features to obtain a text semantic typesetting fusion feature; The text semantic typesetting fusion feature is decoded by the decoder to obtain the certificate information to be extracted.

3. The method according to claim 2, wherein: The text semantic analysis module and the layout analysis module both include multiple transformer layers; the method further includes: Information is exchanged through the transformer layer in the text semantic analysis module and the transformer layer in the layout analysis module, so that both the semantic features and the typesetting features contain the relationship information between the semantics of the text and the typesetting position.

4. The method according to claim 1, before analyzing the information extraction prompt information, the text, and the typesetting position information using the information extraction model, further comprising: Prompt information is obtained from a certificate information extraction request sent by a client, and structured information extraction prompt information is generated based on the prompt information according to a preset prompt information structure.

5. A method for training an information extraction model, comprising: Obtaining sample prompt information and a sample certificate image marked with a certificate information tag, wherein the sample prompt information is used to describe the information type of the certificate information to be extracted from the sample certificate image; Identifying sample text in the sample ID image and sample layout position information of the sample text in the sample ID image; The sample prompt information, the sample text, and the sample typeset position information are analyzed by the information extraction model to be trained to obtain the predicted license information to be extracted by the information extraction model; wherein the information extraction model extracts the predicted license information based on text semantic typeset fusion features, and the text semantic typeset fusion features are obtained by fusing semantic features corresponding to the sample text and typeset features corresponding to the sample typeset position information; The predicted license information is compared with the license information label, and the parameters of the information extraction model are adjusted based on the comparison result to obtain a trained information extraction model.

6. The method according to claim 5, wherein: Before analyzing the sample prompt information, the sample text, and the sample typesetting position information by the information extraction model to be trained, the method further includes: Obtain corpus data from the information publishing platform; Arranging the corpus data in a plurality of typeset modes to obtain a plurality of sample images; Identifying sample text in the sample image and sample layout position information of the sample text in the sample image; Based on the sample text and the sample typesetting position information, the initial information extraction model is trained to obtain a pre-trained information extraction model; The pre-trained information extraction model is used as the information extraction model to be trained.

7. The method according to claim 6, wherein: The initial information extraction model is trained based on the sample text and the sample typesetting position information to obtain a pre-trained information extraction model, including: Performing mask processing on the sample text and the sample typesetting position information to obtain pre-training data; Inputting the pre-trained data into an initial information extraction model to obtain prediction results of the complete sample text and complete sample typesetting position information output by the initial information extraction model; The prediction result is compared with the sample text and the sample typesetting position information, and the parameters of the initial information extraction model are adjusted based on the comparison result to obtain a pre-trained information extraction model.

8. The method according to claim 6, wherein: The corpus data includes attributes and attribute values ​​of multiple objects; the corpus data is arranged in multiple typeset ways to obtain multiple sample images, including: Arrange the attributes and attribute values ​​of each object in a variety of layouts to obtain multiple sample images; The initial information extraction model is trained based on the sample text and the sample typesetting position information to obtain a pre-trained information extraction model, including: Inputting the target attribute prompt information, the sample text, and the sample typesetting position information into an initial information extraction model to obtain a predicted attribute value output by the initial information extraction model; wherein the target attribute prompt information is used to describe the target attribute corresponding to the attribute value to be extracted in the sample image; The predicted attribute value is compared with the attribute value of the target attribute in the sample image, and the parameters of the initial information extraction model are adjusted based on the comparison result to obtain a pre-trained information extraction model.

9. The method according to any one of claims 5 to 8, further comprising: Adjusting the content and / or position of the sample text in the sample certificate image to obtain a new sample image; Determining the sample prompt information and the certificate information label of the newly added sample image based on the sample prompt information and the certificate information label of the sample certificate image; Analyzing the sample prompt information, sample text, and sample layout position information of the newly added sample image through the information extraction model to be trained, and obtaining the predicted license information to be extracted by the information extraction model; The predicted license information of the newly added sample image is compared with the license information label of the newly added license image, and the parameters of the information extraction model are adjusted based on the comparison result to obtain a trained information extraction model.

10. The method according to claim 9, wherein: The sample text in the sample certificate image includes multiple information types and corresponding certificate information, and the marked certificate information tag includes the certificate information corresponding to the target information type; and adjusting the content and / or position of the sample text in the sample certificate image to obtain a new sample image includes at least one of the following steps: For any information type in the sample certificate image and its corresponding certificate information, replace the certificate information with the alternative information of the certificate information to obtain a first newly added sample image; Replacing any information type and corresponding certificate information in the sample certificate image with the alternative information type and corresponding certificate information to obtain a second newly added sample image; Merging at least two adjacent text regions in the sample certificate image to obtain a third newly added sample image; wherein the text region includes any information type or the location of at least part of the text in the certificate information; The text area in the sample certificate image is shifted to obtain a fourth newly added sample image.

11. A method for extracting license information, applied to a server, comprising: In response to a certificate information extraction request for a certificate image sent by a client, obtaining information extraction prompt information; wherein the information extraction prompt information is used to describe the information type of the certificate information to be extracted; Identifying text in the certificate image and typeset position information of the text in the certificate image; The information extraction prompt information, the text, and the typesetting position information are analyzed by an information extraction model to obtain the license information to be extracted by the information extraction model; wherein the information extraction model extracts the license information corresponding to the information type based on a text semantic typesetting fusion feature, and the text semantic typesetting fusion feature is obtained by fusing the semantic features corresponding to the text and the typesetting features corresponding to the typesetting position information; The certificate information to be extracted, which is extracted by the information extraction model, is sent to the client, so that the client can display the certificate information.

12. A certificate information extraction system comprising: Client and server; The client is used to: send a certificate information extraction request to the server; The server is configured to: execute the method according to any one of claims 1 to 11 based on the certificate information extraction request, extract the certificate information to be extracted, and send the certificate information to be extracted to the client; The client is also used to display the received certificate information.

13. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program / instructions, which implements the method according to any one of claims 1 to 11 when the computer program / instructions are executed in a processor.

Citation Information

Patent Citations

  • License key information extraction method based on small sample data

    CN116229494A

  • Identity card information extraction method and device, computer equipment and storage medium

    CN117079292A

  • OCR image sample generation method and apparatus, print font verification method and apparatus, and device and medium

    WO2021212658A1

  • AU2016225819A1