Method, apparatus, electronic device and readable storage medium for extracting table information
By constructing a tabular information extraction model of the multimodal data encoding layer and decoding layer, the problem of poor universality of the tabular information extraction method in the prior art is solved, and efficient information extraction of complex and diverse tables is achieved.
Patent Information
- Application Number
- CN202310193188.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-02-20
AI Technical Summary
In the prior art, the table information extraction method is poor in versatility and is difficult to adapt to complex and diverse table structures and layouts.
The multimodal data encoding layer, entity type decoding layer and entity relationship decoding layer are used to construct a table information extraction model, and the entity type and relationship decoding is realized through the fusion and encoding of image vectors, entity text vectors and entity coordinate vectors.
It realizes information extraction of various style tables, improves extraction efficiency and accuracy, and has strong universality, generalization and scalability.
Smart Images

Figure CN116343248B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and readable storage medium for extracting table information. Background Art
[0002] Table data is an inevitable information storage method. In actual business, the structure and layout of tables are often complex and diverse, which brings great difficulties to the extraction of table information.
[0003] In the prior art, table information can be extracted by manual processing, but this method usually requires a large amount of manpower and material resources for information comparison and processing; the entity element extraction method based on rules can also be used to extract table entity information, but this method can generally only perform entity element extraction with a fixed style and cannot extract the relationship information between entity elements; the table entity information can also be extracted by using a Named Entity Recognition (NER) model. Although this method has no strict requirements on the table style, it has a great dependence on the text features of the table. If a certain text type is not covered in the training data, the prediction ability of the model for this type of entity element is poor. Therefore, there is an urgent need for a table information extraction method with strong versatility. Summary of the Invention
[0004] In view of the above situation, embodiments of the present application provide a method, device, electronic device, and readable storage medium for extracting table information, aiming to solve the problem of poor versatility of the table information extraction method in the prior art.
[0005] In a first aspect, embodiments of the present application provide a method for extracting table information, which is implemented by a table information extraction model. The table information extraction model includes a multi-modal data encoding layer, an entity type decoding layer, and an entity relationship decoding layer. Among them, the multi-modal data encoding layer is respectively connected to the entity type decoding layer and the entity relationship decoding layer;
[0006] The method includes:
[0007] Preprocess the table image to obtain an image vector, entity text vectors respectively corresponding to a plurality of target entities, and entity coordinate vectors;
[0008] Based on the multi-modal data encoding layer, fuse and encode the image vector, the entity text vector, and the entity coordinate vector to obtain a multi-modal feature vector;
[0009] Based on the entity type decoding layer, perform type decoding on the multi-modal feature vector to obtain the entity types corresponding to each of the target entities, where the entity types include at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair;
[0010] Based on the entity relationship decoding layer, perform relationship decoding on the multi-modal feature vector to obtain the entity relationships of at least one target entity pair formed by the multiple target entities.
[0011] In a second aspect, an embodiment of the present application further provides a table information extraction device, which is deployed with a table information extraction model. The table information extraction model includes a multi-modal data encoding layer, an entity type decoding layer, and an entity relationship decoding layer. Among them, the multi-modal data encoding layer is respectively connected to the entity type decoding layer and the entity relationship decoding layer;
[0012] The device includes:
[0013] A preprocessing unit, configured to preprocess the table image to obtain an image vector, entity text vectors and entity coordinate vectors respectively corresponding to multiple target entities;
[0014] A multi-modal data encoding unit, configured to fuse and encode the image vector, the entity text vector, and the entity coordinate vector based on the multi-modal data encoding layer to obtain a multi-modal feature vector;
[0015] A type decoding unit, configured to perform type decoding on the multi-modal feature vector based on the entity type decoding layer to obtain the entity types corresponding to each of the target entities, where the entity types include at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair;
[0016] A relationship decoding unit, configured to perform relationship decoding on the multi-modal feature vector based on the entity relationship decoding layer to obtain the entity relationships of at least one target entity pair formed by the multiple target entities.
[0017] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the steps of the above table information extraction method.
[0018] Fourthly, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is enabled to execute the steps of the above-mentioned method for extracting tabular information.
[0019] The above at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects:
[0020] The method for extracting tabular information provided by the embodiment of the present application constructs a tabular information extraction model. First, preprocess the tabular image to obtain an image vector, entity text vectors and entity coordinate vectors respectively corresponding to a plurality of target entities; use the multi-modal data encoding layer of the extraction model to fuse and encode the image vector, entity text vector, and entity coordinate vector to obtain a multi-modal feature vector; then, based on the entity type decoding layer and entity relationship decoding layer of the extraction model, perform type decoding and relationship decoding on the multi-modal feature vector respectively, so as to obtain the entity type corresponding to each target entity and the entity relationship of at least one pair of target entities. The entity type includes at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair. It can be seen that according to the two factors of layout and semantics, the embodiment of the present application divides the entity types of each target entity in various styles of tables into multiple categories. The model learns multi-modal features such as the image visual features, text semantic features, and spatial layout features of the tabular image, uses the complementarity between multiple modalities, eliminates the redundancy between modalities, and thus learns better feature representations, so as to achieve the purpose of understanding the table, and further can identify the entity category of each target entity and the relationship between target entities. Therefore, it is applicable to the extraction of information from various styles of tables, with strong versatility, generalization, and scalability, and greatly solves the problem of difficult extraction of diverse tabular information; in terms of process, it can realize end-to-end automatic extraction of tabular information and improve the extraction efficiency of tabular information. Description of the Drawings
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0022] Figure 1 Shows a schematic flowchart of a method for extracting tabular information according to an embodiment provided by the present application;
[0023] Figure 2 Shows a schematic structural diagram of a tabular information extraction model according to an embodiment provided by the present application;
[0024] Figure 3 The structural schematic diagram of a table-annotated image according to an embodiment provided by the present application is shown;
[0025] Figure 4 The flowchart of a method for extracting table information according to another embodiment provided by the present application is shown;
[0026] Figure 5 The structural schematic diagram of a device for extracting table information according to an embodiment provided by the present application is shown;
[0027] Figure 6 The structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. Detailed implementation manners
[0028] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0029] The technical solutions provided by each embodiment of the present application will be described in detail below with reference to the drawings.
[0030] Before introducing the embodiments of the present application in detail, the following technical terms will be introduced.
[0031] NER: It refers to identifying entities with specific meanings in text, mainly including proper nouns such as person names and place names. The main technical idea of named entity recognition is to identify the entity boundaries and entity categories in text through a deep learning model.
[0032] Multi-modal representation: Unimodal representation refers to representing information as a numerical vector that can be processed by a computer or further abstracted into a higher-level feature vector, while multi-modal representation refers to learning better feature representations by utilizing the complementarity between multiple modalities (such as images, text, speech, etc.) and eliminating the redundancy between modalities.
[0033] Tabular data is an inevitable way of information storage. In actual business, the structure and layout of tables are often complex and diverse, which brings great difficulties to the extraction of table information. In the prior art, manual processing can be used to extract table information, but this method usually requires a large amount of manpower and material resources for information comparison and processing. It is also possible to use a rule-based entity element extraction method to extract table entity information, which mainly combines the text information of the entity elements to be extracted and a fixed format for extraction. For example, to extract the entity element of "contract number", its approximate area can often be located through the text information of "contract number", and then combined with the characteristic that the "contract number" is generally composed of numbers, letters, and special symbols, the entity element of "contract number" that meets this format can be extracted. However, this method can generally only extract entity elements with a fixed style, and cannot extract the relationship information between entity elements. It has poor generalization and scalability, and it is usually difficult to cover all possible formats and styles in actual business, so the extraction accuracy is often not high.
[0034] It is also possible to use an NER model to extract table entity information. The main technical idea is to label the entity elements to be extracted in the training data so that the model can learn the characteristics of these entity elements, so as to achieve the ability to predict the entity elements to be extracted. For example, let the model learn that the "contract number" has the characteristic of being composed of numbers, letters, and special symbols, so that the "contract number" can be recognized based on this characteristic during model prediction. Although this method has no strict requirements on the table style, it has a large dependence on the text characteristics of the table. If a certain text type is not covered in the training data, the model's prediction ability for this type of entity element is poor. If it is to be migrated to a new sample, generally, the data needs to be re-labeled and the model needs to be re-trained. It is also difficult to ensure whether the extraction effect of the previous sample type data will be affected during this process, and it is easy to have the problem of one rising while the other falls. This method has poor generalization and scalability. The above-mentioned rule-based entity element extraction method and the NER model-based method are both developed based on the types of entity elements specified by the business side to be extracted. Therefore, not all elements in the table can be extracted. When the types of elements in the business requirements increase, a new extraction plan needs to be formulated from an engineering perspective.
[0035] Based on this, the present invention proposes a method for extracting table information. In the embodiments of the present application, according to two factors of layout and semantics, the entity types of each target entity in various styles of tables are divided into multiple categories. The model learns multi-modal features such as the image visual features, text semantic features, and spatial layout features of the table image, and utilizes the complementarity between multiple modalities to eliminate the redundancy between modalities, thereby learning better feature representations, so as to achieve the purpose of understanding the table, and further can identify the entity categories of each target entity and the relationships between target entities. Therefore, it is applicable to the information extraction of various styles of tables, with strong generality, generalization, and scalability, and greatly solves the problem of difficult extraction of diverse table information; in terms of process, it can realize end-to-end automated table information extraction and improve the extraction efficiency of table information.
[0036] The following describes the present application in detail through specific embodiments.
[0037] Figure 1 The flowchart of the method for extracting table information according to an embodiment provided by the present application is shown. From Figure 1 it can be seen that the present application at least includes steps S101 to S104:
[0038] Step S101: Preprocess the table image to obtain an image vector, entity text vectors and entity coordinate vectors respectively corresponding to multiple target entities.
[0039] The method for extracting table information of the present application is implemented based on a table information extraction model. Figure 2 The structural diagram of a table information extraction model according to an embodiment provided by the present application is shown. From Figure 2 it can be seen that the table information extraction model 200 includes a multi-modal data encoding layer 201, an entity type decoding layer 202, and an entity relationship decoding layer 203. Among them, the multi-modal data encoding layer 201 is respectively connected to the entity type decoding layer 202 and the entity relationship decoding layer 203.
[0040] In this embodiment, the table image can be a picture in any format, such as jpg format, png format, jpeg format, etc., or a picture generated from a file in any format. For example, a pdf file or a word file containing table data can be converted into a picture.
[0041] The target entity can be the text in each cell of the table. Figure 3 The structural diagram of a table annotation image according to an embodiment provided by the present application is shown. Taking Figure 3 as an example. Figure 3The table in includes 15 target entities, namely: payment status table, project name, A, previous period cumulative, application amount, total application amount, 2, actual payment amount, 3, deductible amount, water fee, 0, reviewer, C, total: 5.
[0042] The table image is preprocessed to obtain an image vector, entity text vectors corresponding to multiple target entities, and entity coordinate vectors. For example, in some embodiments, the table image can be encoded using the image processing function of the TensorFlow tool to obtain an image vector; the table image is processed by optical character recognition (OCR) to obtain a text A, such as "Payment Status Table Project Name A Last Period Cumulative Application Amount Application Amount 2 Actual Payment Amount 3 Deductible Amount Water Fee 0 Auditor C Total: 5", and the coordinate A corresponding to each character in the text A, such as, 付((6,0),(7,-1)), 款((7,0),(8,-1)), ..., 5((16,-5),(17,-6)), where the coordinate A is the coordinate of the upper left corner point and the coordinate of the lower right corner point of each character. Then perform word segmentation on the text A. For example, the jieba word segmentation tool can be used to perform word segmentation, and the word segmentation results are obtained, "Payment Status Table / Project Name / A / Previous Period Cumulative / Application Amount / Total Application Amount / 2 / Actual Payment Amount / 3 / Deductible Amount / Water Fee / 0 / Reviewer / C / Total: 5", and the coordinates B corresponding to each word, such as the payment status table ((6, 0), (11, -1)), etc. Finally, the word segmentation results and the coordinates B corresponding to each word are vectorized. For example, the Word2vec model can be used for vectorization, so as to obtain entity text vectors and entity coordinate vectors corresponding to multiple target entities. Among them, the dimensions of the entity text vector, the entity coordinate vector and the image vector are the same, for example, they can be 512 dimensions, which is not limited in this application.
[0043] In other embodiments, the table image may be first corrected in direction and angle to obtain a positive corrected table image, and then the corrected table image may be preprocessed to obtain an image vector, entity text vectors corresponding to multiple target entities, and an entity coordinate vector.
[0044] It can be seen from this embodiment that by correcting the direction and angle of the table image, the influence of subsequent operations such as OCR recognition, word segmentation and vectorization on the table image can be avoided, so that more accurate multimodal feature data can be obtained to improve the accuracy of the model output.
[0045] Step S102: Based on the multimodal data encoding layer, the image vector, the entity text vector, and the entity coordinate vector are fused and encoded to obtain a multimodal feature vector.
[0046] In this embodiment, the image vector, the entity text vector, and the entity coordinate vector can be fused first. For example, the three vectors can be concatenated or an add operation can be performed on the three vectors to obtain a fused feature vector, and then the fused feature vector is input into the multi-modal data encoding layer. The multi-modal data encoding layer learns the features of the text, layout, and image of the table image to obtain a multi-modal feature vector.
[0047] Step S103: Based on the entity type decoding layer, perform type decoding on the multi-modal feature vector to obtain the entity type corresponding to each target entity. The entity type includes at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair.
[0048] Figure 3 The text within the dashed box is the entity type corresponding to each target entity. The entity type is divided from two aspects: the layout and semantics of the target entity. For example, the entity type corresponding to the target entity "Payment Situation Table" is the file name; the entity types corresponding to the target entities "Project Name", "Total Applied Amount", "Actual Paid Amount", and "Water Fee" are horizontal entity names; the entity type corresponding to the target entity "A" is horizontal entity content; the entity types corresponding to the target entities "Previous Period Accumulated" and "Reviewer" are vertical entity names; the entity types corresponding to the target entities "Applied Amount" and "Deduction Amount" are titles; the entity types corresponding to the target entities "2", "3", and "0" are cross-entity content; the entity type corresponding to the target entity "C" is vertical entity content; the entity type corresponding to the target entity "Total: 5" is the name-content pair.
[0049] After obtaining the multi-modal feature vector, the multi-modal feature vector can be input into the entity type decoding layer. The entity type decoding layer performs type decoding on the multi-modal feature vector to obtain at least one entity type token corresponding to each target entity, and parses and converts each token to obtain the real entity type. For example, in some embodiments, Table 1 below is a workload statistics table. After obtaining the multi-modal feature vector corresponding to Table 1, the multi-modal feature vector can be input into the entity type decoding layer to obtain that the entity type corresponding to the target entity "Workload Statistics Table" is the file name.
[0050] Table 1
[0051] Workload Statistics Table
[0052] In some other embodiments, Table 2 below is a statistical table of project names. After obtaining the multi-modal feature vectors corresponding to Table 2, the multi-modal feature vectors can be input into the entity type decoding layer to obtain that the entity type corresponding to the target entity "project name" is the horizontal entity name, and the entity type corresponding to the target entity "A" is the horizontal entity content.
[0053] Table 2
[0054] Project Name A
[0055] In some further embodiments, the Figure 3 corresponding table image can be preprocessed, fused, and encoded to obtain multi-modal feature vectors. The multi-modal feature vectors are input into the entity type decoding layer to obtain the 8 types of entity types corresponding to 15 target entities as shown in Figure 3 the figure.
[0056] Step S104: Based on the entity relationship decoding layer, perform relationship decoding on the multi-modal feature vectors to obtain the entity relationships of at least one target entity pair formed by multiple target entities.
[0057] While performing type decoding on the multi-modal feature vectors, the multi-modal feature vectors can also be input into the entity relationship decoding layer to perform relationship decoding on the multi-modal feature vectors, obtain the entity relationship tokens of at least one target entity pair formed by multiple target entities, and parse and transform each token to obtain the real entity relationships.
[0058] In some embodiments of the present application, in the above method, the entity relationships include at least one of the following: horizontal entity name - horizontal entity content, vertical entity name - vertical entity content, horizontal entity name - cross-entity content, and vertical entity name - cross-entity content.
[0059] See Figure 3As shown, the entity relationship between the target entity "Project Name" and the target entity "A" is Horizontal Entity Name - Horizontal Entity Content; the entity relationship between the target entity "Total Applied Amount" and the target entity "2" is Horizontal Entity Name - Cross-Entity Content; the entity relationship between the target entity "Actual Paid Amount" and the target entity "3" is Horizontal Entity Name - Cross-Entity Content; the entity relationship between the target entity "Water Fee" and the target entity "0" is Horizontal Entity Name - Cross-Entity Content; the entity relationship between the target entity "Previous Period Accumulative" and the target entity "2" is Vertical Entity Name - Cross-Entity Content; the entity relationship between the target entity "Previous Period Accumulative" and the target entity "3" is Vertical Entity Name - Cross-Entity Content; the entity relationship between the target entity "Previous Period Accumulative" and the target entity "0" is Vertical Entity Name - Cross-Entity Content; the entity relationship between the target entity "Reviewer" and the target entity "C" is Vertical Entity Name - Vertical Entity Content.
[0060] In some embodiments, after obtaining the multi-modal feature vector corresponding to Table 2, the multi-modal feature vector can be input into the entity relationship decoding layer to decode the relationship of the multi-modal feature vector, and the entity relationship of a target entity pair (Project Name, A) is obtained as Horizontal Entity Name - Horizontal Entity Content.
[0061] In some other embodiments, the following Table 3 is the report card of Student A. After obtaining the multi-modal feature vector corresponding to Table 3, the multi-modal feature vector can be input into the entity relationship decoding layer to decode the relationship of the multi-modal feature vector, and two entity relationships are obtained: the entity relationship of the target entity pair (Mathematics, 90) is Horizontal Entity Name - Horizontal Entity Content; the entity relationship of the target entity pair (Chinese, 85) is Horizontal Entity Name - Horizontal Entity Content.
[0062] Table 3
[0063]
[0064] In still some other embodiments, it is possible to Figure 3 perform preprocessing, fusion, and encoding on the corresponding table image to obtain a multi-modal feature vector, input the multi-modal feature vector into the entity relationship decoding layer, and obtain the entity relationships of 8 target entity pairs formed by 15 target entities as shown in Figure 3 .
[0065] From Figure 1As can be seen from the method shown, the method for extracting tabular information provided by the embodiments of the present application constructs a tabular information extraction model. First, the table image is preprocessed to obtain an image vector, entity text vectors and entity coordinate vectors corresponding to multiple target entities respectively. Using the multi-modal data encoding layer of the extraction model, the image vector, entity text vector, and entity coordinate vector are fused and encoded to obtain a multi-modal feature vector. Then, based on the entity type decoding layer and entity relationship decoding layer of the extraction model, the multi-modal feature vector is respectively decoded for type and relationship, so as to obtain the entity type corresponding to each target entity and the entity relationship of at least one target entity pair. The entity type includes at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair. It can be seen that according to the two factors of layout and semantics, the embodiments of the present application divide the entity types of each target entity in various styles of tables into multiple categories. The model learns multi-modal features such as the image visual features, text semantic features, and spatial layout features of the table image, utilizes the complementarity between multiple modalities, and eliminates the redundancy between modalities, so as to learn better feature representations, thereby achieving the purpose of understanding the table, and further being able to identify the entity category of each target entity and the relationship between target entities. Therefore, it is applicable to the information extraction of various styles of tables, with strong generality, generalization, and scalability, and greatly solves the problem of difficult extraction of diverse tabular information. In terms of the process, it can realize end-to-end automatic tabular information extraction and improve the extraction efficiency of tabular information.
[0066] In some embodiments of the present application, in the above method, the method further includes: in response to an entity type input instruction, determining at least one first entity text corresponding to the entity type input instruction and returning it; in response to an entity text input instruction, determining the second entity text specified by the entity text input instruction, and determining and returning a third entity text corresponding to the second entity text according to the entity relationship to which the second entity text belongs.
[0067] The following combines Figure 3 to give an exemplary illustration of this embodiment.
[0068] It is possible to determine at least one first entity text corresponding to the entity type input instruction and return it in response to the entity type input instruction. For example, in some embodiments, if the entity type input instruction is "name-content pair", the first entity text corresponding to "name-content pair" can be determined as "Total: 5". In other embodiments, if the entity type input instruction is "title", the first entity texts corresponding to "title" can be determined as "Applied Amount" and "Deduction Amount".
[0069] In response to an entity text input instruction, the second entity text specified by the entity text input instruction can be determined, and according to the entity relationship to which the second entity text belongs, the third entity text corresponding to the second entity text can be determined and returned. For example, in some embodiments, Figure 3 The corresponding entity relationships include: (Project Name, A): (Horizontal Entity Name, Horizontal Entity Content), (Total Applied Amount, 2): (Horizontal Entity Name, Cross-Entity Content), (Actual Paid Amount, 3): (Horizontal Entity Name, Cross-Entity Content), (Water Fee, 0): (Horizontal Entity Name, Cross-Entity Content), (Previous Period Cumulative, 2): (Vertical Entity Name, Cross-Entity Content), (Previous Period Cumulative, 3): (Vertical Entity Name, Cross-Entity Content), (Previous Period Cumulative, 0): (Vertical Entity Name, Cross-Entity Content), (Reviewer, C): (Vertical Entity Name, Vertical Entity Content). The model can obtain the understanding result of the table according to the above relationships: (Project Name, A), (Total Applied Amount, 2), (Actual Paid Amount, 3), (Water Fee, 0), (Previous Period Cumulative, 2), (Previous Period Cumulative, 3), (Previous Period Cumulative, 0), (Reviewer, C). There is a key-value relationship in each element of the understanding result. If the second entity text specified by the entity text input instruction is "Project Name", then according to the entity relationship to which the second entity text belongs, the third entity text corresponding to "Project Name" can be determined as "A" and returned.
[0070] In some embodiments of the present application, in the above method, the preprocessing of the table image to obtain an image vector, entity text vectors corresponding to multiple target entities, and entity coordinate vectors includes: performing encoding processing on the table image to obtain the image vector; performing text recognition processing on the table image to obtain text information and coordinate information, where the coordinate information is the coordinate of each character in the text information; performing word segmentation processing on the text information to obtain the entity text and entity coordinates corresponding to the multiple target entities respectively, where the entity coordinates are the coordinates of each entity text; and performing vectorization processing on the entity text and the entity coordinates to obtain the entity text vector and the entity coordinate vector.
[0071] In this embodiment, the table image can be encoded to obtain an image vector; the table image can be subjected to text recognition to obtain text information and coordinate information; the text information can be segmented to obtain the entity text and entity coordinates corresponding to multiple target entities respectively; and the entity text and entity coordinates can be vectorized to obtain the entity text vector and the entity coordinate vector.
[0072] For example, in some embodiments, a deep convolutional neural network can be used to encode a table image to obtain an image vector. Perform OCR recognition on the table image to obtain a piece of text B, such as "Leave Statistics Table Employee A None Employee B None Employee C None", and the coordinates C corresponding to each character in the text B, such as, please ((5, 0), length, width), leave ((6, 0), length, width), …, 21 ((9, -3), length, width), where the coordinate C is the coordinate of the upper left corner point of each character and the length and width of the character. Then perform word segmentation on the text B, for example, the Chinese Lexical Analysis (LAC) word segmentation tool can be used for word segmentation to obtain the word segmentation result "Leave Statistics Table / Employee A / None / Employee B / None / Employee C / None", and the coordinates D corresponding to each word, such as Leave Statistics Table ((5, 0), length×5, width). Finally, perform vectorization processing on the word segmentation result and the coordinates D corresponding to each word, for example, the Word2vec model can be used for vectorization processing, so as to obtain the entity text vectors and entity coordinate vectors corresponding to multiple target entities respectively.
[0073] In some embodiments of the present application, in the above method, the table information extraction model further includes a multi-modal pre-training layer; the method further includes: based on the multi-modal pre-training layer, processing the image vector, the entity text vector, and the entity coordinate vector to obtain a multi-modal initial feature vector; using the multi-modal initial feature vector as the input of the multi-modal data encoding layer.
[0074] As Figure 2 shown, the table information extraction model 200 may further include a multi-modal pre-training layer 204.
[0075] In this embodiment, based on the multi-modal pre-training layer, the image vector, the entity text vector, and the entity coordinate vector can be processed to obtain a multi-modal initial feature vector; the multi-modal initial feature vector is input into the multi-modal data encoding layer to obtain a multi-modal feature vector.
[0076] In some embodiments of the present application, in the above method, the multi-modal pre-training layer is generated based on the LayoutXLM model. LayoutXLM is a multi-modal pre-training model for multi-language document understanding, aiming to bridge the language barrier for visually rich document understanding. LayoutLM jointly models and trains text, layout, and image information in a unified framework, so as to better learn the associations between different modalities.
[0077] In some embodiments, an entity text vector and an entity coordinate vector can be input into a multi-modal pre-trained layer generated based on the LayoutXLM model to obtain an output vector. The output vector is added to an image vector to obtain a multi-modal initial feature vector. Then, the multi-modal initial feature vector can be input into a multi-modal data encoding layer to obtain a multi-modal feature vector.
[0078] As can be seen from the above embodiments, by adding the multi-modal pre-trained layer, the prior knowledge of the multi-modal pre-trained layer is increased, so that the multi-modal feature vector can more accurately represent the features of the table image, thereby greatly improving the prediction effect of the table information extraction model.
[0079] In some embodiments of the present application, in the above method, the table information extraction model is trained according to the following method: constructing a training sample set; wherein, the training sample set includes multiple groups of multi-modal data and corresponding sample labels, wherein the multi-modal data includes text data, coordinate data, and image data, and the sample labels include entity type labels and entity relationship labels; obtaining an initial table information extraction model; inputting the training sample set into the initial table information extraction model to obtain entity type prediction values and entity relationship prediction values; based on the entity type prediction values and the entity type labels, according to the first cross-entropy loss function, obtaining a first loss value; based on the entity relationship prediction values and the entity relationship labels, according to the second cross-entropy loss function, obtaining a second loss value; performing weighted summation on the first loss value and the second loss value to obtain a total loss value; and training the initial table information extraction model multiple times according to the total loss value and a preset threshold to obtain the table information extraction model.
[0080] Construct a training sample set. Among them, the training sample set includes multiple groups of multi-modal data and corresponding sample labels. The multi-modal data includes text data, coordinate data, and image data. The sample labels include entity type labels and entity relationship labels. In some embodiments, the training sample set can be defined as Z = {X t ,X p ,X i}, where X t = {x t1 ,x t2 ,...,x tn}, X p = {x p1 ,x p2 ,...,x pn}, X i = {x i1 ,x i2 ,...,x in}, x t1 、xp1 , x i1 is the multimodal data corresponding to the table image 1, x t1 is the text data, x p1 is the coordinate data, x i1 is the image data. The table image 1 can be used as the image data, and OCR recognition is performed on the table image 1 to obtain the text data and the coordinate data for each character in the text data. Then, multiple groups of multimodal data included in the training sample set are characterized as feature vectors to obtain multiple groups of multimodal feature vectors.
[0081] Obtain the initial model for table information extraction, input the training sample set into the initial model for table information extraction to obtain the entity type prediction value and the entity relationship prediction value; based on the entity type prediction value and the entity type label, according to the first cross-entropy loss function, obtain the first loss value. For example, the first cross-entropy loss function can be the following formula (1); based on the entity relationship prediction value and the entity relationship label, according to the second cross-entropy loss function, obtain the second loss value. For example, the second cross-entropy loss function can be the following formula (2).
[0082] In this application, the entity relationship prediction task is more difficult than the entity type prediction task. Therefore, the first loss value and the second loss value can be weighted and summed to control the fitting situation of the two tasks, avoid overfitting in the learning of the entity type prediction task, and accelerate the fitting speed of the entity relationship prediction task. Specifically, the total loss value can be determined according to the following formula (3); finally, based on the total loss value and the preset threshold, the initial model for table information extraction is trained multiple times to obtain the model for table information extraction.
[0083] L E = cross_entropy_loss(Y E , Y Er ) Formula (1);
[0084] Among them, Y E is the set of true entity labels, and Y Er is the set of entity labels predicted by the model.
[0085] L R = cross_entropy_loss(Y R , Y Rr ) Formula (2);
[0086] Among them, Y R is the set of true entity relationship labels, and Y Rr is the set of entity relationship labels predicted by the model.
[0087] L = δ * L E + (1 - δ) * L R Formula (3);
[0088] Among them, δ is a weight factor less than 0.5.
[0089] For example, in some embodiments, the multi-modal feature vector 1 can be input into the initial model for table information extraction to obtain the entity type prediction value 1 and the entity relationship prediction value 1; based on the entity type prediction value 1 and the entity type label 1, according to formula (1), the first loss value 1 is obtained; based on the entity relationship prediction value 1 and the entity relationship label 1, according to formula (2), the second loss value 1 is obtained; if δ is set to 0.2, then according to formula (3), the total loss value 1 can be obtained; finally, according to the total loss value 1 and the preset threshold, the parameters in the initial model for table information extraction are updated to obtain the table information extraction model 1.
[0090] Then, the multi-modal feature vector 2 is input into the table information extraction model 1 to obtain the entity type prediction value 2 and the entity relationship prediction value 2; based on the entity type prediction value 2 and the entity type label 2, according to formula (1), the first loss value 2 is obtained; based on the entity relationship prediction value 2 and the entity relationship label 2, according to formula (2), the second loss value 2 is obtained; according to formula (3), the total loss value 2 can be obtained; finally, according to the total loss value 2 and the preset threshold, the parameters in the table information extraction model 1 are updated to obtain the table information extraction model 2. And so on, the multi-modal feature vector 3 is input into the table information extraction model 2, the multi-modal feature vector 4 is input into the table information extraction model 3,..., until the total loss value is less than the preset threshold, and the final table information extraction model with the best fitting effect is obtained.
[0091] Figure 4 shows a schematic flowchart of a method for extracting table information according to another embodiment provided by the present application. It can be Figure 4 seen that the news recommendation method of this embodiment includes the following steps S401 to step S420:
[0092] Step S401: Construct a training sample set. Among them, the training sample set contains multiple groups of multi-modal data and corresponding sample labels. The multi-modal data includes text data, coordinate data, and image data, and the sample labels include entity type labels and entity relationship labels.
[0093] Step S402: Obtain an initial model for table information extraction. The initial model for table information extraction includes a multi-modal data encoding layer, an entity type decoding layer, an entity relationship decoding layer, and a multi-modal pre-training layer. Among them, the multi-modal pre-training layer is generated based on the layoutXLM model, and the output of the multi-modal pre-training layer is used as the input of the multi-modal data encoding layer. The multi-modal data encoding layer is respectively connected to the entity type decoding layer and the entity relationship decoding layer.
[0094] Step S403: Input the training sample set into the initial model for table information extraction to obtain entity type prediction values and entity relationship prediction values.
[0095] Step S404: Based on the entity type prediction values and entity type labels, obtain the first loss value according to the first cross-entropy loss function.
[0096] Step S405: Based on the entity relationship prediction values and entity relationship labels, obtain the second loss value according to the second cross-entropy loss function.
[0097] Step S406: Perform weighted summation on the first loss value and the second loss value to obtain the total loss value.
[0098] Step S407: According to the total loss value and the preset threshold, train the initial model for table information extraction multiple times to obtain the table information extraction model.
[0099] Step S408: Perform encoding processing on the table image to obtain an image vector.
[0100] Step S409: Perform text recognition processing on the table image to obtain text information and coordinate information. Among them, the coordinate information is the coordinate of each character in the text information.
[0101] Step S410: Perform word segmentation processing on the text information to obtain entity texts and entity coordinates corresponding to multiple target entities. Among them, the entity coordinates are the coordinates of each entity text.
[0102] Step S411: Perform vectorization processing on the entity texts and entity coordinates to obtain entity text vectors and entity coordinate vectors.
[0103] Step S412: Based on the multi-modal pre-training layer, process the image vector, entity text vector, and entity coordinate vector to obtain a multi-modal initial feature vector.
[0104] Step S413: Use the multi-modal initial feature vector as the input of the multi-modal data encoding layer.
[0105] Step S414: Based on the multi-modal data encoding layer, encode the multi-modal initial feature vector to obtain a multi-modal feature vector.
[0106] Step S415: Based on the entity type decoding layer, perform type decoding on the multi-modal feature vector to obtain the entity types corresponding to each target entity. Among them, the entity types include at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair.
[0107] Step S416: Based on the entity relationship decoding layer, perform relationship decoding on the multi-modal feature vectors to obtain the entity relationships of at least one target entity pair formed by multiple target entities. The entity relationships include at least one of the following: horizontal entity name - horizontal entity content, vertical entity name - vertical entity content, horizontal entity name - cross-entity content, and vertical entity content - cross-entity content.
[0108] Step S417: Determine whether there is an entity type input instruction. If so, go to Step S418; if not, go to Step S419.
[0109] Step S418: Determine at least one first entity text corresponding to the entity type input instruction and return it.
[0110] Step S419: Determine whether there is an entity text input instruction. If so, go to Step S420.
[0111] Step S420: Determine the second entity text specified by the entity text input instruction, and based on the entity relationship to which the second entity text belongs, determine the third entity text corresponding to the second entity text and return it.
[0112] Figure 5 FIG. shows a schematic structural diagram of a table information extraction device according to an embodiment provided by the present application. The table information extraction device is deployed with a table information extraction model. For the table information extraction model, see Figure 2 shown; the device 500 includes:
[0113] A preprocessing unit 501, configured to preprocess the table image to obtain an image vector, entity text vectors and entity coordinate vectors respectively corresponding to multiple target entities;
[0114] A multi-modal data encoding unit 502, configured to fuse and encode the image vector, the entity text vector, and the entity coordinate vector based on the multi-modal data encoding layer to obtain a multi-modal feature vector;
[0115] A type decoding unit 503, configured to perform type decoding on the multi-modal feature vector based on the entity type decoding layer to obtain the entity types corresponding to each of the target entities. The entity types include at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair;
[0116] A relationship decoding unit 504, configured to perform relationship decoding on the multi-modal feature vector based on the entity relationship decoding layer to obtain the entity relationships of at least one target entity pair formed by the multiple target entities.
[0117] In some embodiments of the present application, the above-mentioned device further includes an entity determination unit, which is configured to, in response to an entity type input instruction, determine at least one first entity text corresponding to the entity type input instruction and return it; in response to an entity text input instruction, determine a second entity text specified by the entity text input instruction, and determine and return a third entity text corresponding to the second entity text according to the entity relationship to which the second entity text belongs.
[0118] In some embodiments of the present application, in the above-mentioned device, the preprocessing unit 501 is configured to perform encoding processing on the table image to obtain the image vector; perform character recognition processing on the table image to obtain text information and coordinate information, where the coordinate information is the coordinate of each character in the text information; perform word segmentation processing on the text information to obtain entity texts and entity coordinates respectively corresponding to the multiple target entities, where the entity coordinates are the coordinates of each entity text; perform vectorization processing on the entity text and the entity coordinates to obtain the entity text vector and the entity coordinate vector.
[0119] In some embodiments of the present application, in the above-mentioned device, the entity relationship includes at least one of the following: horizontal entity name - horizontal entity content, vertical entity name - vertical entity content, horizontal entity name - cross-entity content, and vertical entity name - cross-entity content.
[0120] In some embodiments of the present application, in the above-mentioned device, the table information extraction model further includes a multi-modal pre-training layer; the above-mentioned device further includes a multi-modal pre-training unit, which is configured to process the image vector, the entity text vector, and the entity coordinate vector based on the multi-modal pre-training layer to obtain a multi-modal initial feature vector; use the multi-modal initial feature vector as the input of the multi-modal data encoding layer.
[0121] In some embodiments of the present application, in the above-mentioned device, the multi-modal pre-training layer is generated based on the layoutXLM model.
[0122] In some embodiments of the present application, the above device further includes a training unit for constructing a training sample set; wherein, the training sample set includes multiple groups of multimodal data and corresponding sample labels, the multimodal data includes text data, coordinate data, and image data, and the sample labels include entity type labels and entity relationship labels; obtaining an initial model for table information extraction; inputting the training sample set into the initial model for table information extraction to obtain entity type prediction values and entity relationship prediction values; based on the entity type prediction values and the entity type labels, according to the first cross-entropy loss function, obtaining a first loss value; based on the entity relationship prediction values and the entity relationship labels, according to the second cross-entropy loss function, obtaining a second loss value; performing weighted summation on the first loss value and the second loss value to obtain a total loss value; according to the total loss value and a preset threshold, training the initial model for table information extraction multiple times to obtain the table information extraction model.
[0123] It should be noted that the above-mentioned extraction device for table information can implement the foregoing extraction method for table information one by one, which will not be elaborated here.
[0124] Figure 6 shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 6 shown, at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0125] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0126] A memory for storing programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0127] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a news recommendation device at the logical level. The processor executes the program stored in the memory and is specifically used to execute the foregoing method.
[0128] The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0129] The electronic device can execute the table information extraction method provided in multiple embodiments of the present application and be implemented as a table information extraction device in Figure 5 the functions of the illustrated embodiments, which will not be elaborated herein in the embodiments of the present application.
[0130] The embodiments of the present application also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute the table information extraction method provided in multiple embodiments of the present application.
[0131] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0132] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0133] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0135] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0136] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0137] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0138] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity or device including the element.
[0139] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0140] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for extracting tabular information, characterized in that, The method is implemented by a table information extraction model, which includes a multi-modal data encoding layer, an entity type decoding layer, and an entity relationship decoding layer. Among them, the multi-modal data encoding layer is respectively connected to the entity type decoding layer and the entity relationship decoding layer; The method includes: Preprocessing the table image to obtain an image vector, entity text vectors corresponding to multiple target entities respectively, and entity coordinate vectors; Based on the multi-modal data encoding layer, fusing and encoding the image vector, the entity text vector, and the entity coordinate vector to obtain a multi-modal feature vector; Based on the entity type decoding layer, performing type decoding on the multi-modal feature vector to obtain the entity types corresponding to the target entities, where the entity types include at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair; Based on the entity relationship decoding layer, performing relationship decoding on the multi-modal feature vector to obtain the entity relationships of at least one target entity pair formed by the multiple target entities; The table information extraction model is trained according to the following method: Construct a training sample set; among them, the training sample set contains multiple groups of multi-modal data and corresponding sample labels, the multi-modal data includes text data, coordinate data, and image data, and the sample labels include entity type labels and entity relationship labels; Obtain an initial table information extraction model; Input the training sample set into the initial table information extraction model to obtain entity type prediction values and entity relationship prediction values; Based on the entity type prediction values and the entity type labels, according to the first cross-entropy loss function, obtain a first loss value; based on the entity relationship prediction values and the entity relationship labels, according to the second cross-entropy loss function, obtain a second loss value; Perform weighted summation on the first loss value and the second loss value to obtain a total loss value; According to the total loss value and a preset threshold, train the initial table information extraction model multiple times to obtain the table information extraction model.
2. The method according to claim 1, characterized in that, The method further includes: Responding to an entity type input instruction, determining at least one first entity text corresponding to the entity type input instruction and returning it; Responding to an entity text input instruction, determining a second entity text specified by the entity text input instruction, and according to the entity relationship to which the second entity text belongs, determining a third entity text corresponding to the second entity text and returning it.
3. The method according to claim 1, wherein The preprocessing the table image to obtain an image vector, entity text vectors corresponding to multiple target entities respectively, and entity coordinate vectors includes: Encoding the table image to obtain the image vector; Performing optical character recognition on the table image to obtain text information and coordinate information, where the coordinate information is the coordinate of each character in the text information; Performing word segmentation on the text information to obtain entity texts and entity coordinates corresponding to the multiple target entities respectively, where the entity coordinates are the coordinates of each entity text; Vectorize the entity text and the entity coordinates to obtain the entity text vector and the entity coordinate vector.
4. The method according to claim 1, wherein The entity relationship includes at least one of the following: horizontal entity name - horizontal entity content, vertical entity name - vertical entity content, horizontal entity name - cross-entity content, and vertical entity name - cross-entity content.
5. The method according to claim 1, wherein The table information extraction model further includes a multi-modal pre-training layer; The method further includes: Based on the multi-modal pre-training layer, process the image vector, the entity text vector, and the entity coordinate vector to obtain a multi-modal initial feature vector; Use the multi-modal initial feature vector as the input of the multi-modal data encoding layer.
6. The method according to claim 5, characterized in that, The multi-modal pre-training layer is generated based on the layoutXLM model.
7. An extraction device for tabular information, characterized in that, The device is deployed with a table information extraction model, which includes a multi-modal data encoding layer, an entity type decoding layer, and an entity relationship decoding layer. Among them, the multi-modal data encoding layer is respectively connected to the entity type decoding layer and the entity relationship decoding layer; The device includes: A preprocessing unit for preprocessing a table image to obtain an image vector, an entity text vector and an entity coordinate vector corresponding to each of a plurality of target entities; A multi-modal data encoding unit for fusing and encoding the image vector, the entity text vector, and the entity coordinate vector based on the multi-modal data encoding layer to obtain a multi-modal feature vector; A type decoding unit for performing type decoding on the multi-modal feature vector based on the entity type decoding layer to obtain the entity type corresponding to each of the target entities. The entity type includes at least one of the following: file name, title, horizontal entity name, horizontal entity content, vertical entity name, vertical entity content, cross-entity content, and name-content pair; A relationship decoding unit for performing relationship decoding on the multi-modal feature vector based on the entity relationship decoding layer to obtain the entity relationship of at least one target entity pair formed by the plurality of target entities; A training unit for constructing a training sample set. The training sample set contains multiple groups of multi-modal data and corresponding sample labels. The multi-modal data includes text data, coordinate data, and image data. The sample labels include entity type labels and entity relationship labels; obtain an initial table information extraction model; input the training sample set into the initial table information extraction model to obtain entity type prediction values and entity relationship prediction values; based on the entity type prediction values and the entity type labels, according to the first cross-entropy loss function, obtain a first loss value; based on the entity relationship prediction values and the entity relationship labels, according to the second cross-entropy loss function, obtain a second loss value; perform weighted summation on the first loss value and the second loss value to obtain a total loss value; according to the total loss value and a preset threshold, train the initial table information extraction model multiple times to obtain the table information extraction model.
8. An electronic device, including: A processor; And A memory arranged to store computer-executable instructions which, when executed, cause the processor to perform the steps of the method for extracting tabular information according to any one of claims 1 to 6.
9. A computer-readable storage medium storing one or more programs which, when executed by an electronic device including a plurality of application programs, cause the electronic device to perform the steps of the method for extracting tabular information according to any one of claims 1 to 6.
Citation Information
Patent Citations
Entity relationship joint extraction method and device, storage medium and terminal
CN114840680A
Data processing method and device, electronic equipment and storage medium
CN115115913A