Data processing method, table processing method, device, equipment and storage medium
Patent Information
- Application Number
- CN202211518049.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-11-30
AI Technical Summary
表格数量庞大且结构多样,人工从这些表格中获取所需信息需要巨大的成本
[0025] The technical solutions of this disclosure can improve the diversity of training data for table pre-trained models, thereby improving the generalization ability of table pre-trained models.
Smart Images

Figure CN115860136B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning and NLP (Natural Language Processing). Background Technology
[0002] The web (World Wide Web) and industry documents contain a wide variety of tabular data. These tables are numerous and structurally diverse, making it extremely costly to manually extract the necessary information from them. Therefore, the Natural Language Processing (NLP) field has defined numerous table understanding tasks, such as cell classification, classifying relationships between cells, and question-and-answer based cell location. To accomplish these tasks, pre-trained table models can be used to process the tables. These models can comprehensively utilize the table's structural information and the textual information within the table to obtain corresponding semantic representations, thereby effectively improving the processing performance of downstream table understanding tasks.
[0003] Since pre-trained table models cannot acquire patterns and knowledge not learned during the pre-training phase, the training dataset used for pre-training is crucial to the generalization ability of these models. In practical applications, the training data for pre-trained table models is typically real, simple tables. Summary of the Invention
[0004] This disclosure provides a data processing method, a table processing method, an apparatus, a device, and a storage medium.
[0005] According to one aspect of this disclosure, a data processing method is provided, comprising:
[0006] Retrieve data on multiple entity relationships;
[0007] Based on the aforementioned entity relationship data, multiple first tables are constructed;
[0008] Based on the plurality of first tables, a training dataset is obtained; wherein, the training dataset is used to train the pre-trained model of the obtained tables.
[0009] According to another aspect of this disclosure, a table processing method is provided, comprising:
[0010] The target table is processed using a pre-trained table model to obtain a semantic representation of the target table; wherein the pre-trained table model is obtained based on the training dataset of any embodiment of this disclosure;
[0011] The table comprehension task is performed based on semantic representation, and the task processing results are obtained.
[0012] According to one aspect of this disclosure, a data processing apparatus is provided, comprising:
[0013] The data acquisition module is used to acquire data on relationships between multiple entities.
[0014] The table construction module is used to construct multiple first tables based on the multiple entity relationship data;
[0015] The dataset determination module is used to obtain a training dataset based on the plurality of first tables; wherein the training dataset is used to train a pre-trained model of the obtained tables.
[0016] According to another aspect of this disclosure, a form processing apparatus is provided, comprising:
[0017] A semantic representation output module is used to process a target table using a table pre-training model to obtain a semantic representation of the target table; wherein the table pre-training model is obtained based on the training dataset in any embodiment of this disclosure;
[0018] The task processing module is used to perform a table understanding task based on the semantic representation and obtain the task processing result.
[0019] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0020] At least one processor; and
[0021] The memory is communicatively connected to the at least one processor; wherein,
[0022] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0023] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0024] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0025] The technical solutions of this disclosure can improve the diversity of training data for table pre-trained models, thereby improving the generalization ability of table pre-trained models.
[0026] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0027] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0028] Figure 1 This is a schematic flowchart of a data processing method provided in an embodiment of this disclosure;
[0029] Figure 2 This is a schematic diagram showing the quantile distribution of the number of cells in various tables in this embodiment of the disclosure;
[0030] Figure 3 This is a schematic diagram illustrating an application example of an embodiment of this disclosure.
[0031] Figure 4 This is a schematic flowchart of a table processing method provided in another embodiment of this disclosure;
[0032] Figure 5 This is a schematic block diagram of a data processing apparatus provided in an embodiment of the present disclosure;
[0033] Figure 6 This is a schematic block diagram of a data processing apparatus provided in another embodiment of the present disclosure;
[0034] Figure 7 This is a schematic block diagram of a data processing apparatus provided in yet another embodiment of this disclosure;
[0035] Figure 8 This is a schematic block diagram of a table processing apparatus provided in an embodiment of the present disclosure;
[0036] Figure 9 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0037] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0038] Figure 1 This is a schematic flowchart of a data processing method provided in one embodiment of this disclosure. Figure 1 As shown, the method may include:
[0039] Step S110: Obtain multiple entity relationship data;
[0040] Step S120: Construct multiple first tables based on multiple entity relationship data;
[0041] Step S130: Based on multiple first tables, a training dataset is obtained, which is used to train the pre-trained model of the obtained tables.
[0042] In this embodiment of the disclosure, entity relationship data is used to represent the relationship between two entities. Exemplarily, entity relationship data is triple data, and the structure of this triple data can be (Subject, Predicate, Object) or SPO data. For example, if two entities include "Li Bai" and "Quiet Night Thoughts," and the relationship between these two entities is "Li Bai's representative work is Quiet Night Thoughts," then this relationship can be represented using SPO data (Li Bai, representative work, Quiet Night Thoughts).
[0043] In some examples, the relationship between two entities is often an attribute relationship between the entities. For instance, the structure of the SPO data mentioned above can often be understood as (entity, attribute, attribute value). For example, the SPO data (Li Bai, masterpiece, Quiet Night Thoughts) can also be understood as: the attribute value of the entity "Li Bai" for the attribute "masterpiece" is "Quiet Night Thoughts". Therefore, in some examples, entity relationship data can also be called entity attribute data.
[0044] For example, in this embodiment of the disclosure, the first table is a table constructed in response to the need for table pre-training, that is, the first table is not a real table. Optionally, the first table can be constructed by filling the table template with the data of each entity relationship according to a pre-configured table template.
[0045] For example, a table template can be pre-defined with cell ranges for filling in the subject information, predicate information, and object information of the SPO data. In practical applications, the subject, predicate, and object information of each SPO data point are filled into the corresponding cell ranges to obtain the first table.
[0046] Optionally, the multiple first tables may include tables of different types, or multiple tables of the same type. Here, different types of tables may refer to tables that are populated using different templates. In specific implementations, at least a portion of the data from multiple entity relationship data can be populated into different table templates to obtain multiple tables of different types. Alternatively, at least a portion of the data from multiple entity relationship data can be grouped, and each group of data can be populated into the same table template to obtain multiple tables of the same type, with each table corresponding to a group of data.
[0047] For example, in step S130 above, multiple first tables can be aggregated to obtain a training dataset. Alternatively, multiple first tables can be aggregated with multiple real tables to obtain a training dataset. Based on this, the training dataset in this embodiment of the disclosure includes at least multiple first tables.
[0048] The training dataset in this embodiment is used to train a pre-trained table model. Exemplarily, this pre-trained table model can be used to obtain a corresponding semantic representation for an input table. This semantic representation is used to complete a table understanding task. The table understanding task can include self-supervised and supervised tasks. Supervised tasks include, for example, question-based cell retrieval, question parsing, table retrieval, cell type classification, cell pair relationship classification, and table type classification. Self-supervised tasks include, for example, cell-level cloze, cell value recovery, and corrupt cell detection.
[0049] Optionally, in this embodiment of the disclosure, the constructed multiple tables can be subjected to different processing operations corresponding to different tasks, thereby obtaining corresponding training data subsets, so that the table pre-trained model can output accurate semantic representations for the patterns and related knowledge of different tasks.
[0050] As can be seen, the data processing method provided in this disclosure uses entity relationship data to construct a table, and obtains a training dataset for a table pre-training model based on the constructed table. Compared with using only real tables to construct the training dataset for a table pre-training model, the method in this disclosure can improve the diversity of training data for the table pre-training model, thereby improving the generalization ability of the table pre-training model.
[0051] In some embodiments of this disclosure, a method for obtaining the aforementioned entity relationship data is also provided. Optionally, in step S110 above, obtaining multiple entity relationship data includes:
[0052] Extract multiple key-value pairs from the entity description information;
[0053] Based on the entity described by the entity description information and each key-value pair in the multiple key-value pair data, the entity relationship data corresponding to each key-value pair data is obtained.
[0054] For example, entity description information can include any document, page, etc., used to describe the entity. For instance, entity description information can include the entity's online encyclopedia page, which contains a large amount of KV (Key-Value) data that can be converted into SPO data.
[0055] For example, the entity description information for an entity "knowledge graph" contains multiple key-value pairs, as follows:
[0056] Alternative name: Scientific Knowledge Graph;
[0057] Foreign name: Knowledge Graph;
[0058] Applications: Theory and Methodology and Bibliometric Citation Analysis.
[0059] Using the entity "knowledge graph" and the key-value data above, multiple entity relationship data can be constructed, as follows:
[0060] (Knowledge graph, also known as scientific knowledge graph);
[0061] (Knowledge Graph)
[0062] (Knowledge graphs, applications, theories and methods, and bibliometric citation analysis).
[0063] By adopting the above-mentioned method of acquiring entity relationship data, high-quality entity relationship data can be obtained, which is conducive to constructing a first table that conforms to the real situation. This improves the diversity of the training dataset for the table pre-training model and enhances the quality of the training dataset.
[0064] It is understood that in some embodiments, key-value (KV) data used to construct entity relationship data can also be extracted in combination with other methods. For example, some industry websites contain comprehensive industry KV pages, from which KV data can be extracted. Optionally, both the aforementioned KV pages and online encyclopedia pages can be obtained through web scraping.
[0065] Optionally, in some embodiments, step S120 above, which involves constructing multiple first tables based on multiple entity relationship data, may include:
[0066] Based on multiple entity relation data, determine M entity relation data sets corresponding to M subject information respectively, where M is an integer greater than or equal to 2;
[0067] Based on a set of M entity relationship data, construct multiple first tables.
[0068] Specifically, multiple entity relationship data obtained from various channels can be aggregated, and then divided according to the subject information in each entity relationship data set to obtain multiple entity relationship data sets corresponding to multiple subject information sets. Each entity relationship data set includes multiple entity relationship data sets with the same subject information, and the entity relationship data in different entity relationship data sets correspond to different subject information sets.
[0069] According to the above implementation method, aggregating entity relationship data corresponding to the same subject information together is beneficial to constructing a first table that is more in line with the real situation based on the correlation between data during the table construction process, thereby improving the quality of the training dataset.
[0070] Below are several exemplary ways to construct the first table based on a collection of entity relationship data.
[0071] Example 1: In this example, multiple first tables can include relational tables, i.e., SQL (Structured Query Language) tables. A relational table can contain information about multiple entities, and multiple entities have the same set of attributes. The entity rows and attribute columns intersect, or the entity columns and attribute rows intersect to fill the attribute values of the entities.
[0072] Specifically, in this example, based on M entity relationship data sets, multiple first tables are constructed, including:
[0073] From M entity relation data sets, identify K entity relation data sets that have at least N identical predicate information, where N is an integer greater than or equal to 1, K is an integer greater than or equal to 2 and K is less than or equal to M;
[0074] Based on the K subject information and N identical predicate information corresponding to the K entity relationship data sets, populate the table header information in the table template;
[0075] Using the object information from the K entity relationship datasets, fill the table value area in the table template to obtain a relationship table with K subject information.
[0076] The subject information is the S data in the SPO data, the predicate information is the P data in the SPO data, and the object information is the O data in the SPO data.
[0077] For example, let's set N=5, and assume that the M entity relation data sets include an entity relation data set with TV program A as the subject, an entity relation data set with TV program B as the subject, and an entity relation data set with serialized novel C as the subject. The predicate information of multiple entity relation data in the TV program A entity relation data set includes the last episode time, broadcast status, type, online streaming platform, premiere time, and director; the predicate information of multiple entity relation data in the TV program B entity relation data set includes the last episode time, broadcast status, type, online streaming platform, premiere time, and type; and the predicate information of multiple entity relation data in the serialized novel C entity relation data set includes the last episode time, author, serialization platform, and work type. It can be seen that the entity relation data sets of TV program A and TV program B have 5 identical predicate information (last episode time, broadcast status, type, online streaming platform, and premiere time). Therefore, we can fill the table template based on the entity relation data sets of TV program A and TV program B, resulting in the relation table shown in Table 1.
[0078]
[0079] In this example, SPO datasets that share at least N identical P values are aggregated together, and then a relational table is constructed based on these aggregated SPO datasets. This allows for the construction of a relational table that reflects real-world scenarios and possesses a certain degree of complexity. A pre-trained table model trained on this relational table can improve the ability of pre-trained table models to handle complex relational tables.
[0080] Example 2: In this example, multiple first tables can include horizontally stacked tables. A horizontally stacked table is a table whose left and right parts have the same structure, and it can contain information about multiple entities, with multiple entities having the same set of attributes.
[0081] Specifically, in this example, based on M entity relationship data sets, multiple first tables are constructed, including:
[0082] From M entity relation data sets, identify K entity relation data sets that have at least N identical predicate information, where N is an integer greater than or equal to 1, K is an integer greater than or equal to 2 and K is less than or equal to M;
[0083] Divide the K entity relation data sets into L groups of entity relation data sets; where L is an integer greater than or equal to 2;
[0084] Based on each entity relation data set in the L sets of entity relation data sets, a relation table of multiple subject information corresponding to each entity relation data set is obtained;
[0085] A horizontally stacked table is obtained by horizontally combining the relation tables corresponding to each set of entity relation data.
[0086] This example is similar to Example 1, aggregating SPO (Entity-Relationship) datasets that share at least N Ps. The difference is that this example divides these SPO datasets into L groups, each group containing at least one SPO dataset. For each group of SPO datasets, relational tables can be constructed as in Example 1, and then the L relational tables are horizontally combined to obtain a horizontally stacked table. The value of L can be preset or randomly selected from multiple values; for example, L can be randomly selected as 2 or 3, resulting in a horizontally stacked table of 2 or 3 relational tables.
[0087] For example, if N=2 and L=2, the constructed horizontal stacked table can be as shown in Table 2:
[0088]
[0089] In this example, the SPO dataset sharing N identical P values is grouped, and then relational tables are constructed separately for each group. These groups are then horizontally combined to obtain a horizontally stacked table. This method allows for the construction of relational tables that reflect real-world scenarios and possess a certain level of complexity. A pre-trained table model trained using this method can improve the ability of a pre-trained table model to handle complex horizontally stacked tables.
[0090] Example 3: In this example, multiple first tables may include entity tables. An entity table describes the attribute information of an entity, and its header and value usually appear in pairs.
[0091] Specifically, in this example, multiple first tables are constructed based on M entity relationship data sets, including: an entity table corresponding to the i-th subject information is constructed based on the entity relationship data set corresponding to the i-th subject information among the M subject information, where i is a positive integer less than or equal to M.
[0092] In practical applications, for a certain entity relationship data set, the predicate information and object information of each entity relationship data can be paired up and filled into a table template to construct an entity table with subject information corresponding to that entity relationship data set.
[0093] For example, for a serialized novel C, its corresponding entity relation data set includes multiple attributes of the serialized novel C, and the constructed entity table can be shown in Table 3:
[0094]
[0095] Table 3
[0096] For example, the first Y entity relationship data derived from the entity description information in the entity relationship data set can be filled into the first Y rows of the entity table so that the key information is located in the first Y rows of the entity table, where Y is a positive integer, such as 2 or 3. Other entity relationship data can be filled in randomly.
[0097] For example, as shown in Table 3, when filling the predicate and object information in each entity relation data into the table template in pairs, the filling can be based on header alignment and header misalignment.
[0098] In this table header alignment, the P data in the SPO data is filled into columns with odd-numbered serial numbers, and the O data can be filled into 2n-1 columns to the right of the P data, where n is an integer, meaning the number of columns occupied by the O data is odd. For example, in rows 1 to 3 of Table 3, the P data is filled into column 1, and the O data occupies 3 columns; in rows 4 and 5 of Table 3, the P data is filled into columns 1, 3, and 5, and the O data occupies 1 or 3 columns.
[0099] The misalignment of headers refers to the fact that the columns containing P and O data in the SPO data are not limited in terms of the number of columns. For example, in row 5 of Table 3, P data is filled in columns 1 and 4, and O data occupies 2 columns.
[0100] In practical applications, a probability p can be set. Based on probability p, entity relationship data is filled into the table template with headers aligned. Based on probability (1-p), entity relationship data is filled into the table template with headers misaligned. Here, p is a positive number less than 1, for example, p = 0.8.
[0101] In this example, an entity table is constructed based on the entity relationship data set corresponding to the same entity. This allows for the creation of entity tables that accurately reflect real-world scenarios. The pre-trained table model trained using this entity table can improve its ability to handle complex entity tables.
[0102] Example 4: In this example, multiple first tables can include vertically stacked tables. A vertically stacked table is a table whose upper and lower parts have the same structure, and it can contain information about multiple entities.
[0103] Specifically, in this example, multiple first tables are constructed based on M entity relationship data sets, including: vertically combining the entity tables corresponding to each subject information in the M subject information to obtain a vertically stacked table.
[0104] In other words, a vertically stacked table can be obtained by vertically combining multiple entity tables. Optionally, in each entity table of the vertically stacked table, the entity name can be set as a whole row of merged cells and placed in the first row of the entity table.
[0105] For example, a vertically stacked table can be as shown in Table 4 below, which contains an entity table for TV program A and an entity table for serialized novel C:
[0106]
[0107] Table 4
[0108] In this example, a vertically stacked table is constructed based on entity tables of different entities. This allows for the creation of a realistically simulated vertically stacked table with a certain degree of complexity. The pre-trained table model trained on this constructed vertically stacked table can improve the ability of pre-trained table models to handle complex vertically stacked tables.
[0109] Examples 1 to 4 above provide multiple implementation methods for constructing a first table based on an entity relationship data set. It can be understood that the above examples can be implemented individually or in combination, so that the training dataset of the table pre-training model contains diverse tables.
[0110] Optionally, a hierarchical table can be constructed based on the hierarchical relationship between predicate information in the entity relationship data to further enhance the table diversity in the training dataset. For example, step S120 above, constructing multiple first tables based on multiple entity relationship data, may include: determining the hierarchical relationship between multiple predicate information; and filling the multiple entity relationship data corresponding to the multiple predicate information into the table template based on the hierarchical relationship between the multiple predicate information to obtain the hierarchical table.
[0111] This involves the hierarchical relationship between multiple predicate information, that is, the hierarchical relationship between multiple attributes. For example, the attribute "film crew members" in a movie generally includes sub-attributes such as "director of photography" and "screenwriter".
[0112] For example, the hierarchical relationship of multiple predicate information can be determined by crawling web pages. For instance, some online encyclopedia pages contain multi-level K data and corresponding V data, and the hierarchical relationship of multiple predicate information can be obtained based on the multi-level K data in such online encyclopedia pages.
[0113] In practical applications, based on the hierarchical relationship between multiple predicate information, multiple predicate information can be used as table headers, and the object information from the corresponding multiple entity relationship data can be filled into the table template as table values. For example, a hierarchical table can be shown in Table 5 below:
[0114]
[0115] Table 5
[0116] In this example, a hierarchical table is constructed using the hierarchical relationships between multiple predicate information, thereby further increasing the complexity of the first table. The training dataset is obtained using this hierarchical table, and the pre-trained table model trained based on this dataset can handle complex hierarchical tables, thus further improving the generalization ability of the pre-trained table model.
[0117] As can be seen, in the process of constructing the first table, at least one table hyperparameter is needed to specify the size of the table. This table hyperparameter can be a quantitative parameter, such as the number of entities in the relation table, the number of shared attributes in the entity relation data used to construct the relation table, the number of horizontally combined relation tables in a horizontally stacked table, and the number of vertically combined entity tables in a vertically stacked table. In practical applications, the table hyperparameter can be determined through hyperparameter iteration.
[0118] Specifically, before constructing multiple first tables based on multiple entity relationship data, the above method may also include:
[0119] Based on multiple entity relationship data, X iterations are performed to determine the table hyperparameters, which include at least one quantity parameter for constructing multiple first tables; where X is an integer greater than or equal to 2.
[0120] Among them, the j-th iteration operation in the X iteration operations includes:
[0121] Based on multiple entity relationship data and the table hyperparameters updated in the (j-1)th time, construct multiple third tables;
[0122] Based on the size information of each of the multiple third tables, determine the distribution of size information of the multiple third tables;
[0123] Based on the distribution of scale information, the table hyperparameters are updated for the jth time; where j is an integer greater than or equal to 1.
[0124] It should be noted that for the case where j=1, the table hyperparameters updated in the (j-1)th time (i.e. the 0th time) can be the pre-set initial values of the table hyperparameters.
[0125] Optionally, the number of iterations X can be a predetermined value or a value determined by the iteration process. When X is a predetermined value, when the number of iterations reaches X, i.e., j = X, the last updated table hyperparameters are the table hyperparameters used to construct the first table. When X is determined by the iteration process, when the distribution of scale information meets preset conditions (e.g., a distribution of scale information close to that of a real table), the last updated table hyperparameters are the table hyperparameters used to construct the first table.
[0126] In practical applications, the update strategy for table hyperparameters can be determined by comparing the size distribution of multiple constructed third tables with that of the real table.
[0127] For example, table size information can refer to the number of cells in the table. Figure 2 The diagram illustrates the quantile distribution of cell counts for various types of tables. Curve 201 represents the cell count distribution of a real table from an online encyclopedia page, curve 202 represents the cell count distribution of a real table from an industry standard document, curve 203 represents the cell count distribution of a constructed relational table, and curve 204 represents the cell count distribution of a constructed entity table. It can be seen that compared to curves 201 and 202 (the two real table distribution curves), curve 203 has a larger quantile, indicating that the constructed relational table has too many cells, which does not conform to the actual data distribution. Curve 204 has a smaller quantile, indicating that the number of cells is too few, which also does not conform to the actual data distribution. Based on this conclusion, it is necessary to adjust the table hyperparameters and perform secondary data reconstruction.
[0128] The above implementation method ensures that the size distribution of the constructed first table conforms to the real distribution through a distribution control mechanism, thereby improving the realism of the first table and correspondingly improving the quality of the training dataset and the accuracy of the semantic representation output by the table pre-trained model.
[0129] Optionally, during the above hyperparameter iteration process, the table size information includes at least one of the following: number of cells, number of rows, number of columns, number of character elements in cells, and number of character elements in the table.
[0130] By setting one or more scale information to control the iteration of table hyperparameters, it is possible to further ensure that the scale distribution of the constructed first table conforms to the true distribution.
[0131] The above examples illustrate how the table is constructed based on entity relationship data in this disclosure embodiment. In practical applications, the training dataset is not limited to containing the constructed first table; it may also contain a real second table.
[0132] For example, the above data processing method may further include: obtaining standard documents for the target industry; and extracting multiple second tables from the standard documents. Correspondingly, obtaining a training dataset based on the multiple first tables includes: obtaining a training dataset based on the multiple first tables and the multiple second tables.
[0133] It is understood that the second table in this embodiment is a real table, which can be extracted from standard documents in a specific industry. For example, the crawled standard document can be parsed into structured data and then the structured second table can be extracted.
[0134] In practical applications, there can be one or more target industries. For example, target industries may include at least one of the following: power and energy, finance, healthcare, and JG industries. When there are multiple target industries, corresponding standard documents can be crawled from web pages for each target industry, and then real tables can be extracted from each standard document.
[0135] Since standard documents represent unified technical requirements within a specific scope, are geared towards professionals, possess strong industry semantics, and contain numerous tables with rich and complex structures, they are crucial for improving the pre-trained model's ability to understand industry-specific and complex tables. Secondly, the presence of tabular data in standard documents is highly likely to be accompanied by explanatory / descriptive text; therefore, introducing standard documents as a data source can enhance the pre-trained model's ability to jointly model tables and documents.
[0136] It is understandable that, in practical applications, the documents used to extract real tables are not limited to standard documents. For example, tables can also be extracted from various types of documents involved in the industry, such as manuals, ledgers, contracts, announcements, and plans.
[0137] For example, the priority of each document type can be predetermined, and multiple second tables can be sampled from tables in documents of different priorities using different sampling ratios. For instance, the higher the priority, the higher the sampling ratio. This makes the tables in the training dataset more diverse.
[0138] For example, the priority of standard documents is set to the highest. The priority settings for other types of documents can be determined with reference to the following content.
[0139] 1. Instruction Manual: The instruction manual is open to the public. Through manual analysis of some data, it was found that the table structure in the instruction manual is relatively simple and has strong universality. Therefore, it can be assumed that through data learning in the general domain, the tables in the instruction manual type of document can be understood. Therefore, it is set as a low priority.
[0140] 2. Ledger: The ledger refers to the detailed record of the operation process. It is intended for professionals. The document is full of tables, with some content to be filled in and some left blank. It is a very typical discrete table. In addition, the ledger table does not involve text interaction and is set as medium priority.
[0141] 3. Contracts: Contracts refer to transaction agreements, which are geared towards the general public (but lean towards professionals). Depending on the subject matter, the table may include knowledge from other industries. The table structure is relatively simple and is set to medium priority.
[0142] 4. Announcements: Similar to the instruction manual, set to low priority.
[0143] 6. Plan: A plan refers to a specific program, similar to a standard, and is set as a high priority.
[0144] According to the above exemplary method, the training dataset may include a first table constructed based on web encyclopedia pages in a general domain, or a second table extracted based on standard documents of the target industry. As can be seen, the data sources in the training dataset are rich, which can ensure sufficient training volume for the table pre-trained model and greatly improve the generalization ability of the table pre-trained model.
[0145] As seen in the examples above, the training dataset can include multiple constructed tables, as well as real tables obtained through other means. The pre-trained model using this training dataset can output accurate semantic representations for the processing patterns of self-supervised tasks.
[0146] Optionally, in this embodiment of the disclosure, a method for labeling data in the training dataset can also be provided for supervised tasks.
[0147] For example, for a real second table, in some embodiments, the above data processing method may further include:
[0148] For each second table in the training dataset, retrieve the corresponding reference text in the standard documents;
[0149] The supervision label information corresponding to the second table is determined based on the cited text;
[0150] Based on the second table and the corresponding supervision label information, a first subset of training data for the supervised task is obtained.
[0151] The cited text can refer to the text in a standard document that introduces a table by referencing its table number, such as the statement "Table 1 is a description of the components in the lighting fixture". This cited text allows identification of the second table's name. Using the second table's name as a supervisory label can improve the semantic representation ability of the pre-trained table model in table retrieval tasks.
[0152] In practical applications, for the structured second table extracted from the standard document, information such as the table name, cell type, and cell relationships can be obtained by extracting the table's structural information.
[0153] For example, the table name can be obtained through the labels in the second table. Using the table name as the supervision label information of the second table can improve the semantic representation ability of the table pre-trained model in the table localization task.
[0154] For example, the header of the second table can be identified by the markings in the second table, thus determining whether each cell in the table is a header or a value, thereby identifying the type of each cell. Using the cell type as the supervision label information of the second table can improve the semantic representation ability of the table pre-trained model in the cell type classification task.
[0155] For example, the headers in the second table can be identified by the markings in the second table, thereby determining the pairwise relationship between the headers and values in the second table. Furthermore, this can be extended into a pair of cell relationship labels through rules, which can improve the semantic representation ability of the table pre-trained model in the cell relationship classification task.
[0156] Alternatively, the table headers and values can be identified in the second table through table CVT extraction, thereby constructing corresponding supervision label information for different table understanding tasks.
[0157] For example, in some embodiments, for the constructed second table, the above data processing method may further include:
[0158] For each first table in the training dataset, obtain the corresponding entity relationship data;
[0159] Based on entity relationship data, determine the supervision label information corresponding to the first table;
[0160] Based on the first table and the supervision label information, a second subset of training data for the supervised task is obtained.
[0161] Since the first table is constructed based on entity relationship data, the corresponding entity relationship data can be obtained for the first table. The structured information of SPO in the entity relationship data can be used to construct supervision label information, thereby improving the accuracy of supervision label information and ensuring sufficient training for supervised tasks.
[0162] Optionally, for supervised table comprehension tasks involving text interaction, determining the supervision label information corresponding to the first table based on entity relationship data may include:
[0163] Based on entity relationship data and preset question-and-answer templates, construct the query statements and answers corresponding to the first table;
[0164] Use the answers as supervisory label information for the first table and the query statement.
[0165] For example, the question-and-answer template may include the following templates:
[0166] (1) SP-O: This means that the query statement contains S data and P data, and the answer contains O data. For example, for entity relation data (Li Bai, representative works, Quiet Night Thoughts), the query statement can be constructed as "What is Li Bai's representative work?" and the answer is "Quiet Night Thoughts".
[0167] (2) SO-P: This means that the query statement contains S data and O data, and the answer contains P data. For example, for entity relationship data (Li Bai, representative works, Quiet Night Thoughts), the query statement can be constructed as "What is the relationship between Li Bai and Quiet Night Thoughts", and the answer is "Li Bai's representative work is Quiet Night Thoughts".
[0168] (3) OP-S: This means that the query statement contains O data and P data, and the answer contains S data. For example, for entity relation data (Li Bai, representative works, Quiet Night Thoughts), the query statement can be constructed as "Whose representative work is Quiet Night Thoughts", and the answer is "Li Bai".
[0169] (4) S: That is, the question is about the S data, such as "What is Li Bai's representative work?" or "What are Li Bai's hobbies?"
[0170] (5) P: That is, the question is about the data P, such as "What are the representative works of the poets of the Tang Dynasty respectively".
[0171] In addition, some query statements related to Boolean operations, numerical comparisons, numerical calculations, and date comparisons in the cells of the table can be constructed, and the corresponding answers can be marked in the first table to improve the diversity of the supervision label information.
[0172] Based on the above question-and-answer template, various forms of query statements and answers can be constructed for each first table, thereby improving the semantic representation ability of the table pre-trained model in question-and-answer related table understanding tasks (such as cell location, question parsing, and table location).
[0173] Optionally, for supervised table comprehension tasks without text interaction, based on entity relationship data, the supervision label information corresponding to the first table is determined, including:
[0174] Based on the information types of multiple pieces of information in the entity relationship data, determine the cell types of multiple cells in the first table and / or the relationships between multiple cells;
[0175] Use the cell types of multiple cells and / or the relationships between multiple cells as the supervisory label information for the first table.
[0176] For example, for S and P data in entity relationship data, the corresponding cell type can be determined as table header. For O data in entity relationship data, the corresponding cell type can be determined as table value. For S and P data, or S and O data, or P and O data in the same entity relationship data, their corresponding pair of cells can be marked as related. For information in different entity relationship data, their corresponding pair of cells can be marked as unrelated.
[0177] As can be seen, based on the above method, type labels and relationship labels can be constructed for multiple cells or pairs of cells in each first table, thereby improving the semantic representation ability of the table pre-trained model in classification-related tasks (such as cell type classification, cell relationship classification, and table type classification).
[0178] To facilitate understanding of the above data processing methods, Figure 3 A schematic diagram of an application example is shown. For example... Figure 3 As shown, in the application example, the data processing method includes the following steps:
[0179] 1. Define the domain: The domain that the data construction needs to cover, including general domains and multiple target industries.
[0180] 2. Define the data source: Determine the quality and acquisition difficulty of the data in each field, and comprehensively select and construct the data source. For example, select general-domain web pages and industry standard documents as data sources. General-domain web pages can extract structured data (such as key-value pairs) and semi-structured data, while industry standard documents can extract structured data (such as complete tables) and unstructured data.
[0181] 3. Data Construction: Determine the mining scheme and strategy for each data source.
[0182] 4. Tagging: Depending on the data source type, supervised tagging tasks are performed in different ways.
[0183] 5. Distributed control: For key parameters, the control strategy constructs data that approximates the distribution of real data.
[0184] 6. Output: Output data and label results for table pre-training. The output data includes simple tables and complex tables. Simple tables include relational tables, while complex tables include entity tables, stacked tables, and hierarchical tables, etc.
[0185] Figure 4 A flowchart illustrating a table processing method according to another embodiment of this disclosure is shown. Figure 4 As shown, the method may include:
[0186] Step S410: Process the target table using a table pre-training model to obtain the semantic representation of the target table; wherein, the table pre-training model is obtained based on the training dataset obtained in any embodiment of this disclosure;
[0187] Step S420: Perform a table understanding task based on semantic representation to obtain the task processing result.
[0188] The target table is the table to be processed. Tasks involving understanding the target table can include self-supervised and supervised tasks. Supervised tasks include, for example, question-based cell retrieval, question parsing, table retrieval, cell type classification, cell pair relationship classification, and table type classification. Self-supervised tasks include, for example, cell-level cloze, cell value recovery, and corrupted cell detection.
[0189] Since the semantic representation on which the table understanding task is performed is based on a table pre-trained model, and the table pre-trained model is trained on the training dataset obtained in the foregoing embodiments of this disclosure, the table pre-trained model has strong generalization ability and can accurately output semantic representations for various types of table understanding tasks, thereby improving the processing effect of table understanding tasks.
[0190] According to embodiments of this disclosure, this disclosure also provides a data processing apparatus. Figure 5 A schematic block diagram of a data processing apparatus provided according to an embodiment of the present disclosure is shown. Figure 5 As shown, the data processing apparatus may include:
[0191] Data acquisition module 510 is used to acquire data on multiple entity relationships;
[0192] Table building module 520 is used to build multiple first tables based on multiple entity relationship data;
[0193] The dataset determination module 530 is used to obtain a training dataset based on multiple first tables; wherein the training dataset is used to train a pre-trained model of the obtained tables.
[0194] Figure 6 A schematic block diagram of a data processing apparatus provided in another embodiment of this disclosure is shown. Figure 6 As shown, in some embodiments of this disclosure, the data acquisition module of the data processing apparatus may include:
[0195] The key-value pair extraction unit 611 is used to extract multiple key-value pair data from the entity description information;
[0196] The data construction unit 612 is used to obtain entity relationship data corresponding to each key-value pair based on the entity described by the entity description information and each key-value pair data in multiple key-value pair data.
[0197] Optionally, such as Figure 6 As shown, in some embodiments of this disclosure, the table construction module of the data processing apparatus may include:
[0198] The set determination unit 621 is used to determine, based on multiple entity relationship data, M entity relationship data sets corresponding to M subject information respectively, where M is an integer greater than or equal to 2;
[0199] The table acquisition unit 622 is used to construct multiple first tables based on M entity relationship data sets.
[0200] Optionally, table retrieval unit 622 is specifically used for:
[0201] From M entity relation data sets, identify K entity relation data sets that have at least N identical predicate information, where N is an integer greater than or equal to 1 and K is an integer greater than or equal to 2;
[0202] Based on the K subject information and N identical predicate information corresponding to the K entity relationship data sets, populate the table header information in the table template;
[0203] Using the object information from the K entity relationship datasets, fill the table value area in the table template to obtain a relationship table with K subject information.
[0204] Optionally, table retrieval unit 622 is specifically used for:
[0205] Divide the K entity relation data sets into L groups of entity relation data sets; where L is an integer greater than or equal to 2;
[0206] Based on each entity relation data set in the L sets of entity relation data sets, a relation table of multiple subject information corresponding to each entity relation data set is obtained;
[0207] A horizontally stacked table is obtained by horizontally combining the relation tables corresponding to each set of entity relation data.
[0208] Optionally, table retrieval unit 622 is specifically used for:
[0209] Based on the entity relation data set corresponding to the i-th subject information among M subject information, construct an entity table corresponding to the i-th subject information; where i is a positive integer less than or equal to M.
[0210] Optionally, table retrieval unit 622 is specifically used for:
[0211] A vertically stacked table is obtained by vertically combining the entity tables corresponding to each of the M subject information.
[0212] Figure 7 A schematic block diagram of a data processing apparatus provided in another embodiment of this disclosure is shown. Figure 7 As shown, in this embodiment, the table construction module includes:
[0213] Relation determination unit 711 is used to determine the hierarchical relationship between multiple predicate information;
[0214] The table filling unit 712 is used to fill the table template with multiple entity relationship data corresponding to multiple predicate information based on the hierarchical relationship between multiple predicate information to obtain a hierarchical table.
[0215] Optionally, such as Figure 7 As shown, the data processing apparatus may further include:
[0216] The hyperparameter iteration module 720 is used to perform X iterations based on multiple entity relationship data to determine the table hyperparameters. The table hyperparameters include at least one quantity parameter for constructing multiple first tables; where X is an integer greater than or equal to 2.
[0217] Among them, the j-th iteration operation in the X iteration operations includes:
[0218] Based on multiple entity relationship data and the table hyperparameters updated in the (j-1)th time, construct multiple third tables;
[0219] Based on the size information of each of the multiple third tables, determine the distribution of size information of the multiple third tables;
[0220] Based on the distribution of scale information, the table hyperparameters are updated for the jth time; where j is an integer greater than or equal to 1.
[0221] Optionally, the scale information includes at least one of the following: number of cells, number of rows, number of columns, number of character elements in cells, and number of character elements in the table.
[0222] Optionally, such as Figure 7 As shown, the data processing apparatus may further include:
[0223] Standard acquisition module 730 is used to acquire standard documents for the target industry;
[0224] Table extraction module 740 is used to extract multiple second tables from a standard document;
[0225] Specifically, the dataset determination module in the data processing device can be used for:
[0226] The training dataset is obtained based on multiple first tables and multiple second tables.
[0227] Optionally, such as Figure 7 As shown, the data processing device may also include a marking module 750.
[0228] In one example, the marking module 750 is used for:
[0229] For each second table in the training dataset, retrieve the corresponding reference text in the standard documents;
[0230] The supervision label information corresponding to the second table is determined based on the cited text;
[0231] Based on the second table and the corresponding supervision label information, a first subset of training data for the supervised task is obtained.
[0232] In another example, the marking module 750 is used for:
[0233] For each first table in the training dataset, obtain the corresponding entity relationship data;
[0234] Based on entity relationship data, determine the supervision label information corresponding to the first table;
[0235] Based on the first table and the supervision label information, a second subset of training data for the supervised task is obtained.
[0236] For example, the marking module 750 is used for:
[0237] Based on entity relationship data and preset question-and-answer templates, construct the query statements and answers corresponding to the first table;
[0238] Use the answers as supervisory label information for the first table and the query statement.
[0239] For example, the marking module 750 is used for:
[0240] Based on the information types of multiple pieces of information in the entity relationship data, determine the cell types of multiple cells in the first table and / or the relationships between multiple cells;
[0241] Use the cell types of multiple cells and / or the relationships between multiple cells as the supervisory label information for the first table.
[0242] According to embodiments of this disclosure, this disclosure also provides a form processing apparatus. Figure 8 A schematic block diagram of a table processing apparatus provided according to an embodiment of the present disclosure is shown. Figure 8 As shown, the form processing device may include:
[0243] The semantic representation output module 810 is used to process the target table using a table pre-training model to obtain the semantic representation of the target table; wherein the table pre-training model is obtained based on the training dataset in any embodiment of this disclosure;
[0244] The task processing module 820 is used to perform table understanding tasks based on semantic representation and obtain task processing results.
[0245] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0246] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0247] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0248] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0249] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0250] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0251] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as data processing methods or spreadsheet processing methods. For example, in some embodiments, the data processing methods or spreadsheet processing methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the data processing methods or spreadsheet processing methods described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform a data processing method or a table processing method by any other suitable means (e.g., by means of firmware).
[0252] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0253] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0254] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0255] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0256] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0257] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0258] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0259] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A data processing method, comprising: Retrieve data on multiple entity relationships; Entity relation data consists of triples containing subject, predicate, and object information; Based on the aforementioned entity relationship data, multiple first tables are constructed, including: Based on the aforementioned multiple entity relationship data, determine M entity relationship data sets corresponding to each of the M subject information, where M is an integer greater than or equal to 2; Based on the M entity relationship data sets, construct multiple first tables; Based on the plurality of first tables, a training dataset is obtained; wherein, the training dataset is used to train a table pre-training model; wherein, the table pre-training model is used to perform semantic representation of the target table to perform a table understanding task, the table understanding task including at least one of cell classification, cell relationship classification, question-answering based cell location, table location, and question parsing; The construction of multiple first tables based on the M entity relationship data sets includes: From the M entity relation data sets, K entity relation data sets with at least N identical predicate information are determined; the entity relation data sets in the K entity relation data sets contain subject information, predicate information and object information; N is an integer greater than or equal to 1, K is an integer greater than or equal to 2 and K is less than or equal to M; Based on the K entity relationship data sets, multiple first tables are constructed, wherein the multiple first tables include at least one of the following: A relational table containing K subject information; A horizontally stacked table, in which the horizontally stacked table is obtained by horizontally combining the relation tables of subject information; The entity table corresponding to the i-th subject information, where i is a positive integer less than or equal to M; Vertical stacked table, which is obtained by vertically combining the entity tables corresponding to the subject information.
2. The method according to claim 1, wherein, The acquisition of multiple entity relationship data includes: Extract multiple key-value pairs from the entity description information; Based on the entity described by the entity description information and each key-value pair in the plurality of key-value pair data, entity relationship data corresponding to each key-value pair data is obtained.
3. The method according to claim 1, wherein, The steps to obtain the relation table of the K subject information include: Based on the K subject information corresponding to the K entity relationship data sets and the N identical predicate information, fill in the table header information in the table template; Using the object information from the K entity relationship data set, the table value area in the table template is filled to obtain the relationship table of the K subject information.
4. The method according to claim 3, wherein, The steps for obtaining the horizontally stacked table include: Divide the K entity relationship data sets into L groups of entity relationship data sets; where L is an integer greater than or equal to 2; Based on each of the L sets of entity relationship data, a relation table of multiple subject information corresponding to each set of entity relationship data is obtained; A horizontal stacked table is obtained by horizontally combining the relationship tables corresponding to each set of entity relationship data.
5. The method according to claim 1, wherein, The step of obtaining the entity table corresponding to the i-th subject information includes: Based on the entity relation data set corresponding to the i-th subject information among the M subject information, an entity table corresponding to the i-th subject information is constructed; where i is a positive integer less than or equal to M.
6. The method according to claim 5, wherein, The steps for obtaining the vertically stacked table include: A vertically stacked table is obtained by vertically combining the entity tables corresponding to each of the M subject information.
7. The method according to claim 1 or 2, wherein, Based on the multiple entity relationship data, multiple first tables are constructed, including: Determine the hierarchical relationship between multiple predicate information; Based on the hierarchical relationship between the multiple predicate information, the entity relationship data corresponding to the multiple predicate information are filled into the table template to obtain the hierarchical table.
8. The method according to claim 1 or 2, further comprising: Based on the multiple entity relationship data, X iterations are performed to determine the table hyperparameters, which include at least one quantity parameter for constructing the multiple first tables; where X is an integer greater than or equal to 2. Wherein, the j-th iteration operation in the X iteration operations includes: Based on the multiple entity relationship data and the table hyperparameters updated in the (j-1)th time, multiple third tables are constructed; Based on the size information of each of the plurality of third tables, the distribution of the size information of the plurality of third tables is determined; Based on the scale information distribution, the table hyperparameters are updated for the jth time; where j is an integer greater than or equal to 1.
9. The method according to claim 8, wherein, The scale information includes at least one of the following: number of cells, number of rows, number of columns, number of character elements in cells, and number of character elements in the table.
10. The method according to claim 1 or 2, further comprising: Obtain standard documents for the target industry; Extract multiple second tables from the standard document; The training dataset obtained based on the plurality of first tables includes: The training dataset is obtained based on the plurality of first tables and the plurality of second tables.
11. The method of claim 10, further comprising: For each second table in the training dataset, retrieve the reference text corresponding to the second table in the standard document; Based on the referenced text, determine the supervision label information corresponding to the second table; Based on the second table and the corresponding supervision label information, a first subset of training data for the supervised task is obtained.
12. The method according to claim 1 or 2, further comprising: For each first table in the training dataset, obtain the corresponding entity relationship data; Based on the entity relationship data, determine the supervision label information corresponding to the first table; Based on the first table and the supervision label information, a second subset of training data for the supervised task is obtained.
13. The method according to claim 12, wherein, The step of determining the supervision label information corresponding to the first table based on the entity relationship data includes: Based on the entity relationship data and the preset question-and-answer template, construct the query statement and answer corresponding to the first table; The answer is used as the supervision label information for the first table and the query statement.
14. The method according to claim 12, wherein, The step of determining the supervision label information corresponding to the first table based on the entity relationship data includes: Based on the information types of multiple pieces of information in the entity relationship data, determine the cell types of multiple cells in the first table and / or the relationships between the multiple cells; The cell types of the multiple cells and / or the relationships between the multiple cells are used as the supervision label information of the first table.
15. A table processing method, comprising: The target table is processed using a pre-trained table model to obtain a semantic representation of the target table; wherein the pre-trained table model is obtained based on the training dataset as described in any one of claims 1-14; Based on the semantic representation, a table comprehension task is performed to obtain the task processing result.
16. A data processing apparatus, comprising: The data acquisition module is used to acquire data on relationships between multiple entities. Entity relation data consists of triples containing subject, predicate, and object information; The table construction module is used to construct multiple first tables based on the multiple entity relationship data; The dataset determination module is used to obtain a training dataset based on the plurality of first tables; wherein the training dataset is used to train a table pre-training model; wherein the table pre-training model is used to perform semantic representation of the target table to perform a table understanding task, the table understanding task including at least one of cell classification, cell relationship classification, question-answering based cell location, table location, and question parsing; The table construction module includes: a set determination unit, used to determine M entity relationship data sets corresponding to M subject information respectively based on multiple entity relationship data, where M is an integer greater than or equal to 2; and a table acquisition unit, used to construct multiple first tables based on the M entity relationship data sets. Specifically, the table acquisition unit is used to determine K entity relationship data sets with at least N identical predicate information from the M entity relationship data sets; the entity relationship data sets in the K entity relationship data sets contain subject information, predicate information, and object information; N is an integer greater than or equal to 1, K is an integer greater than or equal to 2 and K is less than or equal to M; based on the K entity relationship data sets, multiple first tables are constructed, wherein the multiple first tables include at least one of the following: a relationship table of K subject information; a horizontally stacked table, wherein the horizontally stacked table is obtained by horizontally combining the relationship tables of subject information; an entity table corresponding to the i-th subject information, wherein i is a positive integer less than or equal to M; and a vertically stacked table, wherein the vertically stacked table is obtained by vertically combining the entity tables corresponding to subject information.
17. A form processing apparatus, comprising: A semantic representation output module is used to process a target table using a pre-trained table model to obtain a semantic representation of the target table; wherein the pre-trained table model is obtained based on the training dataset as described in any one of claims 1-14; The task processing module is used to perform a table understanding task based on the semantic representation and obtain the task processing result.
18. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.
20. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-15.
Citation Information
Patent Citations
Information processing method and device
CN108694206A
Table information extraction method and device, equipment and medium
CN114818710A
Electronic component knowledge reasoning and knowledge graph construction method and system and medium
CN115080759A