Pre-training method and device based on table replacement invariance

By introducing a permutation invariance mechanism and a comparison learning method in the tabular pre-training method, the problem of insufficient stability and generalization ability when the sequence of feature columns changes in the prior art is solved, and higher robustness and lower data annotation cost are achieved.

CN119990085APending Publication Date: 2025-05-13INSTITUTE OF COMPUTING INNOVATION ZHEJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510160821.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing tabular pre-training methods have limitations when dealing with changes in the sequence of feature columns, resulting in insufficient stability and generalization capabilities of the model, and require a large amount of labeled data for training.

Method used

A pre-training method based on the Transformer model is adopted, a permutation invariance mechanism is introduced, and a positive and negative sample pair is constructed through a comparative learning method, unsupervised pre-training is performed, and a small amount of labeled data is used to align the model in the second-stage alignment training.

Benefits of technology

It improves the robustness of the model to adjust the order of table feature columns, reduces the cost of data annotation, and enhances the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990085A_ABST
    Figure CN119990085A_ABST
Patent Text Reader

Abstract

The invention discloses a pre-training method and device based on table replacement invariance. The method comprises the following steps: in a first stage, constructing positive and negative sample pair data according to row and column replacement invariance in a table, and then constructing a pre-training task by using a comparative learning method; in order to enable the pre-training model to adapt to various downstream tasks, the second stage is that the table is aligned with the downstream tasks, and the downstream tasks of the table comprise table questions and answers, table classification, table data generation, table abstract extraction and the like. And according to different downstream tasks, performing joint alignment training on the pre-training model and the language large model with the tasks so as to obtain the pre-training model capable of adapting to various downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology and the field of computer natural language large models, and in particular to a pre-training method and device based on table permutation invariance. Background Art

[0002] With the advent of the big data era, tabular data, as a highly structured data form, is widely used in various industries, including finance, healthcare, e-commerce, supply chain management, and government statistics. Tabular data usually contains multiple feature fields (columns) and records (rows), and plays a vital role in corporate decision-making and analysis. In recent years, with the development of machine learning and deep learning technologies, how to effectively use tabular data to improve model performance has become an important research direction in the fields of data science and artificial intelligence.

[0003] In the processing and analysis of tabular data, in traditional machine learning models such as XGBoost, the change of the order of rows and columns in the table will not affect the results; from the perspective of data engineers, the data features of the rows and columns in the table are changed, and the content and information represented in the table remain the same. However, deep learning models are usually sensitive to the order of input data, especially the order of columns in the table. These models often assume that the input features are arranged in a fixed order, but in actual scenarios, the order of columns in the data table may change due to reasons such as data source, preprocessing method, and data integration. This order dependency limits the robustness and flexibility of the model in practical applications, especially when multiple tabular data sources are integrated or automated data preprocessing is performed, changes in the order of feature columns may cause a significant decrease in model performance.

[0004] In recent years, pre-trained language models have achieved remarkable results in the field of natural language processing. However, there are still some challenges when applying these models directly to tabular data. Tabular data is fundamentally different from text data. The features (columns) and samples (rows) in the table have independent semantic meanings, and the adjustment of the order does not change the actual meaning of the data. Therefore, in tabular data modeling, if the model can be invariant to the permutation of the rows and columns of the table, its robustness in processing different tabular data can be significantly improved.

[0005] Existing table pre-training methods, such as TabTransformer and table-based deep learning models, have improved the processing capabilities of tabular data to a certain extent, but they still have certain limitations when dealing with changes in the order of feature columns. Most of these methods rely on a fixed order of input features for modeling. Once the order of feature columns changes, the prediction results of the model may deviate significantly from the original results, resulting in insufficient stability and generalization of the model. In addition, existing models often require a large amount of labeled data during training, and it is costly to obtain large-scale, high-quality labeled tabular data. Therefore, how to effectively pre-train tabular data and improve the model's robustness to the order of feature columns has become a technical problem that needs to be solved urgently. Summary of the invention

[0006] In order to solve the problems existing in the background technology, the present invention provides a pre-training method and device based on table permutation invariance. By introducing the permutation invariance mechanism, the model can still maintain consistent prediction results when the order of feature columns in the input table is adjusted.

[0007] According to one aspect of an embodiment of the present application, a pre-training method based on table permutation invariance is provided, the method comprising:

[0008] Collect tabular data from multiple sources;

[0009] Clean the tabular data and generate structured information;

[0010] Construct positive and negative sample pairs based on the permutation invariance of rows and columns in the table;

[0011] Perform structured data augmentation on tabular data;

[0012] Use the Transformer model structure without adding position information for pre-training;

[0013] Use the loss function based on contrastive learning to train the model and obtain a pre-trained model;

[0014] Align the pre-trained model with the downstream tasks so that the downstream tasks can better understand the table information, and thus the pre-trained model can adapt to various downstream tasks.

[0015] According to another aspect of an embodiment of the present application, a pre-training device based on table permutation invariance is also provided, including:

[0016] Data collection module, used to collect tabular data from multiple sources;

[0017] Data cleaning module, used to clean table data and generate structured information;

[0018] The sample pair construction module is used to construct positive and negative sample pairs based on the permutation invariance of rows and columns in the table;

[0019] Data enhancement module, used to perform structured data enhancement on tabular data;

[0020] Model pre-training module, used for pre-training using the Transformer model structure without adding position information;

[0021] The model training module is used to train the model using a loss function based on contrastive learning to obtain a pre-trained model;

[0022] The alignment training module is used to align the pre-trained model with the downstream tasks so that the downstream tasks can better understand the table information, so that the pre-trained model can adapt to various downstream tasks.

[0023] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, a pre-training method based on table permutation invariance is implemented.

[0024] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a pre-training method based on table permutation invariance when executing the program.

[0025] The beneficial effects of the present invention are:

[0026] The first-stage table pre-training uses contrastive learning and is an unsupervised training that does not require data labeling. The second-stage joint alignment training only requires labeling of a small amount of data, which can greatly reduce the amount of data labeling.

[0027] By introducing the permutation invariance mechanism, the model can maintain consistent prediction results when the order of feature columns in the input table is adjusted, thereby improving the generalization ability and robustness of the model.

[0028] This pre-trained model can be considered as a table encoder, which can compress the information of the table into a stack of vectors. The vector contains the global information of the table and is input into the downstream model, which greatly reduces the input length of the downstream model and improves the performance of downstream tasks related to the table. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A schematic flow chart of a pre-training method based on table permutation invariance provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0031] The present invention will be further described below in conjunction with the accompanying drawings and specific implementations.

[0032] like Figure 1 As shown, the overall concept of the present invention is:

[0033] Phase 1: Collect, clean, and enhance tabular data, construct positive and negative sample pairs based on the permutation invariance of tabular data rows and columns, and use contrastive learning methods for pre-training.

[0034] Phase 2: The model trained in the first phase is trained with downstream tasks so that the downstream tasks can better understand the table information, so that the pre-trained model can adapt to various downstream tasks.

[0035] Specifically:

[0036] 1. Form data collection

[0037] Collect high-quality tabular data from a variety of sources, and obtain diverse tabular data from internal databases, public datasets, web crawlers, etc. Prioritize datasets with rich features (columns) and a large number of samples (rows) to ensure the generalization ability of the pre-trained model, including corporate financial statements, e-commerce order data, medical data, scientific research experimental data, etc.

[0038] 2. Table data cleaning

[0039] Clean existing tabular data and generate structured information (such as column names, row names, cell contents, table summaries, etc.), including data standardization, handling of missing values, outliers, duplicate rows, etc. Use interpolation or mean filling for missing values. Remove duplicate records and irrelevant feature columns. Delete tables with less than 2 columns and less than 5 rows.

[0040] 3. Construct positive and negative sample pairs for contrastive learning

[0041] The tables corresponding to the positive samples should have the same semantics even if the order of rows and columns is different.

[0042] In an optional embodiment, in order to construct enough positive samples, the same table can be reused to construct multiple positive sample pairs: 1. The order of rows is different, but the order of columns is the same; 2. The order of rows is the same, but the order of columns is different; 3. The order of rows and columns is different. The sample pairs constructed in the above three cases contain the same information in the table.

[0043] Negative sample pairs only need to contain different information in the table. There are two cases: 1. From the same table, extract sub-tables consisting of different rows and columns; 2. Negative sample pairs between different tables.

[0044] Randomly select samples from different tabular data as negative sample pairs to ensure that they are semantically significant different.

[0045] 4. Structured Data Augmentation for Tabular Data

[0046] Make the pre-trained model robust and generalizable to different representations of tables. Specifically, the enhancement of table data includes not only randomly permuting the order of rows and columns, but also the conversion of tables in multiple forms of expression, such as Markdown, plain text (string), Schema, etc. Different forms of expression contain the same table information. These enhancement operations can help the model understand the semantic structure of the table without relying on its specific presentation method, thereby ensuring that the model can work stably and effectively in different environments and application scenarios, and can train the model to maintain accurate understanding capabilities under different input formats.

[0047] 5. Pre-training model structure

[0048] In this embodiment, a Transformer model-based architecture is used for pre-training of tabular data, and positional encoding is deliberately not used. The Transformer model structure establishes a global dependency relationship between the rows and columns of the table through a multi-head attention mechanism. Since no position information is added, the model is insensitive to the position of the table rows and columns, improving the robustness to the replacement of tabular data.

[0049] 6. Model Pre-training

[0050] During the pre-training stage of tabular data, this embodiment introduces a loss function based on contrastive learning, such as InfoNCE Loss, to ensure that the model can maintain consistent understanding and reasoning capabilities in different tabular expressions and permuted data, maximize the similarity of positive sample pairs, and minimize the similarity of negative sample pairs, and finally converge to a one-stage pre-trained model.

[0051] 7. Two-stage pre-training model alignment

[0052] 7.1. Construct a two-stage aligned training dataset, and construct training datasets for table tasks such as table question answering, table classification, table data generation, and table summary extraction.

[0053] 7.2. Joint training of pre-trained models and downstream task models.

[0054] According to an embodiment of the present application, a pre-training device for implementing the above-mentioned pre-training method is also provided, including:

[0055] Data collection module, used to collect tabular data from multiple sources;

[0056] Data cleaning module, used to clean table data and generate structured information;

[0057] The sample pair construction module is used to construct positive and negative sample pairs based on the permutation invariance of rows and columns in the table;

[0058] Data enhancement module, used to perform structured data enhancement on tabular data;

[0059] Model pre-training module, used for pre-training using the Transformer model structure without adding position information;

[0060] The model training module is used to train the model using a loss function based on contrastive learning to obtain a pre-trained model;

[0061] The alignment training module is used to align the pre-trained model with the downstream tasks so that the downstream tasks can better understand the table information, so that the pre-trained model can adapt to various downstream tasks.

[0062] According to an embodiment of the present application, a computer-readable storage medium is also provided, on which a computer program is stored. When the program is executed by a processor, a pre-training method based on table permutation invariance is implemented.

[0063] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0064] According to an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a pre-training method based on table permutation invariance when executing the program.

[0065] The electronic device may be any electronic device in the electronic device group. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.

[0066] In summary, the core technology of the present invention is to introduce a permutation invariance mechanism to allow the pre-trained model to learn the features in the table. Similar to the situation when a person processes a table task, when the order of rows and columns of the input table is adjusted, the prediction results can still be kept consistent, thereby improving the generalization ability and robustness of the model.

[0067] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A pre-training method based on table permutation invariance, characterized in that: The following steps are involved: Collect tabular data from multiple sources; Clean the table data and generate structured information; Construct positive and negative sample pairs based on the permutation invariance of rows and columns in the table; Perform structured data augmentation on tabular data; Use the Transformer model structure without adding position information for pre-training; Use the loss function based on contrastive learning to train the model and obtain a pre-trained model; Align the pre-trained model with the downstream tasks so that the downstream tasks can better understand the table information, and thus the pre-trained model can adapt to various downstream tasks.

2. The pre-training method based on table permutation invariance according to claim 1, characterized in that: The tabular data collected from various sources include corporate financial statements, e-commerce order data, medical data, and scientific research experiment data.

3. The pre-training method based on table permutation invariance according to claim 1 or 2, characterized in that: The cleaning of the table data includes processing missing values, outliers, and duplicate rows; wherein missing values ​​are filled using interpolation or mean value, duplicate records and irrelevant feature columns are removed, and tables with less than 2 columns and less than 5 rows are deleted.

4. The pre-training method based on table permutation invariance according to claim 1, characterized in that: The constructing of positive and negative sample pairs according to the permutation invariance of rows and columns in the table includes: The tables corresponding to the positive samples have the same semantics, even if the order of rows and columns is different; The tables corresponding to negative samples are semantically different.

5. The pre-training method based on table permutation invariance according to claim 1 or 4, characterized in that: The structured data enhancement of the table data includes randomly replacing the order of rows and columns or converting the table into multiple expression forms.

6. The pre-training method based on table permutation invariance according to claim 1, characterized in that: The Transformer model structure without adding position information establishes a global dependency between rows and columns of a table through a multi-head attention mechanism, thereby improving the robustness to table data replacement.

7. The pre-training method based on table permutation invariance according to claim 4 or 7, characterized in that: The model training using the loss function based on contrastive learning includes maximizing the similarity of positive sample pairs while minimizing the similarity of negative sample pairs.

8. A pre-training device based on table permutation invariance, characterized in that: include: Data collection module, used to collect tabular data from multiple sources; Data cleaning module, used to clean table data and generate structured information; The sample pair construction module is used to construct positive and negative sample pairs based on the permutation invariance of rows and columns in the table; Data enhancement module, used to perform structured data enhancement on tabular data; Model pre-training module, used for pre-training using the Transformer model structure without adding position information; The model training module is used to train the model using a loss function based on contrastive learning to obtain a pre-trained model; The alignment training module is used to align the pre-trained model with the downstream tasks so that the downstream tasks can better understand the table information, so that the pre-trained model can adapt to various downstream tasks.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the pre-training method based on table permutation invariance as described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the pre-training method based on table permutation invariance as described in any one of claims 1 to 7 is implemented.