Large model training data set construction method based on table type data

Through automated processing methods based on table header association and deformable template model, the problems of low efficiency and insufficient diversity of table data processing in the construction of large model training data sets are solved, and efficient and accurate Q&A pair generation is achieved, enhancing the adaptability and performance of the model.

CN120353893APending Publication Date: 2025-07-22NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510426471.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When building large model training data sets, the prior art faces the problems of low efficiency, high cost, insufficient automation and insufficient diversity in table data processing, especially in vertical applications, which are difficult to adapt to table format changes and complex query needs.

Method used

The automated processing method based on table header association and the extraction template are adopted, combined with the deformable template model and dynamic adjustment mechanism, and the table format is automatically adapted to the field recognition module, and the data set diversity is increased using natural language data enhancement strategies to achieve efficient and accurate question-and-answer pair generation.

Benefits of technology

It improves the efficiency and accuracy of data processing, enhances the generalization ability of the model, can adapt to different table formats and complex query needs, generates diverse question-and-answer pairs, and improves the application performance of the model in vertical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353893A_ABST
    Figure CN120353893A_ABST
Patent Text Reader

Abstract

The invention discloses a large model training data set construction method based on table type data. The method comprises the steps that 1, data preprocessing is carried out; step 2, constructing an automatic processing method and an extraction template based on header association; step 3, constructing a deformable template model; and step 4, binding question types and fields. According to the method, manual intervention is reduced by applying an automatic processing method based on header association and an extraction template, a question and answer template is preset, key fields are matched by using a regular expression, a table format is automatically adapted through a field identification module, and then data quality is guaranteed through a data quality evaluation and feedback mechanism; a deformable template model is provided to be combined with a dynamic adjustment mechanism, and field identification, dynamic adjustment and mapping modules cooperate to ensure accurate query; heterogeneous table data is standardized, problem types and corresponding fields are automatically bound to improve the processing efficiency, and the diversity of a data set is increased by using a data enhancement strategy based on a natural language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a large model training data set based on tabular data, belonging to the technical fields of data processing and artificial intelligence. Background Art

[0002] The development of artificial intelligence has entered the stage of generative large models. Generative large models represented by Deepseek and GPT are accelerating the evolution of AI from perception and understanding to generation and creation, continuously giving birth to new products, new business forms, and new models.

[0003] However, although generative large models have shown great potential in multiple fields, their applications in vertical fields still face many challenges. Most existing large models lack a deep understanding of industry expertise and industry background, which makes it difficult for them to directly meet industry needs when dealing with professional tasks. By integrating the characteristics of strong interactivity, powerful generative ability, flexible adaptability, and self-learning and self-optimization of large models into the business of specific vertical fields, it can not only help users better understand industry knowledge but also accelerate the application of large models in the industry. When training and fine-tuning large models in vertical fields, a large amount of domain training data is required to help the model learn specific domain knowledge, thereby improving the model's understanding and application ability in specific domains.

[0004] The sources and forms of fine-tuning data are diverse, and tabular data is a common form. Tabular data has the characteristics of high structuralization, large information density, and strong data correlation, and is widely used in various professional fields. However, when extracting tabular data to construct a large model training data set, data processing faces many challenges: (1) Existing question-and-answer systems rely on manual annotation of tabular data to construct data sets, which is inefficient and costly. Large-scale tables need to be annotated row by row and field by field, consuming a lot of time and effort. Annotation also requires professional knowledge, and differences in understanding among different annotators easily lead to errors and inconsistencies, affecting data quality. (2) There are various types and formats of tables, and traditional regular expression methods have poor flexibility and require manual writing of rules for different table formats. For example, the date formats in industries such as finance and education are different and need to be adapted separately. When the table format changes, the original rules are likely to fail and are difficult to adjust dynamically. In addition, this method only supports simple queries and cannot handle multi-field calculations, conditional filtering, or complex semantic problems. (3) Traditional information extraction methods lack automation and require manual adjustment of rules or templates, lacking self-adaptive ability. When the data volume or distribution changes, they cannot be automatically optimized. Traditional methods do not introduce a data augmentation mechanism, which makes the generated data set difficult to have sufficient diversity, limiting the efficiency and diversity of generating question-and-answer pairs based on tables. The above problems limit the efficiency and diversity of generating question-and-answer pairs based on tables. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method for constructing a large model training dataset based on tabular data, which uses an automated processing method and extraction template based on header association to reduce manual intervention, presets a question and answer template and matches keyword fields with regular expressions, automatically adapts the table format through a field recognition module, and then ensures data quality through a data quality assessment and feedback mechanism; at the same time, a deformable template model combined with a dynamic adjustment mechanism is proposed, and its field recognition, dynamic adjustment, and mapping modules cooperate to ensure accurate query; standardize heterogeneous table data, automatically bind question types to corresponding fields to improve processing efficiency, and use a natural language-based data augmentation strategy to increase the diversity of the dataset.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a method for constructing a large model training dataset based on tabular data, which uses an automated processing method and extraction template based on header association to achieve efficient generation of questions and answers, constructs a deformable template model and combines dynamic adjustment to ensure the diversity of generated data, standardizes heterogeneous data, automatically binds questions and fields, uses a natural language-based data augmentation strategy to increase the diversity of the dataset; the method includes the following steps:

[0007] Step 1, data preprocessing;

[0008] Step 2, construct an automated processing method and extraction template based on header association;

[0009] Step 3, construct a deformable template model;

[0010] Step 4, bind question types to fields.

[0011] Further, in the step 1, the data preprocessing process is: extract tables of different styles and store them in a unified form, where the table includes a table header, a table body, an attribute name, and an attribute value. The format of each table is different, and different types of tables need to be converted into a unified format, which not only facilitates subsequent reading and processing but also eliminates potential problems caused by format differences.

[0012] Further, in the step 2, different automated extraction templates are set according to the type of question, the types of questions are classified, the keywords to be searched are determined, the information in the table is read, and the keywords are extracted with regular expressions to complete the template to generate a dataset.

[0013] Furthermore, the automated processing method based on header association is:

[0014] Taking the first column of the table as the reference column, its table header as the key identifier, and pairing it with the table headers of other columns one by one to form a dynamic field association; for each pair of table headers, extract their corresponding keywords and column contents respectively, and generate structured question and answer pairs according to the preset automated extraction template.

[0015] After generating the Q&A pairs for a certain column, the header of the reference column will continue the same operation with the header of the next column until all columns of the entire table have been traversed and processed. The process is as Figure 4 shown. Through this systematic processing flow, it is possible to achieve comprehensive coverage and in-depth mining of table data, providing high-quality data support for subsequent natural language processing tasks.

[0016] Furthermore, an automated processing method based on header association is used to extract questions and answers. According to the question template, the column where the keyword is located is matched, and the queried keyword is mapped to a unique identifier. Through the unique identifier, it is further mapped to the query table keyword, and the corresponding value of the relevant column in the table is queried through the query table keyword, thereby generating the answer to the question.

[0017] Furthermore, in the process of using the automated processing method based on header association and the extraction template, a field recognition module is also used to dynamically analyze the headers and fields of the table. This module can automatically judge the type and function of the fields according to the overall structure and context information of the table.

[0018] Further, in step 3, the construction of the deformable template model is composed of multiple modules working together, aiming to efficiently and accurately process various types of table data. Specifically, it includes a field recognition module, a dynamic adjustment module, a mapping module, and a question template generation module, where:

[0019] The field recognition module uses natural language processing technology to deeply analyze the table in a unified format, accurately distinguish the field type and function, and at the same time combine the keywords of specific questions to strengthen the capture of key information;

[0020] The dynamic adjustment module monitors the changes of the table in real time according to the results of field recognition. With the help of the rule library and historical data accumulation, once the table structure changes or the field expression differs, it automatically makes adaptive adjustments to the template architecture and regular matching rules;

[0021] The mapping module assigns a unique identifier to the extracted keywords through a hash algorithm, and dynamically optimizes the mapping rules according to the field characteristics, thereby establishing an accurate connection between the keywords and the actual data;

[0022] The question template generation module generates a variety of Q&A templates using a template engine according to the question category, the identified key information, predefined patterns, and semantic rules;

[0023] These templates have multiple different versions for various questions, greatly enhancing the diversity of Q&A pairs and meeting the requirements in different scenarios.

[0024] Furthermore, in step 4, the process of binding problem types to fields is as follows: The system automatically identifies the problem type. By analyzing the keywords in the problem text, it determines what type the problem belongs to. According to the type, it calls the corresponding field binding function to bind the fields related to the problem in the table, extracts data from the table using the template model, and generates corresponding question-and-answer pairs. Through this binding mechanism, the system can quickly locate relevant field information, accurately extract data and generate diverse question-and-answer pairs using the template model, improving the accuracy and efficiency of data processing and meeting complex question-and-answer requirements. During the process of table data processing, in order to achieve efficient extraction of problem template keywords and generation of question-and-answer pairs.

[0025] In summary, for the extracted data, this solution uses natural language data augmentation strategies to further enhance the diversity and richness of the dataset and strengthen the generalization ability of the model. By means of synonym replacement, sentence restructuring, context expansion, etc., the generated question-and-answer pairs are optimized and expanded to generate more variable expression forms, enabling the model to better adapt to inputs with different language styles and expression habits, thereby improving the performance and robustness of the model in practical applications. Finally, the question-and-answer pairs in the generated dataset are verified. Ensure that the question-and-answer pairs meet the requirements of domain fine-tuning in terms of content, logic, and language expression, etc., thus ensuring data quality and providing a basis for subsequent training of large models in vertical domains.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0027] The present invention proposes a method for constructing a large model training data set based on tabular data. Compared with existing technologies, it mainly includes the following innovations: (1) Extracting tables from various documents and storing them in a unified format, which solves the problem that traditional methods are difficult to handle due to diverse table formats, improves the subsequent reading and processing efficiency, and avoids potential errors caused by format differences; (2) Proposing an automated processing method and extraction template based on header association, setting different extraction templates according to the problem type. Traditional methods require manual writing of regular expressions and are difficult to adapt to table changes. However, the present invention can automatically adapt to different table structures, accurately extract key information by dynamically analyzing headers and fields, support complex problem types, and overcome the disadvantages of low efficiency and incomplete reading of table content in traditional methods; (3) Proposing a deformable template model. The model consists of multiple modules working together. The field recognition module accurately discriminates field characteristics, the dynamic adjustment module responds to table structure changes in real time, the mapping module establishes accurate connections, and the problem template generation module improves the diversity of question and answer pairs, significantly enhancing the processing ability of complex tabular data and improving the accuracy and adaptability of data extraction; (4) Proposing a mechanism for binding problem types and fields, realizing fast and accurate matching of questions and table fields, greatly improving the data processing efficiency and accuracy, and using a natural language data augmentation strategy to further enhance the diversity and richness of the data set, enhance the generalization ability of the model, improve the performance and robustness of the model in practical applications, enrich the data set, and enhance the generalization ability of the model. Description of the Drawings

[0028] Figure 1 is a flowchart of the present invention.

[0029] Figure 2 is a schematic diagram of the automated header association and extraction template construction process.

[0030] Figure 3 is an architecture diagram of the deformable template model.

[0031] Figure 4 is a schematic diagram of the automated processing method based on header association. Detailed Embodiments

[0032] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] As Figures 1 to 4, A method for constructing a large model training dataset based on tabular data mainly includes: (1) Using an automated processing method and extraction template based on header association to reduce manual intervention. Preset question-and-answer templates according to table types, match keyword fields with regular expressions. Analyze the table headers with the help of a field recognition module to automatically adapt to the table format. Verify and optimize the templates through a data quality assessment and feedback mechanism to ensure data quality. (2) Proposing a deformable template model and a dynamic adjustment mechanism to enhance flexibility. This model includes modules such as field recognition, dynamic adjustment, and mapping, which are used to identify table header fields, optimize matching rules, and associate target values respectively to ensure accurate queries. (3) Standardize heterogeneous tabular data, unify the format to ensure data consistency, automatically identify problem types and bind corresponding fields, and improve processing accuracy and efficiency. Finally, use a natural language data augmentation strategy to increase the diversity of the dataset.

[0034] The detailed process is as Figure 1 shown below, and further specific introduction is as follows:

[0035] Step 1, Data preprocessing.

[0036] Extract tables from multiple different documents. Let the document set be D = {d1, d2,..., d n}, and each document d i may contain several tables T ij (j = 1, 2,..., m i ). After extracting all the tables, store them in a unified form. A table usually consists of a table header H, a table body B, an attribute name A, and an attribute value V. Since the formats of various tables are different, different types of tables need to be converted into a unified format

[0037] T′ ij = UnifyFormat(T ij ) (1)

[0038] where T ij is the original table, UnifyFormat() is a function for unifying the table format, and T′ ij is the table in the unified format. The unified format is convenient for subsequent reading and processing, and eliminates potential problems caused by format differences.

[0039] Step 2, Construct an automated processing method and extraction template based on header association.

[0040] Set different extraction templates T t according to the question type Q m . First, classify the types of questions, determine the keywords to be searched K, and use the regular expression R m to match the name fields that may appear in the table, which can be expressed as:

[0041]

[0042] Among them, the Match() function represents using regular expressions to match in the uniformly formatted table and returns the set of fields that match. For tables of different types, t takes different values;

[0043] During the whole process, the Field Identification Module (FIM) is used to dynamically analyze the table header and fields of the table. According to the overall structure S of the table and the context information C, the field type F t (such as numerical type, text type, etc.) and function F f (such as key field, auxiliary field, etc.) can be expressed as:

[0044] (F t , F f ) = FIM(S, C, T ij ′) (3).

[0045] Step 3: Construct a deformable template model.

[0046] After completing data preprocessing, automatic template extraction, and field identification, a deformable template model is constructed. This model is composed of four modules working together. For the field identification module, using natural language processing technology, for the uniformly formatted table T ij ′, through part-of-speech tagging and named entity recognition, the field type F t and function F f are accurately discriminated. At the same time, combined with the keywords K of specific problems, the capture of key information is strengthened. As shown in formula 4,

[0047] (F t , F f ) = NLP(T ij ′) ∩ K (4)

[0048] Among them, NLP() represents the natural language processing function.

[0049] Based on the output (F t , F f ) of the field identification module, the changes R b in the table are monitored in real time, and the historical data accumulation is H d . Once the structure of the table changes (such as adding columns, adjusting column order, etc.) or there are differences in field expressions, the template architecture T a and the regular matching rule R m will be automatically adjusted adaptively according to the rule library and historical data, which can be expressed as:

[0050] (T a ′, R′ m) = Adjust(R b , H d , (F t , F f ), Change(T ij ′)) (5)

[0051] Among them, Change(T ij ′) represents the change situation of the detected table T ij ′, and (T a ′, R′ m ) is the adjusted template architecture and regular matching rule.

[0052] For the mapping module, through the hash algorithm Hash, a unique identifier ID is assigned to the extracted keyword K, and the mapping rule M t , F f ) is dynamically optimized according to the field characteristics (F r , thus establishing an accurate connection between the keyword and the actual data, which can be expressed as:

[0053] ID = Hash(K) (6)

[0054] M r ′ = Optimize(M r , (F t , F f )) (7)

[0055] DataLink = Map(ID, M r ′, T ij ′) (8)

[0056] Among them, the Map() function represents establishing a data connection in the table T r ′ according to the unique identifier ID and the optimized mapping rule M ij ′.

[0057] For the question template generation module, according to the question category Q t , the identified key information (obtained by the field recognition module), the predefined pattern P, and the semantic rule S r , a rich variety of Q&A templates are generated using the template engine TE These templates set multiple different versions for various questions, greatly enhancing the diversity of Q&A pairs and meeting the needs of different scenarios, which can be expressed as:

[0058] Q tm = TE(Q t , IdentifiedInfo, P, S r ) (9)

[0059] Among them, IdentifiedInfo represents the identified key information, which is obtained by the field identification module.

[0060] Step 4: Bind the problem type to the field.

[0061] For the binding of the problem type to the field, the system automatically identifies the problem type Q t , for the problem type Q t Based on the field identification results (F t , F f ) of the template model, the corresponding class fields in the table are bound. Let the set of such class fields be N f ,

[0062] N f = Bind(Q t , I c , (F t , F f ), T ij ′) (10)

[0063] Among them, Bind() is used to implement the binding operation between the problem type and the specific fields in the table.

[0064] Use an automated processing method based on header association to extract questions and answers. According to the question template, match the column where the keyword is located. The queried keyword is mapped to a unique identifier, and through the unique identifier, it is further mapped to the query table keyword. By querying the query table keyword, the corresponding value of the relevant column in the table is queried, thereby generating the answer to the question.

[0065] By using natural language data augmentation strategies, further enhance the diversity and richness of the dataset and enhance the generalization ability of the model. Through means such as synonym replacement, sentence restructuring, and context expansion, optimize and expand the generated question-answer pairs to generate more variable expression forms, so that the model can better adapt to inputs with different language styles and expression habits, thereby improving the performance and robustness of the model in practical applications.

[0066] The above shows and describes the basic principles, main features, and advantages of the present invention. Those of ordinary skill in the art should understand that the above embodiments do not limit the protection scope of the present invention in any form. Any technical solutions obtained by means of equivalent replacement and the like fall within the protection scope of the present invention. The parts not involved in the present invention are the same as or can be implemented by the prior art.

Claims

1. A method for constructing a large model training dataset based on tabular data, characterized in that, It includes the following steps: Step 1, data preprocessing; Step 2, construct an automated processing method and extraction template based on header association; Step 3, construct a deformable template model; Step 4, bind problem types to fields.

2. The method for constructing a large model training data set based on tabular data according to claim 1, wherein, In the said Step 1, the data preprocessing process is: extract tables of different styles and store them in a unified form, where the table includes a header, a body, attribute names, and attribute values.

3. A method for constructing a large model training data set based on tabular data according to claim 1, characterized in that, In the said Step 2, set different automated extraction templates according to the type of the problem, determine the keywords to be searched after classifying the problem types, read the information in the table, and use regular expressions to extract the keywords to complete the template to generate a data set.

4. A method for constructing a large model training data set based on tabular data according to claim 3, characterized in that, The automated processing method based on header association is: Take the first column of the table as the reference column, and its header as the key identifier, pair it with the headers of other columns one by one to form a dynamic field association; for each pair of headers, extract their corresponding keywords and column contents respectively, and generate structured question-and-answer pairs according to the preset automated extraction template; After generating the question-and-answer pairs for a certain column, the header of the reference column will continue to perform the same operation with the header of the next column until all columns of the entire table have been traversed and processed.

5. A method for constructing a large model training data set based on tabular data according to claim 4, characterized in that, Use the automated processing method based on header association to extract questions and answers, match the column where the keyword is located according to the question template, the queried keyword is mapped to a unique identifier, and through the unique identifier, it is further mapped to the query table keyword, and the corresponding value of the relevant column in the table is queried through the query table keyword, so as to generate the answer to the question.

6. The method for constructing a large model training data set based on tabular data according to claim 3, wherein In the process of using the automated processing method and extraction template based on header association, use the field recognition module to dynamically analyze the headers and fields of the table. Specifically, the module automatically judges the type and function of the fields according to the overall structure and context information of the table.

7. A method for constructing a large model training data set based on tabular data according to claim 1, characterized in that In the said Step 3, the construction of the deformable template model is composed of multiple modules working together, including a field recognition module, a dynamic adjustment module, a mapping module, and a question template generation module, where: The field recognition module uses natural language processing technology to deeply analyze the table in a unified format, accurately distinguish the field type and function, and at the same time combine the keywords of specific problems to strengthen the capture of key information; The dynamic adjustment module monitors the changes of the table in real time according to the results of field recognition. With the help of the rule library and historical data accumulation, once the table has structural changes or field expression differences, it will automatically adjust the template architecture and regular matching rules adaptively; The mapping module assigns a unique identifier to the extracted keywords through a hash algorithm, and dynamically optimizes the mapping rules according to the field characteristics, so as to establish an accurate connection between the keywords and the actual data; The question template generation module generates a variety of question-and-answer templates using a template engine according to the question category, the identified key information, predefined patterns, and semantic rules.

8. A method for constructing a large model training data set based on tabular data according to claim 1, characterized in that, In the said Step 4, the process of binding problem types to fields is: the system automatically identifies the problem type, judges what type the problem belongs to by analyzing the keywords in the problem text, calls the corresponding field binding function according to the type, binds the fields related to the problem in the table, extracts data from the table using the template model, and generates the corresponding question-and-answer pairs.

Citation Information

Cited By

  • Header identification and conversion method based on large model and RAG enhanced retrieval

    CN121636564A

  • A table header identification and conversion method based on a large model and RAG enhanced retrieval

    CN121636564B

  • Steel industry-oriented large model training data construction method and system

    CN122491431A