Header column identification method based on depth semantics and context self-adaption
Through the header column recognition method of deep semantics and context adaptive, the problem of poor header classification in the prior art is solved, and high accuracy recognition and adaptability of complex headers are achieved to meet the needs of high-precision data processing.
Patent Information
- Application Number
- CN202510686107.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The prior art has limited classification effects when dealing with complex, variable or similar table headers but different categories, especially in long-tail categories and fine-grained classification scenarios, and lacks deep understanding of table header semantics and full utilization of context information.
The header column recognition method based on deep semantics and context adaptability is adopted. The table analysis tool is used to obtain the header text, table body text and structure, and feature encoding and extraction are performed. The header features and scene context features are integrated, and the header semantic classification model is input to the trained header semantic classification model, and the final header semantic classification result is output.
It significantly improves the accuracy of table header recognition and the adaptability of the model, can automatically extract and understand the deep semantics of table headers, accurately judge the actual meaning of table headers under different tasks, and meet the needs of high-precision data processing and intelligent analysis.
Smart Images

Figure CN120197611A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, natural language processing, and intelligent analysis of structured data, and more specifically, to a method for identifying table header columns based on deep semantics and context adaptation. Background Art
[0002] In actual business, spreadsheets (such as Excel, CSV, etc.) are used as core data carriers and are widely applied in various data collection, statistics, and analysis processes. The header row (or multiple rows) of a table plays a key role in semantically describing the data content of each column. For example, fields such as "Sales Amount", "Customer ID", "Order Date", etc. are the basis for subsequent data parsing, classification, and analysis. However, the actual design and expression of table header columns are extremely diverse, mainly reflected in the following aspects:
[0003] 1. Structural complexity: The table header may contain complex structures such as multi-level nesting, merged cells, spanning rows and columns, etc., increasing the difficulty of parsing;
[0004] 2. Naming randomness: Table headers with the same business semantics may have multiple expression forms, including abbreviations, pinyin, internal codes, etc., lacking a unified standard;
[0005] 3. Semantic ambiguity: The same table header text may refer to different meanings in different business scenarios or contexts, and it is necessary to combine the table content for discrimination. For example, "Balance" may refer to either account balance or inventory balance; "Date" may refer to either order date or shipping date;
[0006] 4. Long-tail distribution phenomenon: In addition to common table headers, a large number of table headers belong to low-frequency, fine-grained "long-tail" categories, and it is difficult to comprehensively cover them with a limited number of samples.
[0007] Traditional table header identification methods mostly rely on artificial rules, keyword matching, or templates, lacking a deep understanding of the semantics of table headers. Although such methods are effective in specific formats, when faced with actual tables with variable structures, diverse naming, and semantic ambiguity (such as multi-level table headers, merged cells), their generalization ability and maintenance efficiency are relatively low, the recognition accuracy drops significantly, it is difficult to recognize the equivalence relationship between "Sales Amount" and "Sales Amount", and it is also impossible to distinguish the specific meaning of fields such as "Address" according to the context. In addition, the cost of maintaining and expanding the rule system is high, and it is difficult to adapt to the complex business requirements across different fields and industries, resulting in poor generalization ability and sustainability.
[0008] Early machine learning methods (such as classifiers based on feature engineering) and some deep learning models (such as basic CNN / RNN) have made progress in improving the automation level, but there are still obvious deficiencies in aspects such as deep semantic understanding, context dynamic modeling, and long-tail category discrimination, especially in the semantic modeling of table headers. These methods rely on shallow features, are difficult to capture the deep semantics and fine-grained differences of table header text, and cannot make full use of the language knowledge in large-scale unlabeled text, resulting in limited classification effects when dealing with complex, variable, or semantically similar but different-category table headers, especially performing poorly in long-tail category and fine-grained classification scenarios.
[0009] In recent years, pre-trained language models (such as BERT, etc.) have been introduced into the field of table header recognition, which have improved the semantic representation ability of table headers through methods such as sequence labeling or text classification. However, existing methods often focus on the table header text itself and fail to fully integrate the overall table content and context information, resulting in limitations when dealing with semantic ambiguity, complex structures, and long-tail categories. In addition, the lack of a targeted mechanism to dynamically adapt to different business scenarios and new table header structures affects the generalization ability and practical application effect of the model. For example, it is difficult to distinguish "order date" from "shipment date" based solely on "date", while product or logistics information in table data can provide key discriminative clues. Existing models lack a mechanism to specifically enhance the discriminative power of features when dealing with semantic ambiguity, long-tail categories, and fine-grained classification, and are prone to misclassifying rare or ambiguous table headers as common categories, resulting in a decrease in classification accuracy.
[0010] In addition, some techniques mix table header row positioning with table header column semantic classification, focusing on the boundary recognition of the table header area while ignoring the in-depth discrimination of the specific semantics of each column within the table header area. Although the multi-task or multi-agent learning paradigm can improve the generalization of the model to a certain extent, there are still obvious deficiencies in enhancing the distinguishability of the feature space, adapting to long-tail distributions, and complex business scenarios.
[0011] Therefore, how to deeply understand the semantics of table header columns, make full use of the context information of table content, improve robustness, so as to achieve new automated table header column recognition and classification for naming diversity, semantic ambiguity, and long-tail distribution, so as to meet the actual needs of high-precision data processing and intelligent analysis, especially the application requirements in high-standard scenarios such as evaluation and review, is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0012] In view of the above problems, the present invention provides a method for table header column recognition based on deep semantics and context adaptation to at least solve some of the technical problems mentioned in the above background technology.
[0013] To achieve the above object, the present invention adopts the following technical solutions:
[0014] A method for identifying table header columns based on deep semantics and context adaptation, comprising the following steps:
[0015] Use a table parsing tool to perform structured parsing on the target table file to obtain the target header text, the target table body text, and the target table structure corresponding to the target table file;
[0016] Perform feature encoding on the target header text to obtain target header features;
[0017] Perform feature extraction on the target table body text and the target table structure to obtain target scene context features;
[0018] Fuse the target header features and the target scene context features to generate target comprehensive features;
[0019] Input the target comprehensive features into a trained table header semantic classification model, output the probability distribution of each category corresponding to the target table file, and take the category with the highest probability as the final table header semantic classification result.
[0020] Further, the performing feature encoding on the target header text to obtain target header features specifically includes:
[0021] Perform word segmentation on the target header text, and after standardizing the segmented words, input them into a fine-tuned pre-trained language model to obtain target header features;
[0022] For multi-line or multi-level target header text, after obtaining the target header features of each target header text, perform embedding splicing or weighted fusion on the obtained multiple target header features to obtain the final target header features.
[0023] Further, the performing feature extraction on the target table body text and the target table structure to obtain target scene context features specifically includes:
[0024] Perform text extraction on the target table body text, and after converting the extracted target table content into a standardized Markdown format, input it as content features into a pre-trained open-source large language model text encoder to obtain target scene context features;
[0025] If there is target table structure information, encode the target table structure as additional features and input them together with the content features into a pre-trained open-source large language model text encoder to obtain target scene context features.
[0026] Further, the training steps of the table header semantic classification model include:
[0027] Construct a header semantic classification dataset; the header semantic classification dataset includes multiple header texts, and for each of the header texts, the corresponding table body text, table structure, and category label are included;
[0028] Perform feature encoding on each of the header texts to obtain header features;
[0029] Extract features from each of the table body texts and the corresponding table structures to obtain scene context features;
[0030] Fuse the corresponding header features and scene context features to generate comprehensive features;
[0031] Use the comprehensive features as input, and combine with the corresponding category labels to train a header semantic classification model.
[0032] Further, the steps for constructing the header semantic classification dataset include:
[0033] Perform structured parsing on the obtained multiple original table files to obtain the header area, table body text, and table structure of each original table file;
[0034] Locate the cells of the candidate headers from the header area, and perform preliminary classification on the candidate header texts therein;
[0035] Based on the table body text, verify the preliminary classification results of each candidate header text to obtain the category label of each candidate header text;
[0036] Based on the candidate header texts with existing category labels, perform sample clustering expansion to obtain the expanded table texts and their category labels;
[0037] Use all candidate header texts and expanded table texts as the final header texts, summarize all header texts and their category labels, as well as the corresponding table body texts and table structures, to form a header semantic classification dataset.
[0038] Further, the step of locating the cells of the candidate headers from the header area and performing preliminary classification on the candidate header texts therein specifically includes:
[0039] Based on the table structure information, locate the cells corresponding to the candidate headers from the header area;
[0040] After performing regularization processing on the candidate header texts in the candidate headers, use a preset domain knowledge base to perform preliminary classification on the candidate header texts;
[0041] The basis for the preliminary classification includes: if the candidate header text matches an entry in the preset domain knowledge base, directly use that entry as the category label for the candidate header text; if not, mark the category of the candidate header text as to be determined.
[0042] Further, based on the table body text, verify the preliminary classification results of each candidate header text to obtain the category label for each candidate header text; specifically including:
[0043] For the candidate header text marked as to be determined, combine the candidate header text and the table body text, extract the context text, and construct an input sample.
[0044] Input the input sample in a standardized format into a pre-trained large language model, and output classification suggestions and confidence scores.
[0045] If the confidence score is greater than or equal to the threshold, use the classification suggestion output by the large language model as the category label for the candidate header text to be determined.
[0046] If the confidence score is less than the threshold, add the input sample to the manual review queue for manual annotation as the category label for the candidate header text to be determined.
[0047] Further, based on the candidate header text with existing category labels, perform sample clustering expansion to obtain the expanded table text and its corresponding category labels, specifically including:
[0048] For the candidate header text with existing category labels, extract its text embedding vector.
[0049] Use a clustering algorithm to perform clustering on the text embedding vectors to obtain fine-grained categories and long-tail categories.
[0050] For the fine-grained categories or long-tail categories with the number of samples at the clustering center lower than the preset value, generate synonyms, near-synonym phrases or perform sample clustering expansion through data augmentation to obtain the expanded table text and its category labels.
[0051] Further, use the candidate header text and its category label as the original sample.
[0052] Use the expanded table text and its category label as the expanded sample.
[0053] Use the contrastive learning method to form positive and negative sample pairs from the expanded sample and the original sample, and guide the training of the header semantic classification model.
[0054] Further, the multi-objective loss function in the training process of the header semantic classification model includes a semantic classification loss function, a contrastive loss function, and a context consistency loss function; expressed as:
[0055]
[0056]
[0057]
[0058]
[0059] Among them, represents the multi-objective loss function; represents the semantic classification loss function; represents the contrastive loss function; represents the context consistency loss function; represents the weight of the semantic classification loss function; represents the weight of the contrastive loss function; represents the weight of the context consistency loss function; K represents a total of K categories; represents the true label of the k-th category; represents the predicted label probability of the k-th category; z represents the current sample feature, z + represents the positive sample feature, z − represents the negative sample feature, τ is the temperature parameter; N represents the number of sample pairs; and are respectively the predicted label probabilities of the table header text in different contexts in the i-th pair of samples; D KL represents the KL divergence, that is, the relative entropy.
[0060] According to the above technical solutions, compared with the prior art, the present invention discloses a method for identifying table header columns based on deep semantics and context adaptation, and has the following beneficial effects:
[0061] The present invention can not only automatically extract and understand the deep semantics of the table header, but also make full use of the context information of the table content and business scenarios to accurately determine the actual meaning of the table header under different tasks, thereby significantly improving the recognition accuracy and the adaptability of the model.
[0062] The present invention encodes the features of the table header text by fine-tuning the pre-trained deep network model, and can accurately capture the semantic information of the table header. The low-rank adaptive fine-tuning method effectively reduces the risk of overfitting, improves the model training efficiency, and supports efficient optimization for specific tasks with a small amount of labeled data, eliminating the cumbersome manual feature design.
[0063] The present invention integrates table content, structure, and business scenario information, and uses a large language model to perform comprehensive semantic analysis on table headers and their contexts, enabling the recognition of the polysemy of the same table header in different scenarios, ensuring the accuracy and consistency of classification results, and meeting the actual needs of multi-tasks and multi-domains.
[0064] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained according to the provided accompanying drawings without creative efforts.
[0066] Figure 1 It is a schematic flow diagram of a method for identifying table header columns based on deep semantics and context adaptation provided by an embodiment of the present invention.
[0067] Figure 2 It is a schematic flow diagram of constructing a table header semantic classification data set provided by an embodiment of the present invention.
[0068] Figure 3 It is a schematic diagram of the training steps of a table header semantic classification model provided by an embodiment of the present invention.
[0069] Figure 4 It is a schematic flow diagram of optimizing the loss function for multi-objective table header semantic classification provided by an embodiment of the present invention. Detailed Embodiments
[0070] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0071] The present invention aims to provide a method for identifying table header columns based on deep semantics and context adaptation, specifically to solve the core problems of low recognition accuracy, weak generalization ability, and insufficient processing ability for fine-grained and long-tail categories in the prior art in dealing with complex spreadsheet scenarios such as diverse formats, non-standard naming, semantic ambiguity, and strong context dependence. This method is for the intelligent parsing and semantic normalization of large-scale, heterogeneous table data, and is widely applicable to business scenarios with extremely high requirements for data structuring and semantic consistency, such as financial auditing, tax compliance, supply chain management, and enterprise data governance. See Figure 1As shown in the figure, it includes the following steps:
[0072] S1. Use a table parsing tool to perform structured parsing on the target table file to obtain the target header text, target table body text, and target table structure corresponding to the target table file;
[0073] S2. Perform feature encoding on the target header text to obtain the target header features;
[0074] S3. Extract features from the target table body text and target table structure to obtain the target scenario context features;
[0075] S4. Integrate the target header features and target scenario context features to generate the target comprehensive features;
[0076] S5. Input the target comprehensive features into the trained header semantic classification model, output the probability distribution of each category corresponding to the target table file, and take the category with the highest probability as the final header semantic classification result.
[0077] The above tags S1 - S5 are only for convenience of explanation and do not limit the execution order of each step.
[0078] This method supports full - process automation, without manual rule configuration, template maintenance, or manual proofreading, greatly improving the efficiency and consistency of header semantic understanding. It is especially suitable for scenarios that need to process a large amount of heterogeneous table data, such as financial auditing, tax compliance inspection, supply chain data analysis, enterprise - level data integration and governance, etc., and helps to support the standardized, intelligent management and high - quality analysis and decision - making of data assets. It not only significantly improves the intelligent level and business applicability of table header column semantic classification, but also provides a solid technical support for enterprises to achieve efficient and low - cost data governance and intelligent analysis. Next, each of the above steps will be described separately.
[0079] In the above step S1, use a table parsing tool to perform structured parsing on the target table file to obtain the target header text, target table body text, and target table structure corresponding to the target table file; specifically:
[0080] After using a table parsing tool (such as pandas, openpyxl, etc.) to perform structured parsing on the target table file, automatically identify the target header area, target data area, and target table structure (such as cell merging relationships); then, extract the target header text and target table body text from the target header area and target data area respectively.
[0081] In the above step S2, perform feature encoding on the target header text to obtain the target header features; specifically:
[0082] Perform word segmentation on the target header text. After standardizing the segmented words, input them into a fine-tuned pre-trained language model (such as the BERT model) to obtain target header features;
[0083] For multi-line or multi-level target header text, after obtaining the target header features of each target header text, perform embedding splicing or weighted fusion on the obtained multiple target header features to obtain the final target header features.
[0084] In step S3 above, perform feature extraction on the target table body text and the target table structure to obtain target scenario context features; specifically:
[0085] Extract text from the target table body text (such as the first few rows of data, table title, related descriptions), and after converting the extracted target table content into a standardized Markdown format, input it as content features into the text encoder of a pre-trained open-source large language model (such as the Qwen2.5 model) to obtain target scenario context features;
[0086] If there is target table structure information (such as grouping, merged cells), then encode the target table structure as additional features and input them together with the content features into the text encoder of the pre-trained open-source large language model to obtain target scenario context features.
[0087] In step S4 above, fuse the target header features and the target scenario context features through splicing, weighted average or attention mechanism to generate target comprehensive features;
[0088] In step S5 above, input the target comprehensive features into the trained header semantic classification model. In the header semantic classification model, output the probability distribution of each category corresponding to the target table file through a multi-layer perceptron or a softmax classification head, and take the category with the highest probability as the final header semantic classification result.
[0089] Next, the training process of the header semantic classification model will be described in detail. Specifically, the training process of the header semantic classification model includes a header semantic classification dataset construction stage and a header semantic classification model training stage.
[0090] 1. Header semantic classification dataset construction stage, see Figure 2 as shown:
[0091] In the embodiments of the present invention, it is necessary to construct a header semantic classification dataset for training the subsequent header semantic classification model; the constructed header semantic classification dataset includes multiple header texts, and includes the corresponding table body text, table structure and category label for each header text;
[0092] The specific construction process is as follows:
[0093] (1)Perform structured parsing on the obtained multiple original table files to obtain the header area, table body text, and table structure of each original table file; specifically:
[0094] ① Automatically batch collect original table files from multi-source business systems (such as finance, taxation, supply chain, etc.), supporting multiple formats such as Excel, CSV, and database exports;
[0095] ② Use table parsing tools (such as pandas, openpyxl, etc.) to perform structured parsing on the table files, automatically identify the header area, data area, and table structure (such as cell merging relationships); the table body text can be extracted from the data area;
[0096] ③ Standardize the parsing results into a unified data structure for subsequent processing.
[0097] (2)Locate the cells of the candidate headers in the header area and perform preliminary classification on the candidate header text therein; specifically:
[0098] ① Based on the table structure information (such as the first row, merged cells, bold font, etc.), automatically locate the cells corresponding to the candidate headers in the header area;
[0099] ② After performing regularization processing on the candidate header text (such as removing spaces, unifying case, removing special symbols), use a preset domain knowledge base to perform preliminary classification on the candidate header text;
[0100] The basis for preliminary classification includes: if the candidate header text matches an entry in the preset domain knowledge base, directly use that entry as the category label of the candidate header text; if it does not match, mark the category of the candidate header text as to be determined.
[0101] (3)Based on the table body text, verify the preliminary classification results of each candidate header text to obtain the category label of each candidate header text; specifically:
[0102] ① For the candidate header text marked as to be determined, combine the candidate header text and the table body text to extract context text (such as adjacent fields, table titles, partial data samples) and construct an input sample;
[0103] ② Input the input sample in a standardized format (such as JSON or Markdown) into a pre-trained large language model (such as the Qwen2.5 model) to output classification suggestions and confidence scores;
[0104] If the confidence score is greater than or equal to the threshold (e.g., 0.8), then use the classification suggestion output by the large language model as the category label of the candidate header text to be determined;
[0105] If the confidence score is less than the threshold, add the input sample to the manual review queue for manual annotation as the category label of the candidate header text to be determined.
[0106] ③ For multiple occurrences of the same header in different tables, count the consistency of the classification suggestions output by the large language model. If the consistency is lower than the set ratio (e.g., 80%), then automatically trigger the manual review process;
[0107] For example, for the header "Sales Amount", the system collects all the category labels assigned to it by the large language model. If it is statistically found that the combined frequency of the mainstream category labels (such as "Revenue" or "Total Sales Amount") is lower than the set threshold (e.g., 70%), then the system automatically marks "Sales Amount" as an item requiring manual review, and manual intervention is used to determine its actual meaning in each table.
[0108] Through the above preliminary classification, various types of header columns (such as "Order Amount", "Customer Name", "Shipping Date", etc.) can be accurately mapped to standardized business semantic categories or knowledge ontologies. Even in the face of expressions with the same meaning but different forms (such as "Sales Amount" and "Sales Amount") or words with multiple meanings (such as "Address", "Date", etc. which need to be judged in context), high-precision and fine-grained semantic classification can be achieved.
[0109] By deeply integrating the text information of the table body, dynamically resolving the semantic ambiguity of the header, ensuring that the classification result is highly consistent with the actual business context of the table, breaking through the limitation of relying solely on the isolated analysis of the header text, and significantly improving the adaptability of the model to complex business scenarios.
[0110] (4)Based on the candidate header text with existing category labels, perform sample clustering and expansion to obtain the expanded table text and its category labels;
[0111] ① For the candidate header text with existing category labels, extract its text embedding vector (which can be obtained using models such as BERT);
[0112] ② Use clustering algorithms (such as K-means or DBSCAN) to perform clustering on the text embedding vectors to obtain fine-grained categories (i.e., more specific or detailed semantic categories) and long-tail categories (i.e., header types that appear less frequently in the dataset);
[0113] ③For fine-grained or long-tail categories where the number of clustering center samples is lower than the preset value (e.g., 10), generate synonyms, near-synonym phrases, or achieve sample clustering expansion through data augmentation (such as spelling variants, context replacement) to obtain the augmented table text and its category labels;
[0114] Through this step, it is possible to effectively handle complex-structured tables such as multi-level headers and merged cells, as well as long-tail header columns with non-standard naming and extremely low frequencies, significantly reducing the risks of misclassification and missed classification and enhancing the stability and reliability of overall recognition.
[0115] (5)Use all candidate header texts and augmented table texts as the final header texts, summarize all header texts and their category labels, as well as the corresponding table body texts and table structures, to form a header semantic classification dataset; specifically:
[0116] ①Summarize all headers and their labels that have undergone automatic annotation, manual review, and clustering expansion to form a structured training dataset;
[0117] ②Deduplicate the dataset, detect outliers, and verify label consistency to ensure data quality;
[0118] ③Output a high-quality header semantic classification dataset that can be used for model training.
[0119] In summary, in the embodiments of the present invention, in response to the practical problems such as data scarcity, high annotation cost, and long-tail category distribution in the header semantic classification task, the present invention proposes an end-to-end automatic dataset construction and intelligent annotation system. First, the system automatically extracts candidate header columns from a large number of heterogeneous tables through a rule-guided weak supervision method and an active learning mechanism, and performs preliminary semantic classification in combination with a domain knowledge base. Subsequently, a large language model (LLM) is used to perform semantic consistency verification and automatic annotation on the headers and their context contents, significantly improving the annotation efficiency and accuracy. To further enhance the coverage of long-tail categories, the system introduces a sample expansion strategy based on clustering and contrast learning to automatically discover and supplement low-frequency and fine-grained category samples, ensuring the representativeness and diversity of the dataset. The entire dataset construction process is highly automated and can dynamically adapt to different business scenarios and table structures, providing a high-quality and low-bias data foundation for subsequent model training.
[0120] 2. The training stage of the header semantic classification model; see Figure 3 as shown:
[0121] (1)Extract multiple header texts from the constructed header semantic classification dataset, as well as the corresponding table body text, table structure, and category label for each header text;
[0122] (2)Perform feature encoding on each header text to obtain header features; specifically:
[0123] ①Perform word segmentation on the header text. After standardizing the segmented words, input them into a fine-tuned pre-trained language model (such as the BERT model) to obtain the header features;
[0124] ②For multi-line or multi-level header text, after obtaining the header features of each header text, perform embedding splicing or weighted fusion on the obtained multiple header features to obtain the final header features;
[0125] (3)Extract features from each table body text and the corresponding table structure to obtain the scenario context features; specifically:
[0126] ①Extract text from the table body text (such as the first few rows of data, table title, related descriptions), and after converting the extracted table content into a standardized Markdown format, use it as content features and input them into a pre-trained open-source large language model (such as the Qwen2.5 model) to obtain the scenario context features;
[0127] ②If there is table structure information (such as grouping, merged cells), encode the table structure as additional features and input them into the pre-trained open-source large language model together with the content features to obtain the scenario context features.
[0128] (4)Fuse the corresponding header features and scenario context features to generate comprehensive features; specifically:
[0129] ①Fuse the header features and scenario features through splicing, weighted average, or attention mechanism to generate comprehensive features;
[0130] ②If the attention mechanism is used, use the header as the query, the scenario features as the key and value, and calculate the attention weights to enhance the attention to the key context.
[0131] (5)Use the comprehensive features as input, combined with the corresponding class labels, to train the header semantic classification model; specifically:
[0132] ①Initialize the feature memory pool to store the representative header texts and their class labels for each category;
[0133] ②After each round of training, update the mean of the header features for each category in the current batch to the memory pool;
[0134] ③During the training process, compare the current sample features with the same / different category features in the memory pool for subsequent loss calculation.
[0135] In summary, in the embodiments of the present invention, a multi-level and cross-modal deep neural network architecture is adopted at the model level to fuse multi-source information such as header text, table structure, and table content, so as to achieve refined semantic modeling of header columns. The core model is based on a fine-tuned pre-trained language model (such as BERT, RoBERTa, etc.), combined with a table structure encoder and a context-aware module, which can dynamically capture the semantic association between the header and the main content of the table. To improve the adaptability of the model to complex structures and long-tail categories, the system introduces a feature memory pool mechanism to continuously store and retrieve historical feature-label pairs, realizing feature enhancement for low-frequency categories. At the same time, the model supports an end-to-end input-output process, automatically completing header localization, feature extraction, and semantic classification, greatly reducing manual intervention and improving the automation and generalization capabilities of the system.
[0136] In the embodiments of the present invention, a large language model is adopted to deeply analyze the context information of the table content, realizing semantic understanding in the header recognition process. By analyzing specific business scenarios (such as finance, taxation, etc.) through context information, the header content is accurately recognized to ensure high accuracy in the recognition process. Through this process, the system can better process complex table data, improving the accuracy and robustness of recognition.
[0137] In another embodiment, it further includes:
[0138] Taking the candidate header text and its category label as the original sample; taking the augmented table text and its category label as the augmented sample; using a contrastive learning method (such as SimCLR), forming positive and negative sample pairs with the augmented sample and the original sample, thereby guiding the training of the header semantic classification model. In another embodiment, a multi-objective header semantic classification loss function design method that takes into account the main task and long-tail adaptation is as Figure 4 shown. Specifically, the multi-objective loss function of the above header semantic classification model during training includes a semantic classification loss function, a contrastive loss function, and a context consistency loss function; specifically including:
[0139] 1. Semantic classification loss function:
[0140] For each training sample, the cross-entropy loss function is used as the semantic classification loss function, expressed as:
[0141]
[0142] Among them, represents the semantic classification loss function; K represents a total of K categories; represents the true label of the kth category; represents the predicted label probability of the kth category;
[0143] 2. Contrastive loss function (InfoNCE Loss):
[0144] For each training batch, sample positive sample pairs (the same category) and negative sample pairs (different categories), and calculate the cosine similarity between the feature vectors respectively;
[0145] (2) Calculate the contrastive loss function, expressed as:
[0146]
[0147] where, represents the contrastive loss function; z represents the feature of the current sample, z + represents the feature of the positive sample, z − represents the feature of the negative sample, and τ is the temperature parameter;
[0148] 3. Context consistency loss function:
[0149] (1) For multiple samples of the same table header in different contexts, calculate the class probability distributions of their model outputs respectively;
[0150] (2) Use KL divergence or mean squared error to constrain the output distributions of the same table header in different contexts to be consistent; expressed as:
[0151]
[0152] where, represents the context consistency loss function; N represents the number of sample pairs; and are the predicted label probabilities of the table header text in the i-th pair of samples in different contexts respectively; D KL represents KL divergence, that is, relative entropy, which is used to measure the distance between distributions.
[0153] 4. Joint optimization objective and model training:
[0154] (1) The multi-objective loss function is a multi-objective weighted combination; expressed as:
[0155]
[0156] where, represents the multi-objective loss function; represents the weight of the semantic classification loss function; represents the weight of the contrastive loss function; represents the weight of the context consistency loss function;
[0157] (2) Use an optimizer such as Adam, and perform backpropagation and parameter update based on the final loss function until the loss converges.
[0158] In summary, to address the problems of fine-grained discrimination and long-tail distribution in header semantic classification, the embodiments of the present invention innovatively design a loss function system for multi-objective joint optimization. In addition to the traditional cross-entropy loss, the system introduces a discriminative loss based on contrastive learning, which effectively improves the discriminative ability of the model in fine-grained and long-tail categories by maximizing the consistency of homogeneous header features and minimizing the similarity of heterogeneous header features. The feature memory pool and contrastive loss work together to enable the model to continuously optimize the feature space structure during training and enhance the generalization ability for new and rare headers. In addition, context consistency constraints are embedded in the loss function to ensure the semantic adaptability and robustness of the model in different business scenarios.
[0159] Next, to further illustrate the actual application effects of the present invention, the following two specific embodiments are given in combination with the model training and testing processes:
[0160] Embodiment 1: Training and testing of a header recognition model based on financial statements; specifically:
[0161] In this embodiment, first, about 5,000 financial forms are collected from public data sources such as annual reports and quarterly reports of listed companies, covering various types such as balance sheets, income statements, and cash flow statements. Through a combination of automated annotation tools and manual review, semantic labels such as "total assets", "total liabilities", and "operating income" are annotated for header cells to form a structured training header semantic dataset. The model uses a fine-tuned BERT large language model, inputs the header text and the content of its adjacent cells, and performs feature encoding and semantic classification. During training, cross-entropy loss and contrastive loss are jointly optimized to improve the recognition ability of mainstream and long-tail categories. After training, it is evaluated on an independent test set of 1,000 financial statements, and the header recognition accuracy of the model reaches 96.2%, which is significantly higher than 82.5% of the traditional rule-based method.
[0162] Embodiment 2: Training and testing of a header recognition model for supply chain management; specifically:
[0163] In this embodiment, 3,000 supply chain-related forms such as purchase orders, inventory tables, and shipping orders from 10 manufacturing enterprises are collected. Using an automated annotation system and combining with the enterprise business knowledge base, semantic annotations are made for header cells such as "material code", "supplier name", and "warehousing date". During the model training stage, a context awareness mechanism is integrated, and the header text, the content of adjacent cells, and table structure information are input into a fine-tuned RoBERTa model together. A multi-objective loss function is used during training, taking into account both the main task accuracy and the long-tail category discrimination ability. In the testing stage, header recognition is performed on 500 unseen supply chain forms, and the model accuracy reaches 93.8%, which can adapt to forms of different enterprises and different formats and accurately classify various business fields.
[0164] In summary, through automated dataset construction, a highly integrated deep model architecture, and an innovative multi-objective loss function design, the embodiments of the present invention systematically improve the automation, accuracy, and adaptability of table header column semantic classification, providing a solid technical foundation for the intelligent parsing of large-scale, heterogeneous tabular data and high-quality data governance.
[0165] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0166] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying table header columns based on deep semantics and context adaptation, characterized in that, It includes the following steps: Use a table parsing tool to perform structured parsing on the target table file to obtain the target header text, target table body text, and target table structure corresponding to the target table file; Perform feature encoding on the target header text to obtain target header features; Extract features from the target table body text and target table structure to obtain target scenario context features; Fuse the target header features and the target scenario context features to generate target comprehensive features; Input the target comprehensive features into the trained header semantic classification model, output the probability distribution of each category corresponding to the target table file, and take the category with the highest probability as the final header semantic classification result.
2. The method for identifying table header columns based on deep semantics and context adaptation according to claim 1, wherein The performing feature encoding on the target header text to obtain target header features; Specifically includes: Perform word segmentation on the target header text. After standardizing the segmented words, input them into the fine-tuned pre-trained language model to obtain target header features; For multi-line or multi-level target header text, after obtaining the target header features of each target header text, perform embedding splicing or weighted fusion on the obtained multiple target header features to obtain the final target header features.
3. The method for identifying table header columns based on deep semantics and context adaptation according to claim 1, wherein The extracting features from the target table body text and target table structure to obtain target scenario context features; Specifically includes: Extract text from the target table body text, and after converting the extracted target table content into a standardized Markdown format, input it into the pre-trained open-source large language model text encoder as content features to obtain target scenario context features; If there is target table structure information, encode the target table structure as additional features and input them into the pre-trained open-source large language model text encoder together with the content features to obtain target scenario context features.
4. The method for identifying table header columns based on deep semantics and context adaptation according to claim 1, characterized in that The training steps of the header semantic classification model include: Construct a header semantic classification dataset; the header semantic classification dataset includes multiple header texts, and includes the table body text, table structure, and category label corresponding to each header text; Perform feature encoding on each header text to obtain header features; Extract features from each table body text and the corresponding table structure to obtain scenario context features; Fuse the corresponding header features and scenario context features to generate comprehensive features; Use the comprehensive features as input, combined with the corresponding category labels, to train the header semantic classification model.
5. The method for identifying table header columns based on deep semantics and context adaptation according to claim 4, wherein The construction steps of the header semantic classification dataset include: Perform structured parsing on the obtained multiple original table files to obtain the header area, table body text, and table structure of each original table file; Locate the cells of the candidate headers in the header area and perform preliminary classification on the candidate header texts therein; Based on the table body text, verify the preliminary classification results of each candidate header text to obtain the category label of each candidate header text; Based on the candidate header texts with existing category labels, perform sample clustering expansion to obtain the expanded table text and its category label; Take all candidate header texts and extended table texts as the final header texts, summarize all header texts and their category labels, as well as the corresponding table body texts and table structures, to form a header semantic classification dataset.
6. The method for identifying table header columns based on deep semantics and context adaptation according to claim 5, wherein Locating the cells of the candidate headers in the header area and preliminarily classifying the candidate header texts therein specifically includes: Based on the table structure information, locate the cells corresponding to the candidate headers in the header area; After normalizing the candidate header texts in the candidate headers, use a preset domain knowledge base to preliminarily classify the candidate header texts; The basis for the preliminary classification includes: if the candidate header text matches an entry in the preset domain knowledge base, directly use the entry as the category label of the candidate header text; if it does not match, mark the category of the candidate header text as to be determined.
7. The method for identifying table header columns based on deep semantics and context adaptation according to claim 6, wherein Based on the table body text, verify the preliminary classification results of each candidate header text to obtain the category label of each candidate header text; specifically includes: For the candidate header texts marked as to be determined, combine the candidate header text and the table body text, extract the context text, and construct an input sample; Input the input sample in a standardized format into a pre-trained large language model, and output classification suggestions and confidence scores; If the confidence score is greater than or equal to the threshold, use the classification suggestion output by the large language model as the category label of the candidate header text to be determined; If the confidence score is less than the threshold, add the input sample to the manual review queue and be manually labeled as the category label of the candidate header text to be determined.
8. The method for identifying table header columns based on deep semantics and context adaptation according to claim 6, wherein Based on the candidate header texts with existing category labels, perform sample clustering expansion to obtain extended table texts and their corresponding category labels, specifically including: For the candidate header texts with existing category labels, extract their text embedding vectors; Use a clustering algorithm to cluster the text embedding vectors to obtain fine-grained categories and long-tail categories; For the fine-grained categories or long-tail categories with the number of samples at the clustering center lower than the preset value, generate synonyms, near-synonym phrases or perform sample clustering expansion through data augmentation to obtain extended table texts and their category labels.
9. The method for identifying header columns based on deep semantics and context adaptation according to claim 5, characterized in that: Take the candidate header texts and their category labels as the original samples; Take the extended table texts and their category labels as the extended samples; Using the contrast learning method, form positive and negative sample pairs with the extended samples and the original samples, and guide the training of the header semantic classification model.
10. The method for identifying table header columns based on deep semantics and context adaptation according to claim 9, wherein The multi-objective loss function in the training process of the header semantic classification model includes a semantic classification loss function, a contrast loss function, and a context consistency loss function; expressed as: ; ; ; ; Among them, represents the multi-objective loss function; represents the semantic classification loss function; represents the contrastive loss function; represents the context consistency loss function; represents the weight of the semantic classification loss function; represents the weight of the contrastive loss function; represents the weight of the context consistency loss function; K represents there are K categories in total; represents the true label of the k-th category; represents the predicted label probability of the k-th category; z represents the current sample feature, z + represents the positive sample feature, z − represents the negative sample feature, τ is the temperature parameter; N represents the number of sample pairs; and are respectively the predicted label probabilities of the table header text in the i-th pair of samples under different contexts; D KL represents the KL divergence, that is, the relative entropy.
Citation Information
Patent Citations
Header classification and header column semantic recognition method based on multi-task deep neural network
CN111523420A
Electric power field table column labeling method based on text classification
CN113486177A
Table information extraction method and device, equipment and medium
CN114818710A
Meter header identification method and device based on large language model, equipment and medium
CN118052213A
Cross-page table discrimination method based on double semantics
CN119202813A
Cited By
Commission settlement and reconciliation method and device for social e-commerce, computer equipment and storage medium
CN120707206A
Multi-agent collaborative reasoning method and system based on memory sharing and conflict resolution
CN121881288A
A multi-agent collaborative reasoning method and system based on memory sharing and conflict resolution
CN121881288B
Target file header and multi-file header similarity identification and associated data filling method based on large model
CN122065785A
Big model-based target file table header and multi-file table header similarity identification and associated data filling method
CN122065785B