Header field intelligent benchmarking method, system and device based on semantic index segmentation

By using a semantic index-based segmentation method, leveraging fine-tuning of large language models and BERT models, and combining vector library retrieval and regularity rule processing of header fields, the problem of low efficiency in header field standardization in existing technologies is solved, achieving efficient and accurate data standardization and intelligent benchmarking.

CN120910054AActive Publication Date: 2025-11-07ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Patent Information

Application Number
CN202511449009.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-07
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

In data governance, existing technologies suffer from inefficient standardization processes for header fields, and struggle to handle large-scale, complex, and heterogeneous data sources. In particular, when dealing with data items that are ambiguous, ambiguous, or synonymous, the matching accuracy is low, failing to meet the needs of modern enterprises for data standardization and intelligence.

Method used

We employ a semantic index-based segmentation method, which decomposes the header field into pairs of qualifiers and data elements using a large language model. By combining fine-tuning of the BERT model and vector library retrieval, we achieve automated field decomposition and accurate mapping. Furthermore, we process non-Chinese fields using regular expressions and dynamically update the thesaurus to adapt to different scenarios.

Benefits of technology

It significantly improves the efficiency and accuracy of data standardization, reduces manual intervention, enhances the accuracy of segmenting complex fields, is suitable for multilingual and multi-scenario data governance environments, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910054A_ABST
    Figure CN120910054A_ABST
Patent Text Reader

Abstract

The invention discloses a header field intelligent benchmarking method and system based on semantic index segmentation, relates to the technical field of data management, and aims to solve the problems that in the prior art, header field standardization depends on manpower, efficiency is low, and semantic fuzzy and non-Chinese fields are difficult to process. An intelligent benchmarking scheme fusing semantic comprehension and vector matching is provided. According to the method, header fields are decomposed through a large language model to generate an initial word bank, and a historical benchmarking database is constructed to realize rapid matching; semantic segmentation is carried out by adopting a fine tuning BERT model, and precise benchmarking is completed in combination with vector library retrieval and Top-K recommendation; code set recognition is achieved by applying regular rules and semantic logic for non-Chinese fields. The system comprises a lexicon construction module, a history matching module, a semantic segmentation module, a vector matching module and the like. The device comprises hardware units such as a processor and a memory. According to the invention, automatic standardized processing of header fields is realized, and the data governance efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data governance, and particularly relates to a table header field intelligent matching method and system based on semantic index segmentation. BACKGROUND

[0002] In the field of data governance, the standardization of table header fields is a key link to realize data integration, exchange and analysis. In the prior art, the matching of data elements and qualifiers mainly depends on manual work, and maintenance personnel need to compare and input data items and standard fields one by one. This method is not only low in efficiency, especially when facing large-scale data sources, but also prone to problems such as omission, error and the like due to the large number of fields or naming differences, which seriously affects the accuracy and efficiency of data standardization.

[0003] At present, part of the system attempts to realize the semi-automatic matching of data fields through a rule engine, but there are obvious limitations in actual application. When dealing with large-scale and complex heterogeneous data sources, these schemes often fail to achieve ideal intelligent matching results due to the lack of deep semantic understanding ability and context perception mechanism. Especially when dealing with data items with ambiguity, ambiguity or synonymous expressions, the matching accuracy of traditional methods decreases significantly. For fields lacking Chinese names or using abbreviations, existing technologies are difficult to effectively identify and process. These problems are particularly prominent in industries such as finance and medicine that require high data accuracy, seriously restricting the efficient circulation and value mining of data.

[0004] The root cause of the above problems lies in the fact that traditional methods cannot effectively analyze the semantic structure of table header fields. Data items are usually composed of qualifiers and data elements, but existing technologies lack the ability to automatically decompose complex fields into these two components. At the same time, the construction and maintenance of standard dictionaries also face challenges, and it is difficult to cover the term expression variants in different business scenarios. In addition, for the identification of special types such as code set fields, existing schemes mostly rely on fixed rule matching, which lacks flexibility and scalability. These technical defects have led to a long-term high-cost and low-efficiency state of data governance, which cannot meet the needs of modern enterprises for data standardization and intelligentization.

[0005] In actual application, these problems further manifest as long data integration period, large amount of manual correction work, and inconsistent standardization results. In particular, in the scenario of cross-system data migration or heterogeneous data source integration, the standardization of table header fields often becomes a key bottleneck for project implementation. Therefore, a new solution that combines semantic understanding, intelligent matching and historical data is urgently needed to improve the automation level and accuracy of data field standardization and meet the needs of modern data governance. SUMMARY

[0006] To solve the above technical problems, the application provides a table header field intelligent matching method and system based on semantic index segmentation to solve the problems existing in the prior art.

[0007] In a first aspect, to achieve the above object, the application provides a table header field intelligent matching method based on semantic index segmentation, comprising the following steps:

[0008] Collecting table header field data, decomposing it into limited words and data elements by a large language model to generate an initial word library;

[0009] Storing the initial word library in a relational database to build a historical matching module;

[0010] For Chinese table header fields, searching the historical database for matching standard terms based on the similarity of field names and table names;

[0011] If the historical matching fails, the field is divided into limited words and data elements by a semantic segmentation model, candidate combinations are retrieved from a vector library, and Top-K recommended results are output;

[0012] For non-Chinese fields, sample values are extracted and mapped to standard data elements by applying regular rules and semantic logic;

[0013] The data matched are stored in the historical database to update the word library.

[0014] Optionally, the process of generating the initial word library comprises:

[0015] Extracting table header field data and converting it into JSON format;

[0016] Inputting the large language model to decompose it into limited words and data element pairs;

[0017] High-quality term pairs are selected through semantic verification and stored in categories.

[0018] Optionally, the process of building the historical matching module comprises:

[0019] Storing the approved limited words and data element pairs in the relational database according to the field name and table name;

[0020] Creating full-text index and vector index for the field name, table name, limited word, and data element.

[0021] Optionally, the working process of the semantic segmentation model comprises:

[0022] Fine-tuning the BERT model to identify the boundaries of limited words and data elements in the field;

[0023] Outputting the segmentation results using the BIO tag system;

[0024] Combining the regular rule library to check edge cases.

[0025] Optionally, the process of retrieving candidate combinations in the vector library includes:

[0026] Encoding the split determiner and data element into vectors;

[0027] Retrieving similar candidates in the Faiss library and calculating cosine similarity;

[0028] Extending the candidate set by synonym replacement and generating Top-K recommendations.

[0029] Optionally, the mapping process of the non-Chinese field includes:

[0030] Extracting field sample values and normalizing processing;

[0031] Applying regular rules to match dates, numbers, and email formats in turn;

[0032] Performing semantic logic rule verification on unmatched samples.

[0033] In a second aspect, the present application also provides a table header field intelligent matching system based on semantic index segmentation, which is used to implement a table header field intelligent matching method based on semantic index segmentation. The system includes:

[0034] A word library construction module for collecting table header field data and decomposing it into determiner and data element pairs through a large language model to generate an initial word library;

[0035] A historical matching database module for storing standard term pairs in the initial word library and establishing full-text index and vector index for field name, table name, determiner, and data element;

[0036] A historical matching module for retrieving matched standard terms from the historical database according to the similarity of field name and table name;

[0037] A semantic segmentation and vector matching module for segmenting the field into determiner and data element through a fine-tuned BERT model and retrieving candidate combinations in the vector library to output Top-K recommendation results;

[0038] A non-Chinese field processing module for extracting non-Chinese field sample values, applying regular rules and semantic logic to map them into standard data elements;

[0039] A dynamic updating module for storing the completed matching data into the historical database to update the word library.

[0040] In a third aspect, the present application also provides a table header field intelligent matching device based on semantic index segmentation, which is used to implement a table header field intelligent matching method based on semantic index segmentation. The device includes:

[0041] A processor is configured to execute the steps of the method.

[0042] A memory is configured to store an initial vocabulary, a historical matching database, and a vector index.

[0043] An input / output interface is configured to receive table header field data and output a matching result.

[0044] A communication module is configured to interact with external data sources and a governance platform to obtain and update data.

[0045] In a fourth aspect, the present application further provides a computer terminal device, comprising:

[0046] One or more processors;

[0047] A memory coupled to the processor and configured to store one or more programs;

[0048] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method for intelligent matching of table header fields based on semantic index segmentation in the first aspect.

[0049] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method for intelligent matching of table header fields based on semantic index segmentation in the first aspect.

[0050] Compared with the prior art, the present application has the following advantages and technical effects:

[0051] The present application provides a method and system for intelligent matching of table header fields based on semantic index segmentation. The present application realizes automatic decomposition and accurate matching of table header fields by combining a semantic segmentation model with vector matching technology. The fast retrieval based on a historical database reduces the computational overhead of repeated matching, and the fine-tuning of the BERT model improves the segmentation accuracy of complex fields. The regular rules and semantic logic verification for non-Chinese fields effectively solve the identification problem of code set fields. The dynamically updated vocabulary mechanism ensures the adaptability of the system to newly added terms. The overall scheme significantly improves the efficiency of data standardization while reducing the need for manual intervention, and is suitable for multi-language and multi-scenario data governance environments. BRIEF DESCRIPTION OF DRAWINGS

[0052] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments thereof and their descriptions are used to explain the present application and do not constitute improper limitations on the present application. In the drawings:

[0053] Figure 1 The flowchart of the steps of the embodiments of the present application;

[0054] Figure 2 Structure diagram of an embodiment of the present application;

[0055] Figure 3 System structure diagram of an embodiment of the present application;

[0056] Figure 4 Device structure diagram of an embodiment of the present application. DETAILED DESCRIPTION

[0057] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0058] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0059] Embodiment one

[0060] As shown in the accompanying drawings, Figure 1 The present embodiment provides a table header field intelligent matching method based on semantic index segmentation, which includes:

[0061] Collecting table header field data, decomposing it into limited words and data element pairs through a large language model to generate an initial word library;

[0062] Storing the initial word library in a relational database to build a historical matching module;

[0063] For Chinese table header fields, based on the similarity of field names and table names, search the historical database for matching standard terms;

[0064] If the historical matching fails, use a semantic segmentation model to segment the field into limited words and data elements, retrieve candidate combinations from a vector library, and output Top-K recommended results;

[0065] For non-Chinese fields, extract sample values and apply regular rules and semantic logic to map them to standard data elements;

[0066] Store the matched data in the historical database to update the word library, and the above content corresponds to the structure diagram Figure 2 .

[0067] Specifically, the following steps are included:

[0068] S1: Collect typical table header data, and use a large language model to decompose each table header field data into limited words and data element data pairs, select data that conforms to the semantics as initial limited words and data elements, and form an available initial word library.

[0069] S2: Using the initial vocabulary generated in S1, arrange the approved qualifiers and data elements into a standard format, store them in a relational database, and create indexes for commonly used search fields to optimize query efficiency. This database will serve as the basis data for the historical benchmarking module.

[0070] S3: For Chinese table header fields, quickly retrieve the most matching qualifiers and data elements from the historical benchmarking database based on field name and table name similarity, and realize the quick identification and benchmarking of repeated fields.

[0071] S4: When matching cannot be completed through historical benchmarking data, the system will enter the "full-space benchmarking module". The core of this module is a downstream semantic segmentation model built based on existing benchmarking data. By fine-tuning the pre-trained BERT model, it adapts to the structured task of "field qualifier and data element segmentation". After segmentation, candidate items are retrieved in the vector library and replaced by synonyms, and finally KxK candidate combinations are sorted to output Top-K recommended results. For all data that completes full-space benchmarking, store it in the historical benchmarking database.

[0072] S5: For non-Chinese fields, the system will enter the "other content benchmarking module". This module targets fields without Chinese names, first performs segmentation and vector matching to locate candidate data elements, then samples 10 field values to calculate cosine similarity, and judges and maps to code set data elements. For all data that completes other content benchmarking, store it in the historical benchmarking database.

[0073] As an implementation in this embodiment, the process of generating an initial vocabulary includes:

[0074] Extract table header field data and convert it to JSON format;

[0075] Input large language model to decompose into qualifier and data element pairs;

[0076] Filter high-quality term pairs through semantic verification and store them in categories.

[0077] Specifically, S1 includes:

[0078] S1.1: Systematically extract table header field data from the target data source. Ensure that the collected table header data is representative and covers multiple scenarios. During the collection process, perform preliminary cleaning on the table header data to remove duplicate fields, blank fields, or obviously invalid fields. Finally, organize a structured table header field data set, such as in list or table form, for subsequent processing.

[0079] S1.2: Convert the collected table header field dataset into a JSON structure. For each table header field, supplement the necessary context information to help the model better understand the semantics of the field. Check the consistency of the data format to ensure that there are no coding errors, missing values, or format abnormalities, etc. After completing the preprocessing, generate a well-formatted table header field dataset to ensure that it can be directly input into a large language model for analysis.

[0080] S1.3: Input the preprocessed table header field dataset into a large language model, and configure the model to perform semantic analysis tasks. The task of the model is to decompose each table header field into a pair of qualifiers and data elements. In the model configuration, it is explicitly required to output structured decomposition results in the form of JSON format key-value pairs. For complex or ambiguous table header fields, the model should attempt to infer a reasonable decomposition method based on the context. Finally, generate the preliminary qualifiers and data elements corresponding to each table header field.

[0081] S1.4: Perform semantic verification on the decomposition results output by the large language model to ensure that each qualifier and data element accurately reflects the semantics of the table header field. For decomposition results that are semantically inaccurate, unreasonable, or do not conform to the context, they are filtered out through automatic rules. Retain high-quality qualifiers and data elements to form a filtered decomposition result list that is semantically correct and usable.

[0082] S1.5: Organize the filtered qualifiers and data elements into an initial vocabulary and construct a structured term set. Classify and organize the qualifiers and data elements by data type, industry domain, or semantic function to improve the retrieval and use efficiency of the vocabulary. Add meta-information to each term, including source table, semantic description, and applicable scenarios, to support subsequent vocabulary expansion and maintenance work.

[0083] As an optional implementation in this embodiment, the large language model mentioned above can use a deep search large language model.

[0084] As an embodiment of the present embodiment, the construction process of the historical benchmarking module includes:

[0085] Store the approved qualifiers and data element pairs by field name and table name in a relational database;

[0086] Create full-text indexes and vector indexes for field names, table names, qualifiers, and data elements.

[0087] Specifically, S2 includes:

[0088] S2.1: Extract the approved qualifiers and data elements from the initial vocabulary generated in S1, ensuring that each entry contains the field name, qualifier, data element, and their meta-information. Format standardize these entries into a unified JSON format, preparing for storage in the historical benchmarking database.

[0089] S2.2: Store the formatted qualifiers and data elements into a relational database according to attributes such as field name, table name, qualifier ID, data element ID, etc. Ensure that the database table structure supports efficient storage and querying, including all necessary fields and their meta-information. Verify the integrity of the storage process to ensure no data loss or format errors.

[0090] S2.3: Create full-text indexes and vector indexes for commonly used search fields in the historical benchmarking database to improve query efficiency. Use vector indexes to support semantic similarity matching, optimizing the retrieval performance of the subsequent automatic matching module. Test the effectiveness of the indexes to ensure fast response to diverse query requirements.

[0091] S3: If the benchmarking field is not in Chinese, execute S5 Other Content Benchmarking Module, otherwise execute "Historical Benchmarking Module": From the historical full-space benchmarking data and benchmarking database, quickly retrieve the most matching qualifiers and data elements based on field name and table name similarity, achieving fast identification and benchmarking of duplicate fields.

[0092] As an embodiment in this embodiment, the working process of the semantic segmentation model includes:

[0093] Fine-tuning the BERT model to identify the boundaries of qualifiers and data elements of fields;

[0094] Using the BIO tag system to output segmentation results;

[0095] Combining regular rule library to check edge cases.

[0096] Specifically, the S4 includes:

[0097] S4.1: Extract part of the data from the constructed historical benchmarking database to construct semantic index segmentation training data set. Construct [data field, segmentation index] training data, where "data field" is the data field to be segmented, and "segmentation index" is the index position that can accurately segment the data field into "qualifier" and "data element". Use pre-trained large language models, including but not limited to BERT, to fine-tune downstream tasks, so that the BERT model has the ability to recognize semantic segmentation of data fields.

[0098] S4.2: The fine-tuned BERT semantic segmentation model is used to perform sequence labeling on the input data items, automatically segmenting the original fields into two parts: "qualifiers" and "data elements". The model outputs labeled sequences using the BIO tagging system, which greatly improves the segmentation accuracy of complex and long fields.

[0099] S4.3: To address potential edge cases in BERT segmentation, secondary validation and correction are performed using a maintained regular expression rule library. For example, regular expression rules are used to fine-tune tags for common formats such as dates, currencies, and units of measurement.

[0100] As one implementation method in this embodiment, the process of retrieving candidate combinations from the vector library includes:

[0101] Encode the segmented qualifiers and data elements into vectors;

[0102] Search for similar candidates in the Faiss database and calculate the cosine similarity.

[0103] Expand the candidate set by synonym replacement and generate Top-K recommendations.

[0104] Specifically, the detailed steps include:

[0105] S4.4: The segmented qualifying words and data elements are mapped into vectors using the m3e encoder and stored in the Faiss vector library. For each input segment, a nearest neighbor search is performed to retrieve the sets of the most similar qualifying word vectors. and data element vector set Similarity is calculated using cosine similarity:

[0106] ;

[0107] in, For a limited set of word vectors One of the elements, For data element vector set One of the elements;

[0108] S4.5: For the retrieved data element candidates, further utilize rule matching (such as root and prefix / suffix rules) and synonym clustering in the vector space to replace or supplement the synonym list, thereby enhancing the coverage and robustness of the data element representation.

[0109] S4.6: Take the first part of the qualifier and the data element respectively. Each candidate is paired with another candidate to generate a total of [number] candidates. For each pair of combinations (composite items), calculate the overall similarity score with the original items, sort by score, and output the Top-ranked items with the highest similarity. The combination is recommended as the final benchmark, and the data fields, qualifiers, data elements, and metadata of the benchmark will be stored in the historical benchmark database.

[0110] As an embodiment in this embodiment, the mapping process of the non-Chinese field includes:

[0111] Extract field sample values and normalize them;

[0112] Apply regular rules to match dates, numbers, and email formats in sequence;

[0113] Perform semantic logic rule verification on unmatched samples.

[0114] Specifically, the S5 includes:

[0115] S5.1: For fields without Chinese names, randomly extract non-empty records, remove obvious outliers (such as entries with excessively short or long lengths), and perform uniform blank and case normalization on each sample to obtain a sample set to be identified .

[0116] S5.2: For the sample set, apply predefined regular rules in sequence:

[0117] Date (such as \d{4}-\d{2}-\d{2}, \d{2} / \d{2} / \d{4});

[0118] Email (such as \b[\w.%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b);

[0119] Mobile phone number (such as ^1[3-9]\d{9}$);

[0120] ID number (such as \d{17}[0-9Xx]);

[0121] For each sample record, record its matched data element type and matching proportion matching sample number . If the matching proportion of a certain category is , it is directly mapped to the data element.

[0122] S5.3: For samples not fully covered by regular expressions, apply semantic logic rules in sequence:

[0123] Name library comparison: fuzzy match the sample with a common Chinese name library or English name library;

[0124] Gender / dictionary: check if the sample value hits "male", "female", or the name of each nationality;

[0125] Region / area mapping: compare the sample with a list of region and area names;

[0126] Other enumerated fields are determined by rule table or small mapping table.

[0127] S5.4: For each candidate data element type , calculate its "regular matching score" and "semantic rule score", and calculate the comprehensive score in a weighted manner , the comprehensive score is , the weight coefficient is , the regular matching score calculation function is , the semantic rule score calculation function is, both of which are calculated by vector cosine similarity matching. Select the candidate with the highest score and as the final mapping, otherwise mark it as "unidentified". In addition, the completed data field, qualifier, data element and metadata are stored in the historical mapping database. More specific implementation process includes:

[0128] S1.1: From multiple common business systems, including customer management, order processing, financial accounting and product configuration modules, through metadata scanning and structured export, typical table header fields are systematically extracted to ensure coverage of different business scenarios and naming styles. The extracted fields are preliminarily cleaned to remove duplicate, blank or invalid fields, and the field names, table names and sample values are unified and sorted to form a structured table header field dataset, meeting the needs of semantic diversity and coverage.

[0129] S1.2: Convert the table header field dataset to JSON format and supplement context information, including field value examples and system sources, to enhance the semantic understanding ability of large language models. The system automatically performs semantic segmentation and named entity recognition. Check the data format consistency to ensure no encoding errors or missing values, and generate a format specification field dataset to adapt to model input.

[0130] S1.3: Input the preprocessed field dataset into the large language model, configure the model to perform semantic analysis tasks, and automatically decompose each table header field into qualifier and data element pairs. The model infers the decomposition method of complex or ambiguous fields based on the context, and outputs structured JSON format key-value pairs (such as {"qualifier": "registered person", "data element": "name"}). Through semantic similarity calculation, filter the decomposition results that meet the field semantics and eliminate unreasonable pairs to form a candidate term set.

[0131]

[0132] ​S1.4: Organize the filtered qualifying words and data elements into an initial thesaurus, categorize them by data type or business scenario, and add metadata. Through automated semantic validation rules, standardize terminology granularity and naming conventions, and handle synonymous or polysemous cases. The thesaurus is stored in JSON or database format, generating the first standardized thesaurus.

[0133] S2.1: Extract approved qualifiers and data elements from the initial thesaurus generated in S1.4. For example, the header field "Cumulative Order Amount" is broken down in S1 into {"Qualifier": "Cumulative", "Data Element": "Order Amount"}, along with metadata (e.g., Source Table: Order Table, Semantic Description: Indicates Total Order Amount). The system organizes these term pairs into a unified JSON structure, for example: { "field_name": "Cumulative Order Amount", "table_name": "Order Table", "qualifier": "Cumulative", "data_element": "Order Amount", "metadata": { "source": "Order Management System", "description": "Total Order Amount"}}. Similarly, other fields such as "Registration Time" are processed {"Qualifier": "Registration", "Data Element": "Time"}, ensuring all term pairs have a consistent format for database storage.

[0134] S2.2: Import the organized terminology into a relational database. The table structure includes columns for field names, table names, qualifier IDs, data element IDs, and metadata. Taking "Cumulative Order Amount" as an example, the system generates a record with the following fields: Field Name: Cumulative Order Amount; Table Name: Order Table; Qualifier ID: Q001 (corresponding to "Cumulative"); Data Element ID: D001 (corresponding to "Order Amount"); Metadata: {"source": "Order Management System", "description": "Total Order Amount"}. Similarly, a record is generated for "Registration Time" (Qualifier ID: Q002, Data Element ID: D002). The system automatically verifies the stored procedure, checks the record integrity, ensures no data loss or format errors, and forms an initial historical benchmark database.

[0135] S2.3: Create full-text and vector indexes for field names, table names, qualifiers, and data elements in the database. Full-text indexes support fast keyword matching; for example, entering "order amount" will retrieve records for "cumulative order amount." Vector indexes are based on a semantic embedding model, embedding "order amount" as a vector, and support semantic similarity queries, such as matching the synonymous field "total order amount." The system tests index performance to ensure that queries for "cumulative order amount" return results in milliseconds, optimizing the retrieval efficiency of the subsequent automatic matching module.

[0136] S3.1: The system receives the new Chinese header field "Total Sales Amount" and extracts the table name information: Sales Record Table. Through automated preprocessing, the fields are converted to standard JSON format.

[0137] S3.2: Retrieve records from the historical benchmarking database that exactly match "Total Sales Amount". The system uses full-text indexing to precisely match the field name "Total Sales Amount" and the table name "Sales Record Table". Assuming the database contains the record stored in S2: {"field_name": "Total Sales Amount", "table_name": "Sales Record Table", "qualifier": "Total", "data element": "Sales Amount", "qualifier ID": "Q003", "data element ID": "D004"}, the system confirms an exact match. If no exact match is found for the field name or table name, the system skips subsequent steps, and the record is marked as unmatched in the log.

[0138] S3.3: For records with an exact match, the system directly uses the qualifier and data element results from the database. For example, after a record is matched for "total sales amount", the output is: qualifier: total (ID: Q003); data element: sales amount (ID: D004). If no matching record is found, the system does not generate a new term, skips the module, and logs the data to track unmatched fields.

[0139] If the field name is in English or an abbreviation, such as "reg_dt" or "pay_amt", the system skips the Chinese matching process and directly proceeds to the "Other Content Matching Module" (i.e., the subsequent S5 process). This judgment process is automatically completed by the field language feature detection model to achieve dynamic traffic separation between Chinese and non-Chinese paths, ensuring that fields of different languages ​​can enter the most suitable matching module.

[0140] S4.1: Extract partial data from the historical benchmark database constructed in S2, such as the record "Cumulative Order Amount": {"Qualifier": "Cumulative", "Data Elements": "Order Amount"}. The system constructs a training dataset in the format [data field, segmentation index], where "data field" is "Cumulative Order Amount" and "segmentation index" is [0, 2, 4] (representing "Cumulative" [0:2], "Order Amount" [2:4]). Other fields (such as "Registration Time": [0, 2, 4]) are processed in a similar manner to generate a training dataset containing 10,000 records. A pre-trained BERT model is fine-tuned and optimized for sequence labeling tasks, enabling the model to learn to accurately segment the Chinese header field into qualifyers and data elements.

[0141] S4.2: The fine-tuned BERT model uses the BIO tag system (B-Modifier, I-Modifier, B-Data Element, I-Data Element, O-Other) for sequence labeling. The model is fine-tuned based on the training dataset constructed in S4.1 (containing 10,000 [data field, split index] records, e.g. "cumulative order amount" corresponds to index [0, 2, 4]) to optimize parameters to identify the semantic boundaries of Chinese table header fields. For "cumulative order amount", the model predicts the label character by character, outputting the following sequence: cumulative: B-Modifier (indicating the start of the modifier); count: I-Modifier (indicating the continuation of the modifier); order: B-Data Element (indicating the start of the data element); single: I-Data Element; gold: I-Data Element; amount: I-Data Element. The model improves semantic understanding of complex fields through a multi-layer Transformer structure, combining context information (such as "order table" implying a sales scenario), to ensure accurate splitting. The system splits "cumulative order amount" into two parts according to the BIO tag sequence: modifier: "cumulative" (composed of B-Modifier and I-Modifier, character position [0:2]); data element: "order amount" (composed of B-Data Element and I-Data Element, character position [2:6]). The splitting result is output in JSON format.

[0142] S4.3: For potential edge cases of BERT splitting, the system applies a regular expression rule library for secondary verification. For example, for "order amount 2023" which may be mistakenly split into "order" and "amount 2023", the regular rule (such as "\d{4}$" to identify the year) corrects it to "order amount" and "2023" (excluding the date part). For "cumulative order amount", the regular rule confirms that there is no special format (such as date, currency), and the BERT splitting result is retained: {"modifier": "cumulative", "data element": "order amount"}.

[0143] S4.4: The splitting results "cumulative" and "order amount" are mapped to 768-dimensional vectors through the m3e encoder and stored in the Faiss vector library. The Faiss library has pre-stored modifier (such as "cumulative" "total") and data element (such as "order amount" "sales amount") vectors of the historical pair database. For "cumulative", perform nearest neighbor search to retrieve Top-3 modifier vectors (such as "cumulative" "total" "total"); for "order amount", retrieve Top-3 data element vectors (such as "order amount" "sales amount" "transaction amount"). Similarity is calculated using cosine, for example, the similarity between "cumulative" and "total" is 0.93.

[0144] S4.5: The system combines the top K candidates (K=3) of the modifier and data element in pairs to generate KxK=9 synthetic items, and outputs the Top-K recommendation results through comprehensive similarity calculation. From the modifier set and data element set The first three candidates are selected from each, generating nine combinations. For each combination, the system calculates a comprehensive similarity score, with the formula: wherein, is the comprehensive similarity score, is the candidate qualifier vector, is the original data item vector, is the candidate data element vector, and are the cosine similarities of the qualifier and data element, respectively, with the formula: , Similarly; Based on the table name (such as "order table") and the historical benchmark frequency, the business rationality of the combination is evaluated, for example, the "order table" context improves the score of the "order" related combination. The weight is set to , which can be adjusted according to the general field scenario. The system ranks the nine combinations from high to low according to the score, and selects Top-1 (K=1) as the recommended result { "qualifier": "cumulative", "data element": "order amount"} with a similarity of 0.98. The benchmarking result (field name: cumulative order amount, qualifier: cumulative, data element: order amount, metadata: { "table_name": "order table", "sample_value": "123456.78"} is stored in the historical benchmarking database, and the Faiss vector library is updated. Test 100 fields, Top-1 accuracy rate reaches 93%.

[0145] S5.1: The system first performs data preprocessing on the field "OrderID" to generate a reliable sample set for subsequent analysis. Randomly select 10 non-empty records from the "OrderID" field, such as "ORD12345", "2023-ORD-001", "ORD98765", etc. Exclude obvious outliers (such as null values or records containing only one character). The system normalizes each sample, including removing extra spaces, unifying case (e.g. converting "ord12345" to "ORD12345") and standardizing character encoding to UTF-8 to ensure data consistency. The normalized sample set retains 10 valid records, such as "ORD12345", "2023-ORD-001", "ORD98765", "ORD45678", etc. These records reflect the typical format of order numbers, which may contain letters, numbers and separators (such as "-"). The preprocessing process ensures that the sample set can represent the true data distribution of the field, providing a reliable foundation for subsequent regular rules and semantic analysis. The system records the original value and normalized value of each sample, and generates a sample set to be identified for the next step of rule matching.

[0146] S5.2: The system applies a predefined regular expression rule library to the normalized sample set sequentially to identify the data element type of the field "OrderID". The regular rule library is designed for general domains and contains various common formats, such as date (e.g., `\d{4}-\d{2}-\d{2}` or `\d{2} / \d{2} / \d{4}`), mobile phone number (e.g., `^1[3-9]\d{9}$`), ID number (e.g., `\d{17}[0-9Xx]`), and number format (e.g., `[A-Za-z0-9\-]+`). For the "OrderID" sample set, the system matches each record one by one and finds that all samples (e.g., "ORD12345" "2023-ORD-001") conform to the number format regular rule `[A-Za-z0-9\-]+`, which allows combinations of letters, numbers, and hyphens. The matching result shows that 10 out of 10 samples completely conform to the number format, with a matching ratio of 100%. The system records the matching type (number) and matching ratio of each sample and confirms that the field "OrderID" is highly likely to be a number type data element. Since the matching ratio reaches 100%, the system initially maps the field as the standard data element "order number", but still needs to be further verified by subsequent semantic logic rules to ensure accuracy.

[0147] S5.3: For fields not completely covered by regular rules or cases that need further verification, the system applies semantic logic rules to analyze the sample set to exclude other possible semantic types and confirm the mapping result. Although the regular matching ratio of the "OrderID" field has reached 100%, to ensure robustness, the system still performs semantic logic rule checks. Semantic rules include name library comparison (checking if the sample matches common Chinese or English names), gender / ethnicity dictionary matching (checking if it contains "male" "female" or ethnic names), regional / area mapping (comparing with regional / area name lists), and rule table matching for other enumerated fields. For the "OrderID" sample set, the system first compares the common name library (including Chinese names such as "Zhang Wei" and English names such as "John Smith") and finds that samples such as "ORD12345" do not match any name pattern. Gender / ethnicity dictionary checks confirm that samples do not contain "male" "female" or ethnic names. Regional / area mapping also has no matching results because samples do not involve geographic names. The system further checks the mapping table of enumerated fields (such as status fields such as "Active" "Pending"), but the letter-number combination of "OrderID" samples does not conform to the enumerated type characteristics. Since the samples highly conform to the number format and do not hit other semantic rules, the system confirms that "OrderID" is most likely an order number type data element and does not need further expansion analysis.

[0148] S5.4: The system calculates a comprehensive score for the candidate data element type to determine the final mapping result. For the "OrderID" field, the candidate data element type is "order number" based on the results of regular matching and semantic rule analysis. The system calculates a "regular matching score" defined as the proportion of matching samples, i.e. , where is the number of samples matching the regular rule, is the total number of samples. Here, , , . Since the regular matching proportion reaches 100%, there is no need to calculate the semantic rule score (no other candidate types are found by the semantic rule). The comprehensive score formula is: , where the weight , the weight , and (no other semantic types). Therefore, the comprehensive score is: . Since the score exceeds the threshold value 0.8, the system maps "OrderID" to the standard data element "order number". If the comprehensive score is below the threshold value or there are multiple candidate types, the system will be marked as "unidentified". In this example, "order number" is the final mapping result, with a mapping note recording the correspondence between the original field and the standard term.

[0149] Based on this, the embodiment of the present application provides an intelligent standardization method for table header fields based on semantic index segmentation, which significantly improves the efficiency, accuracy and adaptability of field standardization in data governance compared with the prior art. Through the collaborative work of the semantic index segmentation model based on BERT fine-tuning, semantic vector matching and the historical standardization database, the system realizes the automatic decomposition and accurate matching of table header fields, greatly reduces manual intervention and reduces maintenance costs. At the same time, the system effectively handles semantic ambiguity, non-Chinese and code set fields, overcomes the limitations of traditional methods on complex fields, and enhances the adaptation ability in multiple scenarios. The accuracy and recall rate of the qualifier matching are both more than 90% on the top 3 standard qualifiers recommended automatically; the accuracy and recall rate of the data element matching are both more than 95% on the top 3 standard data elements recommended automatically, further verifying the superior performance of the system. In addition, the modular design and dynamic vocabulary update further optimize the query performance and system scalability, promote data standardization and efficient circulation, provide strong support for data integration and analysis, and show significant creativity and practical value.

[0150] Embodiment Two

[0151] In this embodiment, as shown in Figure 4 , an intelligent standardization device for table header fields based on semantic index segmentation is provided, comprising:

[0152] A processor configured to perform the steps of the above-described method for intelligent matching of table header fields based on semantic index segmentation.

[0153] A memory configured to store an initial vocabulary, a historical matching database, and a vector index.

[0154] In addition to the processor and the memory, the system further comprises:

[0155] An input / output interface configured to receive table header field data and output matching results.

[0156] A communication module configured to interact with external data sources and governance platforms to obtain and update data.

[0157] In this embodiment, a computer terminal device is provided, comprising:

[0158] One or more processors;

[0159] A memory coupled to the processor and configured to store one or more programs;

[0160] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-described method for intelligent matching of table header fields based on semantic index segmentation.

[0161] In this embodiment, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-described method for intelligent matching of table header fields based on semantic index segmentation are implemented.

[0162] In this embodiment, an electronic device is also provided, which comprises a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the above-described method for intelligent matching of table header fields based on semantic index segmentation.

[0163] In this embodiment, a computer program product is also provided, which comprises a computer program. When the computer program is executed by a processor, the steps of the above-described method for intelligent matching of table header fields based on semantic index segmentation are implemented.

[0164] The above program can run in a processor, or can also be stored in a memory (or called a computer readable medium), the computer readable medium includes permanent and non-permanent, movable and non-movable media can be realized by any method or technology information storage. Information can be computer readable instructions, data structure, program module or other data. Examples of computer storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage device or any other non-transmission medium, which can be used to store information that can be accessed by a computing device.

[0165] These computer programs can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the flow Figure 1 The flow or multiple flows and / or the function of the block Figure 1 The steps of the function of the block or multiple blocks are specified, and different steps can be realized by different modules.

[0166] Such a device or system is provided in the embodiment. The system is called semantic index segmentation based table header field intelligent matching system, which includes:

[0167] The lexicon construction module is used to collect table header field data and decompose it into limited words and data element pairs through large language model, and generate initial lexicon;

[0168] The historical matching database module is used to store the standard term pairs in the initial lexicon, and establish full text index and vector index for field name, table name, limited word and data element;

[0169] The historical matching module is used to retrieve the matched standard terms from the historical database according to the similarity of field name and table name;

[0170] The semantic segmentation and vector matching module is used to segment the field into limited words and data elements by fine tuning BERT model, and retrieve candidate combinations in the vector library to output Top-K recommended results;

[0171] The non-Chinese field processing module is used to extract non-Chinese field sample values, and map them to standard data elements by applying regular rules and semantic logic;

[0172] A dynamic updating module is configured to store the target data into the historical database to update the vocabulary.

[0173] As an embodiment of the present embodiment, the vocabulary construction module comprises:

[0174] A data extraction unit is configured to extract the table header fields from the data source and perform preliminary cleaning;

[0175] A semantic decomposition unit is configured to decompose the fields into the qualifier and data element pairs by using a large language model.

[0176] A verification storage unit is configured to filter the high-quality term pairs and store them into the initial vocabulary.

[0177] As an embodiment of the present embodiment, the historical target database module comprises:

[0178] A standardization storage unit is configured to store the qualifier and data element pairs into the relational database according to the field name and table name.

[0179] An index optimization unit is configured to create full-text index and vector index for the commonly used fields to accelerate the retrieval.

[0180] As an embodiment of the present embodiment, the semantic segmentation and vector matching module comprises:

[0181] A BERT segmentation unit is configured to fine-tune the BERT model and segment the fields by using the BIO tagging system.

[0182] A vector retrieval unit is configured to encode the segmentation results into vectors and retrieve similar candidates in the Faiss library.

[0183] A recommendation generation unit is configured to calculate the comprehensive similarity of the candidate combinations and output the Top-K results.

[0184] As an embodiment of the present embodiment, the non-Chinese field processing module comprises:

[0185] A sample preprocessing unit is configured to extract the non-Chinese field sample values and perform normalization.

[0186] A regular matching unit is configured to apply predefined rules to identify the date, number, and email format.

[0187] A semantic verification unit is configured to exclude ambiguity by using logical rules and confirm the final mapping.

[0188] As an embodiment of the present embodiment, the dynamic updating module comprises:

[0189] A data writing unit is configured to store the target results into the database according to the field name, qualifier, and data element.

[0190] A vector library updating unit is configured to encode the new term pair into a vector and expand the Faiss library.

[0191] The system or device is used to realize the functions of the method in the above-mentioned embodiments, each module in the system or device corresponds to each step in the method, and has been described in the method and will not be repeated here.

[0192] As shown in Figure 3 A table header field intelligent matching system structure based on semantic index segmentation is provided, comprising:

[0193] A historical matching database is configured to store metadata corresponding to historical matching data.

[0194] A historical matching qualifier / data element library is configured to store qualifiers and data elements corresponding to historical matching data.

[0195] A BERT semantic segmenter is configured to segment original data items.

[0196] A regular rule library is configured to store regular expressions corresponding to different data types.

[0197] A content discriminator is configured to discriminate non-Chinese data item fields.

[0198] An irregular content regular rule library is configured to store regular expressions corresponding to different data types of non-Chinese fields.

[0199] An m3e encoder is configured to encode text into a vector.

[0200] A Faiss vector library is configured to store a vector database.

[0201] A standard qualifier library is configured to store a database of standard qualifiers.

[0202] A standard data element word library is configured to store a database of standard data elements.

[0203] Through the above embodiments, the problem of intelligent matching of table header fields based on semantic index segmentation in the related art is solved, thereby ensuring that the problems in the prior art can be solved.

[0204] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for intelligent matching of table header fields based on semantic index segmentation, characterized in that, The method comprises the following steps: Collecting table header field data, decomposing into qualifier and data element pairs by large language model, and generating initial vocabulary; Storing the initial vocabulary in a relational database and constructing a historical matching module; For Chinese table header fields, retrieving matching standard terms from the historical database based on field name and table name similarity; If historical matching fails, use a semantic segmentation model to segment the field into qualifiers and data elements, retrieve candidate combinations from the vector library, and output Top-K recommended results; For non-Chinese fields, extract sample values and apply regular rules and semantic logic to map them to standard data elements; Store the completed matching data in the historical database to update the vocabulary.

2. The method of claim 1, wherein, The process of generating the initial vocabulary includes: Extracting table header field data and converting it to JSON format; Input large language model to decompose into qualifier and data element pairs; Filter high-quality term pairs through semantic verification and store them in categories.

3. The method of claim 1, wherein, The construction process of the historical matching module includes: Store the approved qualifier and data element pairs in the relational database according to the field name and table name; Create full-text and vector indexes for field name, table name, qualifier, and data element.

4. The method of claim 1, wherein, The working process of the semantic segmentation model includes: Fine-tune the BERT model to identify the boundaries of qualifiers and data elements in the field; Output segmentation results using the BIO tag system; Check edge cases in combination with the regular rule library.

5. The method of claim 1, wherein, The process of retrieving candidate combinations from the vector library includes: Encode the segmented qualifiers and data elements into vectors; Retrieve similar candidates in the Faiss library and calculate the cosine similarity; Expand the candidate set through synonym replacement and generate Top-K recommendations.

6. The method of claim 1, wherein, The mapping process for non-Chinese fields includes: Extract field sample values and normalize them; Apply regular rules to match dates, numbers, and email formats; Perform semantic logic rule verification on unmatched samples.

7. A semantic index segmentation based table header field intelligent matching system, characterized in that, The system includes: A vocabulary construction module for collecting table header field data and decomposing it into qualifier and data element pairs by large language model to generate an initial vocabulary; A historical matching database module for storing standard term pairs in the initial vocabulary and creating full-text and vector indexes for field name, table name, qualifier, and data element; A historical matching module for retrieving matching standard terms from the historical database based on field name and table name similarity; A semantic segmentation and vector matching module for segmenting fields into qualifiers and data elements by fine-tuning the BERT model and retrieving candidate combinations from the vector library to output Top-K recommended results; A non-Chinese field processing module for extracting non-Chinese field sample values and mapping them to standard data elements using regular rules and semantic logic; A dynamic updating module for storing the completed matching data in the historical database to update the vocabulary.

8. A device for intelligent matching of table header fields based on semantic index segmentation, characterized in that The device includes: A processor for executing the steps of the method as claimed in any one of claims 1-6; A memory for storing the initial vocabulary, the historical matching database, and the vector index; An input-output interface for receiving table header field data and outputting matching results; A communication module for interacting with external data sources and governance platforms to obtain and update data.

9. A computer terminal device, characterized by It includes: One or more processors; Memory coupled to the processor for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement steps of the method of any of claims 1-6.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, which when executed by a processor, implements steps of the method of any of claims 1-6.

Citation Information

Patent Citations

  • Data benchmarking method and device and storage device

    CN110795482A

  • Data matching method and device and electronic equipment

    CN114153962A

  • Multi-strategy data governance rule adaptation method and system based on standard data element

    CN117454188A

  • Financial data intelligent entry and verification method

    CN120449835A

  • Automatic data quality rule matching method, system and device for public security service data

    CN120745786A

Cited By

  • Method and device for determining hydropower engineering cost data code

    CN121638169A

  • A method and device for determining the cost data coding of a hydroelectric project

    CN121638169B