Intelligent metadata management method based on large model

Through intelligent metadata governance methods based on large models, and using LoRA technology to fine-tune the data processing module, the existing technology has solved the problem of insufficient accuracy when dealing with pinyin abbreviations and variable naming rules, and achieved higher Chinese description prediction accuracy and automated governance capabilities.

CN120196611APending Publication Date: 2025-06-24CHONGQING YOUNIKONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244911.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art lacks accuracy in automatically restoring the description information of database table names and field names, especially when dealing with pinyin abbreviations and variable naming rules.

Method used

Using intelligent metadata governance methods based on large models, by building data processing modules and fine-tuning using LoRA technology, the model's understanding of complex pinyin abbreviations and variable naming rules is enhanced, and accurate prediction of the Chinese description of the database is achieved.

Benefits of technology

Improve the prediction accuracy of Chinese descriptions of databases, especially when dealing with database tables named with pinyin abbreviation, it significantly improves the accuracy and readability of generating Chinese descriptions, and reduces the dependence of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196611A_ABST
    Figure CN120196611A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent metadata management method based on a large model, and the method comprises the following steps: selecting a plurality of open-source databases or a plurality of enterprise-level databases to form a data set, and synthesizing an instruction fine-tuning training set; constructing a data processing module, performing training by using an instruction fine tuning training set, calculating model generation loss by using a cross entropy loss function, and performing parameter training optimization on the model generation loss by using an Adma optimizer; and finally, combining the defined instruction design module and the trained data processing module to form a metadata processing model. According to the method, the target data can be effectively treated, and the field description and the table name description of the database in the metadata are generated and perfected in the treatment process, so that the Chinese description of the database is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data governance and artificial intelligence, and particularly relates to an intelligent metadata governance method based on a large model. Background Art

[0002] In data governance and database design, the business meanings of table names and field names are crucial for the understanding and use of data. However, in actual development, the design of table names and field names often uses the first letter abbreviations of Chinese pinyin or English abbreviations, making it difficult to directly understand their business meanings.

[0003] Currently, the methods for restoring these description information mainly include manual restoration, rule-driven restoration tools, statistics-based methods, and existing intelligent tools. Manual restoration relies on manual viewing and input, with low efficiency and easy to make mistakes, especially when dealing with large-scale data sets. Although rule-driven restoration tools use predefined rules and templates for restoration, these rules may not cover all cases and lack flexibility to adapt to different data sources and naming habits. Statistics-based methods infer description information by analyzing patterns and frequencies in data, but often cannot accurately understand the specific meaning when dealing with the first letter abbreviations of pinyin, resulting in inaccurate restoration results. Although existing intelligent tools adopt machine learning technology, they may rely on limited training data and predefined models and cannot effectively handle complex pinyin abbreviations and variable naming rules.

[0004] The deficiencies of the existing technology are mainly reflected in the following aspects. First, the restoration of the first letter abbreviations of pinyin is difficult because this restoration requires a deep understanding of the specific context and application field, which is often difficult for existing tools to achieve. Second, the rule-driven method has poor adaptability and cannot flexibly cope with different data sources and changing naming habits. Finally, the intelligent level of traditional methods is low and the effect is not good when dealing with complex first letter abbreviations of pinyin. These deficiencies indicate that the existing technology still faces significant challenges in automatically restoring the description information of table names and field names, and there is an urgent need for more intelligent and flexible solutions. Summary of the Invention

[0005] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to improve the accuracy of predicting the Chinese descriptions of the generated databases.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions: an intelligent metadata governance method based on a large model, including the following steps:

[0007] S100: Select a number of open-source databases to form a dataset D. Take each database in D as a sample. Each sample includes a table name, a description of the table name, column field names, descriptions of the column field names, column field types, and the constituent values of the columns. Among them, if a database does not have a description of the table name and column field descriptions, supplement the description of the table name and column field descriptions of this database. The selection of the database can be open-source databases on the network or enterprise-level private databases.

[0008] S200: Synthesize an instruction fine-tuning training set: Construct an instruction fine-tuning template. The instruction fine-tuning template includes instruction, input, and output. Extract information from each sample according to the instruction fine-tuning template to obtain instruction fine-tuning training samples. All the training samples are synthesized into an instruction fine-tuning training set Q.

[0009] S300: Construct a data processing module M1, and set the original weight parameter of M1 as W0. Randomly assign an initial fine-tuning value ΔW by LoRA, set the learning rate as α, and set the maximum number of training times as N. The data processing module selects and uses the open-source large model Qwen2.5 - 7B, which is an existing technology. LoRA is a model fine-tuning training method and belongs to the existing technology.

[0010] Use Q as the input data of M1, and train M1 using the Adam optimizer. Stop training when the training reaches the maximum number of training times N to obtain the trained data processing module M1'. At this time, the weight fine-tuning value corresponding to M1' is ΔW'. The Adam optimizer is an existing technology.

[0011] S400: Construct a metadata governance model M. The metadata governance model includes an instruction design module and the trained data processing module M1'.

[0012] The instruction design module includes a routing prompt word template and an instruction template. After inputting the sample into the instruction design module, the routing prompt word template is used to determine whether the sample meets the data format required by M1'. If the table name and field names of the sample are in English, it is considered not to meet the requirements, and then the sample is input into the instruction template. The output of the instruction template is the sample that meets the data format required by M1'. Otherwise, it is considered to meet the requirements. In actual applications, the instruction design module can be self-adaptively designed according to specific needs.

[0013] S500: Select an unknown metadata database sample I, input I into M, and obtain the Chinese description prediction value I' of I.

[0014] Among them, if the table name and column field name of the database information input by I are in English, M uses W0 for prediction; if the table name and column field name of the database information input by I are the initials of pinyin, M uses the weight parameter W for prediction; if the table name and field name of the sample are the initials of pinyin, it is directly input into M1′;

[0015] The specific expression of the weight parameter W is as follows:

[0016] W = W0 + ΔW = W0 + AB

[0017] Among them, W, W0 ∈ R d×k , A ∈ R d×r and B ∈ R r×k represents the weight matrix for low-rank adaptation, and r represents the rank.

[0018] Preferably, the specific steps for supplementing the Chinese descriptions of the samples in D in S100 are as follows:

[0019] Judge the types of the table name and column field name of each data sample in D. If the table name and column field name are in English description, translate the English descriptions of the table name and field name of this sample into Chinese descriptions, and then use the pypinyin library and the re library to translate the Chinese descriptions of the table name and column field name into the initials of Chinese pinyin; for example, "YGBH" represents "employee number", and both the pypinyin library and the re library are existing databases and are prior art.

[0020] If the table name and column field name are the initials of pinyin, supplement the Chinese descriptions by means of manual annotation and automatic collection; use prior art to automatically annotate these data, and the purpose of combining a small amount of manual annotation is to ensure the integrity and accuracy of the required database information.

[0021] Preferably, the specific optimization formula of the Adma optimizer in S300 is as follows:

[0022] During the training process, freeze the original weight matrix and participate in forward propagation and backward propagation. The specific formula during forward propagation is:

[0023]

[0024] The specific formula for calculating the gradients of the A and B matrices during backpropagation is:

[0025]

[0026] Among them, x represents the input data, h represents the weights of the last hidden layer of the generated description, represents the loss calculated by the model;

[0027] It is calculated using the cross-entropy loss function, and the specific formula is as follows:

[0028]

[0029] Where represents the probability distribution described by the model prediction, and y i is the encoded vector of the true description. During the training process, the Adam optimizer is used for backpropagation of gradients to fine-tune the model as described above, and the weight file with the minimum loss on the test dataset is saved. During inference, it is merged with the Qwen2.5-7B weight file to complete the inference task.

[0030] Preferably, the specific contents of the routing prompt word template and the instruction template in S400 are as follows:

[0031] The instruction template includes the table name, column field names, column field types, and the constituent values of the columns; the routing prompt word template includes judgment rules and return format requirements.

[0032] The instruction template extracts the data information required by the model from the input samples and converts it into the json format required by the model when outputting. The routing prompt word template determines whether the input data information is in English or pinyin to select and use the corresponding weight parameters for the data processing module.

[0033] Input the table name, column field names, column field types, and the constituent values of the columns of the input table, and write the prompt of the instruction template to let the model accept the input parameters and output the Chinese description information of the table name and field names in the specified format. The Prompt of the instruction template is shown in the figure, and let the model output the Chinese description information in the specified json format; the routing prompt word template uses langchain to build an application program to generate the corresponding template.

[0034] Compared with the prior art, the present invention has at least the following advantages:

[0035] 1. In this technical solution, a large number of public databases and enterprise private databases are jointly tested, and a suitable instruction adjustment module is designed. The prompt information designed in the instruction adjustment module extracts the key information in the database, and then adjusts the weight parameters in the data processing module so that it can adapt to different input data. Such classification adaptation can extract the necessary information to the greatest extent to ensure the accuracy of the final output Chinese description.

[0036] 2. This technical solution uses the LoRA technology to fine-tune the large model. The bypass structure of the LoRA weights enhances the model's understanding ability of complex pinyin abbreviations and variable naming rules, enabling the model to automatically identify and adapt to new naming patterns, improving the accuracy of generating Chinese descriptions, especially for database tables named with pinyin abbreviations.

[0037] 3. This technical solution utilizes the model inference ability and the designed routing prompt words. Based on the basic metadata and a small amount of sample data, it automatically determines the database naming type, selects the corresponding inference model to automatically fill in the business meanings of table names and field names, reduces the dependence on manual intervention, and lowers the labor cost in the data governance process. Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the operation logic of each module in the method of the present invention.

[0039] Figure 2 It is an example of the name prompt word template in the present invention.

[0040] Figure 3 It is an example of the instruction fine-tuning template in the present invention.

[0041] Figure 4 It is an example of the routing prompt word template in the present invention. Detailed Embodiment

[0042] The present invention will be further described in detail below.

[0043] Refer to Figures 1 - 4 , an intelligent metadata governance method based on a large model, including the following steps:

[0044] S100: Select several open-source databases or enterprise-level databases to form a data set D. Take each database in D as a sample, and each sample includes a table name, a table name description, column field names, column field name descriptions, column field types, and the constituent values of the columns; wherein, if the database does not have a table name description and column field descriptions, supplement the table name description and column field descriptions of the database.

[0045] The specific steps for supplementing the Chinese descriptions of the samples in D in S100 are as follows:

[0046] Judge the types of the table names and column field names of each data sample in D. If the table names and column field names are in English descriptions, translate the English descriptions of the table names and field names of this sample into Chinese descriptions, and then use the pypinyin library and the re library to translate the Chinese descriptions of the table names and column field names into the initials of Chinese pinyin.

[0047] If the table names and column field names are the initials of Chinese pinyin, supplement the Chinese descriptions by means of manual annotation and automatic collection.

[0048] S200: Synthesize the instruction fine-tuning training set: Construct an instruction fine-tuning template, which includes instruction, input, and output; Extract information from each sample according to the instruction fine-tuning template to obtain instruction fine-tuning training samples, and synthesize all the training samples into the instruction fine-tuning training set Q;

[0049] S300: Use the open-source large model Qwen2.5-7B to construct a data processing module M1, and set the original weight parameters of M1 as W0. Randomly assign an initial fine-tuning value ΔW by LoRA, set the learning rate as α, and set the maximum number of training times as N;

[0050] Use Q as the input data of M1, and use the Adam optimizer to train M1. Stop training when the training reaches the maximum number of training times N to obtain the trained data processing module M1'; At this time, the weight fine-tuning value corresponding to M1' is ΔW';

[0051] The specific optimization formula of the Adma optimizer in S300 is as follows:

[0052] During the training process, freeze the original weight matrix and participate in forward propagation and backward propagation. The specific formula during forward propagation is:

[0053]

[0054] The specific formula for calculating the gradients of matrices A and B during backpropagation is:

[0055]

[0056] Among them, x represents the input data, h represents the weight of the last hidden layer of the generated description, represents the loss calculated by the model;

[0057] Calculate using the cross-entropy loss function, and the specific formula is as follows:

[0058]

[0059] Among them represents the probability distribution of the model's predicted description, and y i is the encoded vector of the true description.

[0060] S400: Construct a metadata governance model M, and the metadata governance model includes an instruction design module and the trained data processing module M1';

[0061] The instruction design module includes a routing prompt word template and an instruction template. Edit the routing prompt word template and the instruction template according to the illustrated template. After inputting a sample into the instruction design module, the routing prompt word template is used to determine whether the sample meets the data format required by M1'. If the table name and field name of the sample are in English, it is considered not to meet the requirement, and then the sample is input into the instruction template. The output of the instruction template is the sample that meets the data format required by M1'; otherwise, it is considered to meet the requirement.

[0062] The specific contents of the routing prompt word template and the instruction template in S400 are as follows:

[0063] The instruction template includes a table name, column field names, column field types, and constituent values of the columns; the routing prompt word template includes a judgment rule and a return format requirement.

[0064] Figures 2 - 4 This is the template design format of the present invention. Specifically, according to different project requirements, a template format that meets the project needs can be designed. The instruction template extracts the data information required by the model from the input sample and converts it into the json format required by the model when outputting. The routing prompt word template determines whether the input data information is in English or pinyin, so as to determine that the data processing module selects and uses the corresponding weight parameters.

[0065] Input the table name, column field names, column field types, and constituent values of the columns of the input table, and write the prompt of the instruction template to make the model accept the input parameters and output the Chinese description information of the table name and field name in the specified format. The Prompt of the instruction template is as shown in the figure, and make the model output the Chinese description information in the specified json format; the routing prompt word template uses langchain to build an application program to generate the corresponding template.

[0066] S500: Select the unknown metadata database sample I, input I into M, and obtain the Chinese description prediction value I' of I;

[0067] Among them, if the table name and column field names of the database information input by I into M are in English, then M uses W0 for prediction; if the table name and column field names of the database information input by I into M are the initials of pinyin abbreviations, then M uses the weight parameter W for prediction; if the table name and field name of the sample are the initials of pinyin abbreviations, then it is directly input into M1';

[0068] The specific expression of the weight parameter W is as follows:

[0069] W = W0 + ΔW = W0 + AB

[0070] Among them, W, W0 ∈ R d×k , A ∈ R d×r and B ∈ R r×k represents the weight matrix for low-rank adaptation, and r represents the rank.

[0071] Experimental Contents and Results

[0072] Two enterprise-level databases (Database1 and Database2) and six open-source databases, namely PTE, AdventureWorks, TPC-DS, TPC-D, TPC-C, and TPC-H, were selected as the experimental objects. As shown in Table 1, Database1 contains 1,360 tables with a total of 15,441 fields, and the complete business meanings corresponding to the 15,441 fields were determined through Deepseek and manual annotation. Database2 contains 926 tables with a total of 17,154 fields, and the business meanings of 1,921 fields were determined through Deepseek and manual annotation. There are a large number of field names and table names abbreviated by the first characters of pinyin in Database1 and Database. The open-source databases have a total of 158 tables and 838 fields, and the business meanings of 838 fields were determined through Deepseek and manual annotation. See Table 1 for details.

[0073] Table 1 Statistical Information of Experimental Databases

[0074] Database name Number of tables Number of fields Number of marked fields PTE 38 76 76 AdventureWorks 71 486 486 TPC-DS 24 61 61 TPC-D 8 61 61 TPC-C 9 93 93 TPC-H 8 61 61 Database1 1360 15441 15441 Database2 926 17154 17154

[0075] In this experiment, Qwen2.5-7B was used as the base model, and the LoRA technique was used for parameter-efficient fine-tuning, only adjusting the parameters of some layers to improve the training efficiency. During the training process, we used the AdamW optimizer, set the learning rate to 2e-5, the batch size to 16, and trained for 3 rounds to ensure the convergence of the model. The cross-entropy loss was used as the loss function to maximize the probability that the model generates correctly restored text.

[0076] The evaluation metric is accuracy (Accuracy), which is used to measure the proportion of the generated descriptions of the model that exactly match the standard descriptions. The accuracy formula is The BLEU (Bilingual Evaluation Understudy) score was used to measure the similarity between the generated text and the reference text. The calculation formula is:

[0077]

[0078] where p n represents the precision matching rate of n-gram, n takes a maximum value of 4, w n is the weight, w1 = w2 = w3 = w4 = 0.25, and BP is the length penalty term. The recall rate calculation method based on the longest common subsequence (LCS) was adopted, and the ROUGE-L score was calculated using the word order matching between the generated description text and the reference description text. The formula is as follows:

[0079]

[0080] Among them, LCS(X, Y) represents the length of the LCS between the generated text and the reference text.

[0081] In addition to the automated evaluation metrics, the present invention also conducts manual evaluation. Three business experts score the rationality and readability of the model output, with the scoring range from 1 to 5 (1 point represents completely unreadable, and 5 points represents fully meeting the business expectations). By means of manual evaluation, the deficiencies of the automated metrics are supplemented. At the same time, to ensure the accuracy of the experimental results, the dataset is divided into a training set and a test set in an 8:2 ratio. The training set is used for training, and the test set is used to verify the final effect of the model. Finally, the average metrics of 10 experiments are recorded as the experimental results. The experimental results are as follows:

[0082] Table 2 Comparison of the results before and after training of the experimental data

[0083] Evaluation metrics Existing model Improved model Accuracy 45.2% 58.8% BLEU 52.4 73.6 ROUGE-L 52.1 71.5 Business expert score 2.2 / 5 3.8 / 5

[0084] As shown in Table 2, the comparison of various evaluation metrics of the model before and after training is as follows: Before training, the accuracy of the model is 45.2%, the BLEU value is 52.4, the ROUGE-L value is 52.1, and the business expert score is 2.2 / 5; after training, the accuracy of the model is improved to 58.8%, the BLEU value is improved to 73.6, the ROUGE-L value is improved to 71.5, and the business expert score is improved to 3.8 / 5. From each metric, the performance of the model after training has been significantly improved. Among them, the improvement amplitudes of the BLEU value and the ROUGE-L value are particularly significant, indicating that the lexical matching degree, semantic coherence of the text generated by the model, and the similarity with the reference text have been significantly improved. The improvement of the business expert score further verifies the progress of the model in terms of rationality and readability.

[0085] In short, the present invention proposes a metadata governance method and system based on a large model. It can utilize the metadata information in the database to restore the business meanings of table names and field names, and achieve automated data governance, especially for databases named with the initials of pinyin.

[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An intelligent metadata management method based on a large model, characterized by: The steps include: S100: Select several open source databases to form a data set D, and take each database in D as a sample, each sample including a table name, a table name description, a column field name, a column field name description, a column field type, and a column component value; wherein, if the database does not have a table name description and a column field description, the table name description and the column field description of the database are supplemented; S200: Synthesize instruction fine-tuning training set: construct an instruction fine-tuning template, which includes instruction, input, and output; extract information from each sample according to the instruction fine-tuning template to obtain instruction fine-tuning training samples, and synthesize all training samples into instruction fine-tuning training set Q; S300: construct a data processing module M1, and set the original weight parameter of M1 to W0, randomly assign an initial fine-tuning value ΔW by LoRA, set the learning rate to α, and set the maximum number of training times to N; Take Q as the input data of M1, use Adam optimizer to train M1, stop training when the maximum number of training times N is reached, and obtain the trained data processing module M1′; at this time, the weight fine-tuning value corresponding to M1′ is ΔW′; S400: construct a metadata governance model M, which includes an instruction design module and a trained data processing module M1′; The instruction design module includes two parts: a routing prompt word template and an instruction template. After the sample is input into the instruction design module, the routing prompt word template is used to determine whether the sample meets the data format required by M1′. If the table name and field name of the sample are in English, it is considered not to meet the requirements. Then the sample is input into the instruction template, and the output of the instruction template is the sample that meets the data format required by M1′; otherwise, it is considered to meet the requirements. S500: Select an unknown metadata database sample I, input I into M, and obtain a predicted value I' of the Chinese description of I; Among them, if the table name and column field name of the database information input by I to M are in English, M uses W0 for prediction; if the table name and column field name of the database information input by I to M are the pinyin initials, M uses the weight parameter W for prediction; if the table name and field name of the sample are the pinyin initials, directly input M1′; The specific expression of the weight parameter W is as follows: W=W0+ΔW=W0+AB Where W, W0∈R d×k , A∈R d×r and B∈R r×k represents the weight matrix of low-rank adaptation, and r represents the rank.

2. The intelligent metadata management method based on a large model as claimed in claim 1, characterized in that: The specific steps of supplementing the Chinese description of the sample in D in S100 are as follows: Determine the type of the table name and column field name of each data sample in D. If the table name and column field name are in English, translate the English description of the table name and field name of this sample into Chinese description, and then use the pypinyin library and re library to translate the Chinese description of the table name and column field name into Chinese pinyin abbreviations; If the table name and column field name are pinyin abbreviations, manual annotation and automatic collection methods are used to supplement the Chinese description.

3. The intelligent metadata management method based on a large model as described in claim 2, characterized in that: The specific optimization formula of the Adma optimizer in the S300 is as follows: During the training process, the original weight matrix is ​​frozen and participates in forward propagation and backward propagation. The specific formula for forward propagation is: The specific formula for back propagation A and B matrix gradient calculation is: Among them, x represents the input data, h represents the weight of the last hidden layer generated to describe the last layer, Represents the loss calculated by the model; The cross entropy loss function is used for calculation. The specific formula is as follows: in represents the probability distribution described by the model prediction, y i is the encoding vector of the true description.

4. The intelligent metadata management method based on a large model as claimed in claim 3, characterized in that: The specific contents of the routing prompt word template and the instruction template in S400 are as follows: The instruction template includes the table name, column field name, column field type and column composition value; the routing prompt word template includes judgment rules and return format requirements.