Metadata tag identification method based on multi-Embedding model integration
By integrating multiple embedding models and employing a feedback learning mechanism, the shortcomings of traditional metadata classification and grading methods in terms of coverage, accuracy, and automation are addressed, achieving efficient and accurate metadata tag identification and classification.
Patent Information
- Application Number
- CN202511033517.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional metadata classification and grading methods are inadequate in terms of coverage, accuracy, consistency, and automation, especially in handling complex tag systems and cross-language contexts, and lack real-time and intelligent support.
We adopt a multi-Embedding model ensemble approach, combining cross-language and dedicated Chinese models, enhancing semantic parsing through the Transformer architecture, and dynamically adjusting model weights using a feedback learning mechanism to achieve collaborative decision-making among multiple models.
It improves the accuracy and robustness of metadata classification and grading, reduces manpower and time costs, adapts to classification and grading tasks in different fields, and has strong scalability and real-time optimization capabilities.
Smart Images

Figure CN120929981A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of classification and grading technology for enterprise data governance, and more specifically, to a metadata tag recognition method based on the integration of multiple embedding models. Background Technology
[0002] Traditional metadata classification and grading methods primarily rely on regular expression matching strategies and manual sorting. Regular expression matching extracts pattern rules from a subset of samples to match and classify text. However, this method suffers from insufficient coverage and is prone to false positives and false negatives, especially when dealing with complex semantic scenarios. Manual sorting, on the other hand, requires significant manpower, is inefficient, and the consistency and accuracy of classification results are difficult to guarantee due to varying understandings of unstructured data among different employees.
[0003] Furthermore, the classification and grading standards defined in business operations are often not completely mutually exclusive, and the coexistence of single and multiple labels is quite common. This complex labeling system places high demands on traditional text multi-classification models, while existing models often exhibit certain limitations in handling such problems and fail to fully meet practical needs.
[0004] In terms of automation capabilities, traditional methods lack effective mechanisms to dynamically adjust and optimize the classification and grading process, making it difficult to provide real-time and intelligent support for enterprise data governance. At the same time, their ability to capture semantics across languages and in mixed Chinese-English contexts also has certain shortcomings, affecting the accuracy and robustness of metadata classification and grading.
[0005] In summary, traditional methods have room for improvement in terms of classification and grading coverage, consistency, automation level, and adaptability to complex tagging systems. A more efficient and intelligent solution is urgently needed to enhance metadata governance capabilities. Summary of the Invention
[0006] The purpose of this invention is to provide a metadata tag recognition method based on the integration of multiple embedding models, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a metadata tag recognition method based on multi-Embedding model integration, the method comprising the following steps: Multiple pre-trained embedding models were selected. Based on the naming characteristics of database metadata table names and field names, and the linguistic features of semantic description fields, two multilingual embedding models with cross-language capabilities (bge-m3, gte_sentence-embedding_multilingual-base) and two dedicated Chinese embedding models (Conan-embedding-v1, Yuan-embedding-1.0) were configured. The cross-language alignment capability under the Transformer architecture enhances the parsing effect of mixed Chinese and English semantics, and the multi-model complementarity mechanism overcomes the limitations of a single model in domain generalization. Pre-compute the vector representations of all label names in the label library under the four selected pre-trained Embedding models, and construct the label vector library; The system receives metadata (field names and descriptions) from database tables to be classified and graded, and calls four pre-trained embedding models in parallel to generate vector representations of the input metadata. It then calculates the cosine similarity between the input metadata and the tag library vectors in each model space. The formula for calculating cosine similarity is... ,in This represents the vector obtained by encoding the metadata T by the k-th Embedding model. Represents the vector encoded using the same type for candidate label Li, where ⋅ denotes the vector dot product. Represents the magnitude of the vector. This represents the angle between two vectors. This represents the cosine value of the included angle, and its range is [-1, 1]. Indicates the first k The target text calculated by the embedding model T and candidate tags Li The semantic similarity between the two vectors. The closer the value is to 1, the more consistent the directions of the two vectors are, and the more semantically similar they are. The weights of each model are dynamically adjusted based on a feedback learning mechanism, with the initial weights set to equal weights. A weighted comprehensive similarity score between each label and its metadata is calculated, and the label with the highest similarity score exceeding a preset threshold is selected as the final result. A feedback module is introduced. When the system receives actual user feedback (confirmation, rejection, or feedback of a real label), it updates the model weights. It determines whether the user's feedback of a real label is among the Top-3 recommendations of each embedding model and adjusts the weights based on the matching position. The weight increases when the matching position is 1st, 2nd, and 3rd; the weight decreases when there is no match. After collecting T feedbacks over a period of time, each feedback provides an increment to each model. A normalization function ensures that the sum of all model weights is 1.
[0008] As a further improvement to this technical solution, the selection of multiple pre-trained embedding models specifically includes: The naming conventions of database metadata table names and field names were analyzed. Typically, English names are used, while corresponding semantic description fields (such as "comment") store Chinese information. Considering the differences in semantic capture capabilities of different pre-trained models in cross-language scenarios and mixed Chinese-English contexts, two multilingual embedding models and two dedicated Chinese embedding models were configured. The attention mechanism and alignment capabilities of the Transformer architecture were utilized to improve the parsing accuracy of mixed Chinese-English semantics. For multilingual embedding models, the focus is on their semantic alignment capabilities in cross-language scenarios to ensure accurate capture of mixed Chinese and English semantics; for dedicated Chinese embedding models, the focus is on their adaptability in complex semantic scenarios to ensure effective handling of Chinese semantic description fields.
[0009] As a further improvement to this technical solution, the vector representation of all tag names in the pre-computed tag library under the four selected pre-trained embedding models specifically includes: Each label name in the label library is input into one of four pre-trained embedding models to generate a vector representation of each label in a different model space. For each model, the dimensions of the label vectors are ensured to be consistent, and the generated vectors are stored in a label vector library for use in subsequent similarity calculations.
[0010] As a further improvement to this technical solution, the step of receiving the metadata (field name, field description) of the database table fields to be classified and graded, and simultaneously calling four pre-trained embedding models to generate vector representations of the input metadata, specifically includes: The metadata of the database table fields to be classified and graded is preprocessed to extract field names and descriptions. These field names and descriptions are then input into four pre-trained embedding models to generate corresponding vector representations. For each model, it is ensured that the vector dimensions of the input metadata are consistent with the vector dimensions in the label vector library.
[0011] As a further improvement to this technical solution, the calculation of the cosine similarity between the input metadata and the tag library vector in each model space specifically includes: For each model space, calculate the cosine similarity between the input metadata vector and all label vectors in the label library. The formula for calculating the cosine similarity is: ,in This represents the vector obtained by encoding the metadata T by the k-th Embedding model. Represents the vector encoded using the same type for candidate label Li, where ⋅ denotes the vector dot product. Represents the magnitude of the vector. This represents the angle between two vectors. This represents the cosine value of the included angle, and its range is [-1, 1]. Indicates the first k The target text calculated by the embedding model T and candidate tags Li The semantic similarity between the two vectors. The closer the value is to 1, the more consistent the directions of the two vectors are, and the more similar their meanings are.
[0012] As a further improvement to this technical solution, the dynamic adjustment of the weights of each model based on the feedback learning mechanism specifically includes: Initialize the weights of each model to be equal, i.e., each model has an initial weight of 0.25. Calculate the weighted overall similarity between each tag and the metadata, using the formula: ,in This represents the weight of the k-th model. This represents the similarity calculated by the k-th model. The label with the highest weighted overall similarity that exceeds a preset threshold is selected as the final result.
[0013] As a further improvement to this technical solution, the introduction of a feedback module specifically includes: After the system receives actual user feedback, it determines whether the user's actual tag matches any of the top-3 recommendations in each embedding model. If the match is for the 1st position, the corresponding model weight increases; if the match is for the 2nd position, the corresponding model weight increases; if the match is for the 3rd position, the corresponding model weight increases; if there is no match, the corresponding model weight decreases. After T feedback cycles, the weight increments for each model are accumulated, and a normalization function is used to ensure that the sum of all model weights is 1.
[0014] Compared with the prior art, the technical advantages of the present invention are: This invention designs a metadata tag recognition method based on the ensemble of multiple embedding models. By integrating multiple pre-trained embedding models, it solves the limitation of single models in domain generalization. Through a feedback learning mechanism, it dynamically adjusts model weights, avoiding the simplistic "black and white" judgments of traditional methods and improving the accuracy of classification and grading. Furthermore, it supports the integration of any embedding model and can be easily replaced by a more powerful model, demonstrating strong scalability.
[0015] This invention addresses the accuracy and adaptability bottlenecks in database table field metadata tag recognition through a multi-model collaborative decision-making and feedback learning mechanism. Compared to traditional regular expression matching strategies and manual sorting methods, it significantly reduces labor and time costs, adapts to metadata classification and grading tasks in different fields, and exhibits strong robustness and scalability. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall process of the method of the present invention.
[0017] Figure 2 This is a flowchart of vector generation and similarity calculation for multiple embedding models.
[0018] Figure 3 A schematic diagram illustrating the logic of dynamically adjusting model weights for a feedback learning mechanism.
[0019] Figure 4 This is a system architecture block diagram of the present invention. Detailed Implementation
[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0021] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” and “described” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] This invention provides a metadata tag recognition method based on the integration of multiple embedding models, and its specific implementation is described in detail with reference to the accompanying drawings. Figure 1 This is a schematic diagram of the overall process of the method of the present invention, showing the overall steps from selecting a pre-trained Embedding model to the final label recognition result output. Figure 2 This is a flowchart of vector generation and similarity calculation for multi-Embedding models, which details the process of vector generation and cosine similarity calculation for input metadata and tag libraries in different model spaces. Figure 3 This is a logical diagram illustrating how user feedback affects the adjustment of model weights and the process of weight normalization. Figure 4 The system architecture diagram presents a modular design for multi-Embedding model integration, label vector library construction, feedback module, and final classification and grading decision.
[0023] First of all Figure 1In this process, selecting the pre-trained embedding model is a crucial first step. Based on the naming characteristics of database metadata table names and field names, and the linguistic features of semantic description fields, two multilingual embedding models with cross-language capabilities (bge-m3, gte_sentence-embedding_multilingual-base) and two dedicated Chinese embedding models (Conan-embedding-v1, Yuan-embedding-1.0) are configured. These models are chosen based on the cross-language alignment capabilities and attention mechanisms of the Transformer architecture, which enhance the parsing effect of mixed Chinese and English semantics. The metadata input module is responsible for receiving the database table field metadata (field names and field descriptions) to be classified and graded, and inputting this data into the four pre-trained embedding models to generate corresponding vector representations. During this process, the field names and field descriptions are input into the four pre-trained embedding models respectively, ensuring that the vector dimensions of the input metadata are consistent with the vector dimensions in the label vector library. The construction of the label vector library requires inputting all the label names in the label library one by one into four pre-trained embedding models to generate vector representations of each label in different model spaces, and storing them in the label vector library for subsequent similarity calculations.
[0024] exist Figure 2 In this module, the similarity calculation module is responsible for calculating the cosine similarity between the input metadata and the tag library vectors in each model space. For each model space, the cosine similarity between the input metadata vector and all tag vectors in the tag library is calculated one by one. The formula for calculating the cosine similarity is: ,in This represents the vector obtained by encoding the metadata T by the k-th Embedding model. This represents the vector encoded using the same type for candidate label Li, where ⋅ represents the vector dot product. Represents the magnitude of a vector. This represents the angle between two vectors. This represents the cosine value of the included angle, ranging from [-1, 1]. Indicates the first k The target text calculated by the embedding model T and candidate tags Li The semantic similarity between the two vectors. The closer the value is to 1, the more consistent the directions of the two vectors are, and the more similar they are semantically. Through this process, the similarity value between the input metadata in each model space and all tags in the tag library can be obtained.
[0025] exist Figure 3 In this model, the feedback module and the weight adjustment module work together to achieve dynamic weight adjustment based on the feedback learning mechanism. Initially, the weights of each model are set to equal weights, i.e., the initial weights of each model are... The weighted average similarity between each tag and its metadata is 1 / 4. The formula for calculating the weighted average similarity between each tag and its metadata is: ,in This represents the weight of the k-th model. This represents the similarity calculated by the k-th model. The label with the highest weighted overall similarity that exceeds a preset threshold is selected as the final result. After the system receives actual user behavior feedback, it determines whether the user's actual label is among the Top-3 recommendations of each Embedding model and adjusts the weights according to Table 1.
[0026] Table 1:
[0027] If the matching position is 1st, the corresponding model weight increases by + If the matching position is 2nd, the corresponding model weight will increase by + If the matching position is the 3rd, the corresponding model weight will increase by + If a match is missed, the corresponding model weights are reduced. After T feedback cycles, the weight increments for each model are accumulated, and a normalization function is used to ensure that the sum of all model weights equals 1. This process involves the feedback module collecting user feedback information and transmitting it to the weight adjustment module for weight updates, thereby optimizing the model's recommendation performance.
[0028] exist Figure 4 In this system, the architecture is achieved through the collaborative work of multiple modules. The pre-trained embedding model generates vector representations of the input metadata and the label library. The label vector library stores the label vectors. The metadata input module receives and preprocesses the metadata. The similarity calculation module calculates the cosine similarity. The feedback module collects user feedback. The weight adjustment module dynamically adjusts the model weights. The classification and hierarchical decision module selects the final label based on the weighted comprehensive similarity. The connections and collaboration between these modules ensure the efficient operation of the entire system. For example, the metadata input module passes the pre-processed metadata to the pre-trained embedding model, the generated vector representation is passed to the similarity calculation module for similarity calculation, and the result is further passed to the classification and hierarchical decision module for label selection. Simultaneously, the feedback module passes user feedback to the weight adjustment module, dynamically adjusts the model weights, and then re-participates in the similarity calculation and label selection process.
[0029] In practical applications, suppose a company's database contains a large amount of metadata that needs to be classified and graded. The method of this invention first selects a pre-trained embedding model and vectorizes the tag library to construct a tag vector library. Then, the metadata input module receives the metadata of the database table fields to be classified and graded, including field names and descriptions. This information is passed to the pre-trained embedding model to generate vector representations, and the similarity calculation module calculates the cosine similarity with all tags in the tag library. The classification and grading decision module selects the final tags based on the weighted comprehensive similarity, completing the initial classification and grading task. During user interaction, the feedback module collects user feedback on recommended tags (confirmation, rejection, or actual tags) and passes this information to the weight adjustment module for dynamic adjustment of model weights. After a period of feedback accumulation, the system can gradually optimize the recommendation effect and improve the accuracy of classification and grading.
[0030] In the specific implementation of this invention, the selection and configuration of the pre-trained Embedding model 1 are crucial. Analysis of the naming conventions of database metadata table names and field names revealed that English naming is typically used, while the corresponding semantic description fields store Chinese information. Considering the differences in semantic capture capabilities of different pre-trained models in cross-language scenarios and mixed Chinese-English contexts, two multilingual Embedding models and two dedicated Chinese Embedding models were configured. The attention mechanism and alignment capabilities of the Transformer architecture were utilized to improve the parsing accuracy of mixed Chinese-English semantics. For the multilingual Embedding model, its semantic alignment capability in cross-language scenarios was examined to ensure accurate capture of mixed Chinese-English semantics. For the dedicated Chinese Embedding model, its adaptability in complex semantic scenarios was examined to ensure effective processing of Chinese semantic description fields.
[0031] Furthermore, during the construction of the label vector library, it is necessary to ensure the consistency of the label vector dimensions. All label names in the label library are input one by one into four pre-trained embedding models to generate vector representations of each label in different model spaces. For each model, the generated vectors are stored in the label vector library for subsequent similarity calculations. This process requires ensuring the consistency of the label vector dimensions to avoid calculation errors caused by inconsistent dimensions.
[0032] During the metadata input module's operation, the metadata of the database table fields to be classified and graded is preprocessed to extract field names and descriptions. These field names and descriptions are then input into four pre-trained embedding models to generate corresponding vector representations. For each model, it is ensured that the vector dimensions of the input metadata are consistent with those in the label vector library. This process requires strict control over the format and content of the input data to ensure that the generated vectors accurately reflect the semantic information of the metadata.
[0033] During the similarity calculation module's operation, for each model space, the cosine similarity between the input metadata vector and all tag vectors in the tag library is calculated one by one. The formula for calculating cosine similarity is: ,in This represents the vector obtained by encoding the metadata T by the k-th Embedding model. This represents the vector encoded using the same type for candidate label Li, where ⋅ represents the vector dot product. Represents the magnitude of a vector. This represents the angle between two vectors. This represents the cosine value of the included angle, ranging from [-1, 1]. Indicates the first k The target text calculated by the embedding model T and candidate tags Li The semantic similarity between the vectors is calculated. A value closer to 1 indicates that the two vectors are more aligned in direction and therefore more semantically similar. This process requires ensuring computational accuracy to avoid biases in recommendation results due to calculation errors.
[0034] During the operation of the feedback module and weight adjustment module, after the system receives the user's actual behavioral feedback, it determines whether the user's actual label is among the Top-3 recommendations of each embedding model, as shown in Table 1. If the matching position is 1st, the corresponding model weight is increased by +. If the matching position is 2nd, the corresponding model weight will increase by + If the matching position is the 3rd, the corresponding model weight will increase by +. If a match is missed, the corresponding model weights are reduced. After T feedback cycles, the weights of each model are incremented. The expression for the change in the weights of the i-th Embedding model is as follows: The new weight expression for the i-th Embedding model is as follows: In the formula, This represents the old weight values of the i-th Embedding model. This represents the new weight values of the i-th Embedding model. It is a normalization function, and its expression is as follows: In the formula, N refers to the number of embedding models, which is 4 in this case. This indicates the number of times the index is traversed, with a value ranging from 1 to 4. This represents the change in the weight of the j-th Embedding model.
[0035] The normalization function ensures that the sum of all model weights is 1. This process requires ensuring the accuracy and timeliness of feedback information to avoid weight adjustment deviations caused by feedback delays or errors.
[0036] Finally, during the classification and grading decision-making module 7, the weights of each model are dynamically adjusted based on a feedback learning mechanism. The weighted comprehensive similarity between each label and its metadata is calculated, and the label with the highest similarity exceeding a preset threshold is selected as the final result. This process needs to ensure the rationality and accuracy of the decision-making process to avoid classification and grading errors caused by decision-making mistakes.
[0037] To enable those skilled in the art to fully understand and implement this invention, the specific implementation principle of this invention is further explained below in conjunction with a specific application scenario.
[0038] In practical applications, suppose a company's database contains a large amount of table field metadata that needs to be categorized and graded. This metadata includes field names and field descriptions. Field names are usually named in English, while field descriptions are mainly in Chinese, and may involve mixed Chinese and English semantics. The method of this invention first selects a pre-trained embedding model and vectorizes the tag library to construct a tag vector library. Subsequently, the metadata input module receives the database table field metadata to be categorized and grades, and inputs the field names and field descriptions into four pre-trained embedding models to generate vector representations. Finally, the system completes the classification and grading task through the collaborative work of a similarity calculation module, a feedback module, a weight adjustment module, and a classification and grading decision module.
[0039] Step 1: Selection and Configuration of Pre-trained Embedding Models. Based on the characteristics of the database metadata, two multilingual embedding models (bge-m3, gte_sentence-embedding_multilingual-base) and two dedicated Chinese embedding models (Conan-embedding-v1, Yuan-embedding-1.0) are configured. These models are based on the Transformer architecture, possessing cross-language alignment capabilities and attention mechanisms, effectively parsing mixed Chinese and English semantics. For example, the bge-m3 model captures the semantic relationships between field names and field descriptions by aligning the semantic spaces of different languages; the Conan-embedding-v1 model focuses on complex Chinese semantic scenarios, ensuring accurate parsing of field description information. By integrating multiple models, the limitations of single models in domain generalization are overcome, improving the overall classification performance.
[0040] Step 2: Construction and Storage of the Label Vector Library. All label names in the label library are input one by one into four pre-trained embedding models to generate vector representations for each label in different model spaces. For example, for the label "user identity information," the bge-m3 model is used to generate its corresponding vector representation, which is then stored in the label vector library. This process ensures the consistency of the label vector dimensions, avoiding subsequent calculation errors due to inconsistent dimensions. Furthermore, the label vector library is designed to support dynamic expansion; the vector representations can be updated at any time when new labels are added, ensuring the system's flexibility and adaptability.
[0041] Step 3: Preprocessing and Vectorization of Metadata Input Module. The metadata input module receives the database table field metadata to be classified and graded, extracting the field names and descriptions. For example, the field name is "user_id," and the field description is "the user's unique identifier." The field names and descriptions are then input into four pre-trained embedding models to generate corresponding vector representations. For each model, it is ensured that the vector dimension of the input metadata is consistent with the vector dimension in the label vector library. For example, the "user_id" vector generated by the bge-m3 model has a dimension of 768, and the "user's unique identifier" vector generated by the Conan-embedding-v1 model also has 768 dimensions. This process strictly controls the format and content of the input data to ensure that the generated vectors accurately reflect the semantic information of the metadata.
[0042] Step 4: Cosine similarity calculation in the similarity calculation module. For each model space, calculate the cosine similarity between the input metadata vector and all tag vectors in the tag library. For example, in the bge-m3 model space, calculate the cosine similarity between the "user_id" vector and the tag "user identity information" vector. The formula for calculating cosine similarity is... ,in This represents the vector obtained by encoding the metadata T by the k-th Embedding model. This represents the vector encoded using the same type for candidate label Li, where ⋅ represents the vector dot product. Represents the magnitude of a vector. This represents the angle between two vectors. This represents the cosine value of the included angle, ranging from [-1, 1]. This represents the semantic similarity between the target text T and candidate label Li calculated by the k-th embedding model. The closer the value is to 1, the more consistent the directions of the two vectors are, and the more semantically similar they are. Through this process, the similarity value between the input metadata in each model space and all labels in the label library can be obtained. For example, the cosine similarity between "user_id" and "user identity information" is 0.85, indicating that the two are highly semantically related.
[0043] Step 5: Dynamic weight adjustment of the feedback module and weight adjustment module. Initially, all model weights are set to equal weight, i.e., each model's initial weight is 1 / 4. The weighted comprehensive similarity between each tag and metadata is calculated using the following formula: ,in This represents the weight of the k-th model. This represents the similarity calculated by the k-th model. For example, for the label "user identity information," its weighted comprehensive similarity is 0.783. The label with the highest weighted comprehensive similarity that exceeds a preset threshold is selected as the final result. After the system receives actual user feedback, it determines whether the user's actual label is among the Top-3 recommendations of each Embedding model. For example, if the user confirms "user identity information" as the correct label, and this label is among the Top-1 recommendations of the bge-m3 model, then the corresponding model weight is increased. If it's in the Top-2 recommendations of the Conan-embedding-v1 model, then the corresponding model weights will increase. After T feedback cycles, the weight increments for each model are accumulated, and a normalization function is used to ensure that the sum of all model weights equals 1. This process involves collecting user feedback information through the feedback module and transmitting it to the weight adjustment module 6 for weight updates, thereby optimizing the model's recommendation performance.
[0044] Step Six: Final Label Selection in the Classification and Grading Decision Module. After dynamically adjusting the weights of each model based on the feedback learning mechanism, the weighted comprehensive similarity between each label and the metadata is recalculated. The label with the highest similarity exceeding a preset threshold is selected as the final result. For example, after weight adjustment, the weighted comprehensive similarity of "user identity information" increased from 0.85 to 0.90, exceeding the preset threshold of 0.80, and therefore it was selected as the final label. This process ensures the rationality and accuracy of the decision, avoiding classification and grading errors caused by decision-making mistakes.
[0045] Through the collaborative efforts of the above steps, this invention achieves efficient classification and grading of metadata for database table fields. For example, when processing metadata with the field name "transaction_amount" and the field description "transaction amount," the system can accurately match the tag "financial transaction information" and continuously optimize the recommendation effect based on user feedback. After a period of feedback accumulation, the system can gradually improve the accuracy of classification and grading, meeting the actual needs of enterprise data governance.
[0046] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A metadata tag recognition method based on multi-Embedding model ensemble, characterized in that, The steps are as follows: S1. Select multiple pre-trained Embedding models, and configure two multilingual Embedding models with cross-language capabilities and two dedicated Chinese Embedding models. S2. Pre-compute the vector representations of all label names in the label library under the four selected pre-trained Embedding models, and construct the label vector library; S3. Receive the metadata of the database table fields to be classified and graded, including field names and field descriptions, and call four pre-trained Embedding models in parallel to generate vector representations of the input metadata. S4. Calculate the cosine similarity between the input metadata and the tag library vector in each model space; S5. Based on the feedback learning mechanism, dynamically adjust the weights of each model. The initial weights are set to equal weights. Calculate the weighted comprehensive similarity between each label and the metadata. Select the label with the highest similarity and exceeding the preset threshold as the final result. S6. Introduce a feedback module. After the system receives feedback on the user's actual behavior, it updates the model weights. After T feedback cycles, the weight increments of each model are accumulated, and a normalization function is used to ensure that the sum of all model weights is 1.
2. The metadata tag recognition method based on multi-Embedding model integration according to claim 1, characterized in that: The selection of multiple pre-trained embedding models specifically includes analyzing the naming conventions of database metadata table names and field names, combining the differences in semantic capture capabilities of different pre-trained models in cross-language scenarios and mixed Chinese and English contexts, and configuring two multilingual embedding models and two dedicated Chinese embedding models.
3. The metadata tag recognition method based on multi-Embedding model integration according to claim 2, characterized in that: The vector representation of all label names in the pre-computed label library under the four selected pre-trained embedding models is specifically calculated by inputting all label names in the label library into the four pre-trained embedding models one by one, generating a vector representation of each label in different model spaces, and storing it in the label vector library.
4. The metadata tag recognition method based on multi-Embedding model integration according to claim 3, characterized in that: The process of receiving the metadata of the database table fields to be classified and graded specifically includes preprocessing the metadata of the database table fields to be classified and graded, extracting the field names and field descriptions, and inputting the field names and field descriptions into four pre-trained embedding models to generate corresponding vector representations.
5. The metadata tag recognition method based on multi-Embedding model integration according to claim 4, characterized in that: The step of calculating the cosine similarity between the input metadata and the tag library vectors in each model space specifically includes, for each model space, calculating the cosine similarity between the input metadata vector and all tag vectors in the tag library, using the following formula: ,in This represents the vector obtained by encoding the metadata T by the k-th Embedding model. Represents the vector encoded using the same type for candidate label Li, where ⋅ denotes the vector dot product. Represents the magnitude of the vector. This represents the angle between two vectors. This represents the cosine of the included angle, and its range is [-1, 1]. Indicates the first k The target text calculated by the embedding model T and candidate tags Li The semantic similarity between the two vectors is such that the closer the value is to 1, the more consistent the directions of the two vectors are and the more similar their meanings are.
6. The metadata tag recognition method based on multi-Embedding model integration according to claim 5, characterized in that: The dynamic adjustment of model weights based on the feedback learning mechanism specifically includes initializing each model weight to be equal, and calculating the weighted comprehensive similarity between each label and metadata, using the formula: ,in This represents the weight of the k-th model. This represents the similarity calculated by the k-th model.
7. The metadata tag recognition method based on multi-Embedding model integration according to claim 6, characterized in that: The aforementioned feedback module specifically includes determining whether the user's actual behavioral feedback tag is in the Top-3 recommendations of each Embedding model after the system obtains the user's actual behavioral feedback, and adjusting the weight according to the matching position: the weight increases by Δ1 when the matching position is 1st, by Δ2 when the matching position is 2nd, by Δ3 when the matching position is 3rd, and decreases by Δ4 when there is no match.
8. The metadata tag recognition method based on multi-Embedding model integration according to claim 2, characterized in that: The multilingual embedding models include bge-m3 and gte_sentence-embedding_multilingual-base, and the dedicated Chinese embedding models include Conan-embedding-v1 and Yuan-embedding-1.
0.
9. A metadata tag recognition method based on multi-Embedding model integration according to claim 8, characterized in that: The vector dimensions in the tag vector library are consistent with the vector dimensions of the input metadata, ensuring consistency of vector dimensions during similarity calculation.
10. A metadata tag recognition method based on multi-Embedding model integration according to claim 9, characterized in that: After collecting user feedback information, the feedback module transmits it to the weight adjustment module for weight update. The weight adjustment module ensures that the sum of all model weights is 1 through a normalization function.