Data lake metadata governance method based on semantic synthesis and text vectorization
By performing semantic synthesis and text vectorization processing in the data lake, the dynamic semantic commonality of field attribute judgment in the data lake is solved, more accurate metadata management and consistency control are achieved, and the intelligence level of the data lake environment is improved.
Patent Information
- Application Number
- CN202510629028.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-16
AI Technical Summary
It is difficult for the existing technology to effectively capture the dynamic semantic commonality of fields in diversified contexts in data lakes, resulting in a single-sided judgment of field ownership and lack of dynamic judgment of semantic hierarchical relationships, resulting in a high conflict rate across data sources, and a lag in semantic structure optimization, which affects the accuracy and consistency of metadata governance.
Through methods based on semantic synthesis and text vectorization, vector encoding of field content keywords and context words is performed, semantic co-occurrence relationships are identified, semantic hierarchical structures are constructed, combined with semantic path adjustment, field attribution and consistency are optimized, and metadata evolution is dynamically responded to.
It improves the accuracy of field attribution division, enhances the metadata quality control capabilities in multi-source heterogeneous environments, ensures the coherence and accuracy of metadata governance in the structural evolution process, and improves the level of intelligence in the data lake environment.
Smart Images

Figure CN120181094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a data lake metadata governance method based on semantic synthesis and text vectorization. Background Art
[0002] The field of natural language processing technology includes the technology of computer processing, analysis, understanding and generation of text data. The core content of this field mainly involves technical means such as text vectorization, semantic understanding, automatic reasoning and text generation. Natural language processing technology is committed to simulating and analyzing human language through computers, so that computers can understand, generate and process text data. The technical methods in this field are widely used in multiple directions such as information retrieval, machine translation, sentiment analysis, speech recognition and text generation. By adopting various algorithms and models, especially models based on deep learning, natural language processing technology can realize the effective processing and intelligent analysis of large amounts of unstructured text data.
[0003] Among them, the data lake metadata governance method refers to a technology for managing and governing metadata in a data lake through semantic synthesis and text vectorization technology. This method aims at how to effectively manage large-scale heterogeneous data stored in a data lake environment, especially for improving the quality control and utilization efficiency of metadata. By adopting text vectorization technology, metadata is converted into numerical vector form, which facilitates semantic analysis and reasoning. The semantic synthesis method is used to understand, summarize and deduce metadata to ensure that the data in the data lake can be efficiently queried, accessed and managed. The solution of this patent subject involves structured processing of metadata in the data lake, accurate classification and optimized storage of metadata based on semantic understanding methods, and achieving data governance goals through text vectorization.
[0004] Existing technologies for processing data lake metadata rely on static vectorized representations of metadata and pre-defined semantic classification methods. This makes it difficult to effectively capture the dynamic semantic commonalities that arise in diverse contexts, leading to biased field attribution judgments. For example, the differences in the actual meanings of fields with the same name across different data sources are not accurately identified, which can easily lead to classification errors. Existing methods lack a dynamic judgment mechanism for field semantic hierarchical relationships. Semantic structure hierarchies are often based on fixed rules, making it difficult to address hierarchical drift in the context of complex semantic evolution. Field consistency is typically determined based on field name or path similarity, failing to integrate semantic dependencies, resulting in high cross-data source field conflict rates. During structural optimization, existing technologies lag in responding to path changes and lack the ability to dynamically identify high-frequency alternative paths. Field semantic paths cannot adapt to the pace of metadata evolution, weakening the system's adaptability. During field merging, there is a lack of linked judgment criteria for similarity and structure retention, which can easily lead to semantic drift after the field merge, affecting the accuracy and consistency of the governance structure and hindering the improvement of metadata governance capabilities in the overall data lake environment. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the existing technology and propose a data lake metadata governance method based on semantic synthesis and text vectorization.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a data lake metadata governance method based on semantic synthesis and text vectorization, comprising the following steps:
[0007] S1: Based on the field names, data types, and descriptions of the data sources in the data lake, perform vector encoding of field content keywords and adjacent context words, extract semantic co-occurrence groups of fields in differentiated contexts, identify the distance between semantic vectors, determine whether they are within the same field grouping range, and obtain the field semantic aggregation amount;
[0008] S2: Based on the field semantic aggregation amount, extract the semantic center vector of the field group, compare the cosine angle values of the semantic center vectors between the field groups, determine whether the angle value is higher than the field semantic level, and obtain the field semantic hierarchical structure;
[0009] S3: Calling the field semantic hierarchical structure, extracting the contextual semantic structure, semantic path, and semantic dependency relationship between fields in each group in the differentiated data source, and determining whether the semantic dependency exceeds the set consistency determination range. If so, marking the field as a conflicting item and obtaining the metadata consistency conflict rate;
[0010] S4: Based on the metadata consistency conflict rate, collect the replacement field paths in the data lake metadata change log, extract the frequency of occurrence of the replaced fields in the original path, screen the replacement high-frequency paths for structural position replacement, and obtain semantic path adjustment data.
[0011] As a further solution of the present invention, the field semantic aggregation quantity includes field attribution labels, contextual semantic features, and vector aggregation indicators; the field semantic hierarchical structure includes hierarchical grouping identifiers, inter-semantic layer mappings, and field group hierarchical divisions; the metadata consistency conflict rate includes the proportion of conflicting fields, field conflict distribution, and semantic conflict levels; the semantic path adjustment data includes field replacement paths, structural adjustment records, and replacement frequency indicators.
[0012] As a further solution of the present invention, the step of obtaining the field semantic aggregation amount is specifically as follows:
[0013] S111: Based on the field names, data types, and descriptions of the data sources in the data lake, extract keywords and context from the fields. Analyze word frequency and combination structure through vectorized encoding to obtain a local semantic vector for the field.
[0014] S112: Call the local semantic vector of the field, identify the Euclidean distance and cosine similarity according to the semantic vector of the field in the differentiated context, and perform distance threshold judgment for similar fields using the formula:
[0015] ;
[0016] Calculate the field semantic matching value, compare the matching value with the preset semantic attribution benchmark range, group and merge the field vectors that meet the benchmark, and obtain the field semantic attribution grouping quantity;
[0017] in, Represents the semantic matching value of the field, Representative field With fields The local semantic vector of Representative field With fields Semantic energy value in local semantics, For fields With fields The cosine similarity value between the semantic vectors;
[0018] S113: Based on the field semantic grouping quantity, for fields that have not yet been assigned, determine whether the semantic matching values with the fields in the assigned group meet the assignment conditions, retain the fields that do not meet the assignment conditions as independent fields, and obtain the field semantic collection quantity.
[0019] As a further solution of the present invention, the step of obtaining the field semantic hierarchical structure is specifically as follows:
[0020] S211: Based on the field semantic aggregation amount, extract the semantic center vector of each field group, calculate the cosine angle value of the semantic center vector between the field groups, select the field groups whose cosine angle value is lower than the preset field semantic level threshold, and obtain the preliminary semantic similarity between the field groups;
[0021] S212: Call the preliminary semantic similarity between the field groups to determine whether it is higher than the field semantic level division standard value, identify the field groups that meet the standard, and merge them into the same level group using the formula:
[0022] ;
[0023] Get the semantic level merged value of the field;
[0024] in, Represents the semantic level merged value of the field, Representative field In the The semantic center component of the dimensional vector, Representative field In the The semantic center component of the dimensional vector, Represents the number of vector dimensions;
[0025] S213: calling the field semantic level merge value, arranging the merged field level group, analyzing the level structure mapping relationship, and obtaining the field semantic hierarchical structure.
[0026] As a further solution of the present invention, the step of obtaining the metadata consistency conflict rate is specifically as follows:
[0027] S311: calling the semantic level field group defined in the field semantic hierarchical structure, extracting the contextual semantic structure of the field in the differentiated data source, parsing the semantic path, and obtaining the field semantic dependency matrix;
[0028] S312: Based on the field semantic dependency matrix, determine the overlap of field semantic paths, combine the field context semantic structure, and identify the field semantic consistency using the formula:
[0029] ;
[0030] Calculate the field consistency deviation value, filter out the field pairs that exceed the set range, and obtain the field consistency anomaly set;
[0031] in, Represents the field consistency deviation value, represents the contextual semantic features of field k, represents the contextual semantic features of the comparison field k, represents the weight coefficient of field k in the semantic path, and p represents the total number of fields in the field group;
[0032] S313: Based on the field consistency exception set, call the marked fields, determine the semantic path deviation, filter the fields that exceed the judgment range, perform consistency conflict marking, and obtain the metadata consistency conflict rate.
[0033] As a further solution of the present invention, the step of obtaining the semantic path adjustment data is specifically as follows:
[0034] S411: Based on the metadata consistency conflict rate, extract the data lake metadata change log, identify the replacement field path and the corresponding original path occurrence frequency, analyze the replacement frequency in the change log based on the number of occurrences of the original path field, filter the paths with replacement frequencies higher than a set threshold, and obtain high-frequency replacement path data;
[0035] S412: Based on the high-frequency alternative path data, analyze the degree of adjustment of the alternative position of each path in the data lake metadata, and determine the adaptability and position of the alternative path in the structure using the formula:
[0036] ;
[0037] Obtain structure position fitness data;
[0038] in, Represents the structural position fitness data, Representative Path The replacement frequency, Representative Path The frequency of occurrence in the original path, Represents the total number of records in the change log, The log number representing the first appearance of the alternative path in the change log, The log number representing the first appearance of the original path in the change log. Indicates the total number of paths;
[0039] S413: According to the structural position fitness data, the field paths in the replacement path are screened, and the corresponding field positions of the original path are replaced to obtain semantic path adjustment data.
[0040] As a further embodiment of the present invention, the method further comprises step S5:
[0041] S5: Based on the semantic path adjustment data, extract the word vector value of the field in the current collected data, compare the fusion similarity and group retention rate of the field word vector before and after the adjustment, merge the qualified fields into a unified structure group, and obtain the metadata governance fusion structure;
[0042] The metadata governance fusion structure includes a fusion field group, a unified semantic structure, and a word vector similarity matrix.
[0043] As a further solution of the present invention, the steps for obtaining the metadata governance fusion structure are specifically as follows:
[0044] S511: Based on the semantic path adjustment data, extract the fields in the currently collected data, identify the word vector value of each field, analyze the semantic relationship of the fields in the differential data table, and obtain the field word vector data;
[0045] S512: Call the field word vector data, compare the fusion similarity of the field word vectors before and after the adjustment, analyze the degree of fusion matching between the adjusted field and the original structure, identify the group retention rate, and use the formula:
[0046] ;
[0047] Calculate field fusion similarity data, analyze the similarity data, filter fields that meet the grouping threshold, and obtain field group retention rate data;
[0048] in, Represents the field fusion similarity data, and Respectively represent the field word vector values before and after adjustment, Represents the hierarchical depth of the field in the data table. Represents the number of times a field is referenced in the data lake. The number of times the structure of the representative field has changed;
[0049] S513: Based on the field grouping retention rate data, filter out qualified fields, merge them into a unified structure group, and obtain a metadata governance fusion structure.
[0050] Compared with the prior art, the advantages and positive effects of the present invention are:
[0051] The present invention vectorizes the keywords and context of fields in data sources, identifies semantic co-occurrence relationships, and uses semantic distance as the attribution criterion. This improves the depth of recognition of semantic relationships between fields in differentiated contexts, enhances the accuracy of field attribution, and establishes a hierarchical structure based on the cosine angle relationship between the semantic center vectors of each attribution group. This effectively constructs a semantic abstraction and aggregation system between fields, solves the problem of ambiguous field semantic hierarchy in traditional structures, further analyzes the semantic structure and dependency relationships of field contexts in multi-source data, accurately determines field consistency, and improves metadata quality control capabilities in multi-source heterogeneous environments. The structural position of frequently replaced field paths in change logs is adjusted to achieve dynamic optimization of field semantic paths, promoting the enhanced adaptability of semantic structures in actual business changes. Based on the comparison of path adjustment results with word vectors in the current collected data, the fusion similarity and structure retention rate are comprehensively judged to ensure the accuracy of field merging operations and the stability of system structure, and improve the consistency and accuracy of metadata governance during structural evolution. Through multi-dimensional semantic comparison and dynamic merging strategies, the intelligence level and processing efficiency of metadata management in complex data lake environments are significantly improved, and the visibility of semantic level and consistency of data governance are enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the main steps of the present invention;
[0053] Figure 2 This is a flowchart for obtaining the field semantic aggregation amount in the present invention;
[0054] Figure 3 This is a flowchart for obtaining the field semantic hierarchical structure in the present invention;
[0055] Figure 4 This is a flowchart for obtaining metadata consistency conflict rate in the present invention;
[0056] Figure 5 This is a flowchart for obtaining semantic path adjustment data in the present invention;
[0057] Figure 6 This is a flowchart for obtaining the metadata governance fusion structure in the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0059] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0060] Example 1
[0061] See also Figure 1 The present invention provides a technical solution: a data lake metadata governance method based on semantic synthesis and text vectorization, comprising the following steps:
[0062] S1: Based on the field names, data types, and descriptions of the data sources in the data lake, perform vector encoding on the keywords and adjacent context words in the field content. This extracts semantic co-occurrence groups of fields in differentiated contexts, identifies the corresponding distances between the semantic vectors of each group of fields, and determines whether the semantic distance values fall within the range for grouping similar fields. If so, the fields are grouped together; otherwise, they are retained as independent fields, resulting in the field semantic aggregation quantity.
[0063] S2: Based on the field semantic aggregation amount, the semantic center vector of each field group is extracted. The cosine angle values of the semantic center vectors between field groups are compared to determine whether the angle value is higher than the field semantic level. The judgment standard value is divided. If the judgment standard value is met, the fields are merged into the same level group to obtain the field semantic hierarchy structure.
[0064] S3: Call the semantic level field groups defined in the field semantic hierarchy structure, extract the contextual semantic structure, semantic path, and semantic dependency relationship between fields in each group in the differentiated data source, and determine whether they exceed the set field consistency judgment range. If so, mark the field as a conflicting item and obtain the metadata consistency conflict rate;
[0065] S4: Based on the metadata consistency conflict rate, we collect the replacement field paths in the data lake metadata change log, extract the frequency of occurrence of the replaced fields in the original paths, and select the replacement paths with high frequency to replace the structural positions and obtain semantic path adjustment data.
[0066] S5: Adjust the data based on the semantic path, extract the word vector value of the field in the current collected data, compare the fusion similarity and group retention rate of the field word vector before and after the adjustment, merge the qualified fields into a unified structure group, and obtain the metadata governance fusion structure.
[0067] The field semantic collection quantity includes field attribution labels, contextual semantic features, and vector aggregation indicators. The field semantic hierarchical structure includes hierarchical grouping identifiers, inter-semantic layer mapping, and field group hierarchical division. The metadata consistency conflict rate includes the proportion of conflicting fields, field conflict distribution, and semantic conflict level. The semantic path adjustment data includes field replacement paths, structural adjustment records, and replacement frequency indicators. The metadata governance fusion structure includes fused field groups, unified semantic structures, and word vector similarity matrices.
[0068] See also Figure 2 ,The specific steps for obtaining the field semantic collection amount are:
[0069] S111: Based on the field names, data types, and descriptions of the data sources in the data lake, extract keywords and context from the fields. Analyze word frequency and combination structure through vectorized encoding to obtain a local semantic vector for the field.
[0070] During semantic analysis of fields in the data lake, keywords and related context words are identified in each field. For example, when processing sales data, keywords are "sales" and "region," and context words include descriptive terms such as "growth" and "decline." Vocabulary selection is based on the business meaning of the field and the actual content of the data. Efficient vectorized encoding technology is used to convert the vocabulary into a mathematically processable vector form. For example, using word embedding technology such as Word2Vec, each word is converted into a multi-dimensional numerical vector. The importance of this step lies in laying the foundation for subsequent semantic analysis and ensuring that the vector of each word can reflect the characteristics of its semantic content. Based on the vector, a local semantic vector matrix is constructed for each field. This matrix is calculated based on the word frequency and word vector within the field. The result will directly affect the accuracy of field classification, and ultimately generate a local semantic vector for the field. The result provides a quantitative perspective to understand and compare the semantic similarity of different fields.
[0071] S112: Call the local semantic vector of the field, identify the Euclidean distance and cosine similarity based on the semantic vector of the field in the differentiated context, and determine the distance threshold of the same type of fields using the formula:
[0072] ;
[0073] Calculate the field semantic matching value, compare the matching value with the preset semantic attribution benchmark range, group and merge the field vectors that meet the benchmark, and obtain the field semantic attribution grouping quantity;
[0074] in, Represents the semantic matching value of the field, Representative field With fields The local semantic vector of Representative field With fields Semantic energy value in local semantics, For fields With fields The cosine similarity value between the semantic vectors;
[0075] For fields With fields The local semantic vector group in different context scenarios needs to detect the semantic matching value between two fields. During the execution process, the field is first collected. and fields The local semantic vector of The vector is , field The vector is ;
[0076] First calculate the square of the Euclidean distance, that is ;
[0077] Calculate semantic energy value , using the vector element square sum method, we get , similarly ;
[0078] Then calculate the cosine similarity value , by dividing the vector dot product by the module length product, the vector dot product is , the module lengths are , the cosine similarity is ;
[0079] Substituting the values into: ;
[0080] The semantic attribution benchmark range is set to 0.05, and the value , meet the attribution conditions, the field With fields Classify them into the same group and repeat the above calculation process. After traversing all field combinations and classifying and merging them, the field semantic grouping quantity is obtained.
[0081] The benefit of the formula is that by combining the square of the Euclidean distance with the square root of the semantic energy and introducing the cosine similarity product, the matching value comprehensively reflects the semantic differences and similarities between fields, thereby more reasonably distinguishing the field ownership relationship. With fields The fields are grouped into the same field group, which is then used to generate the field semantic grouping quantity.
[0082] S113: Based on the number of field semantic groupings, for fields that have not yet been assigned, determine whether their semantic matching values with the fields in the assigned group meet the assignment conditions, retain the fields that do not meet the assignment conditions as independent fields, and obtain the number of field semantic groupings;
[0083] In the process of determining whether a field should be retained as an independent field, an important step is to use the field semantic attribution grouping volume calculated previously as a reference to check those fields that have not yet been classified. For example, a field with unique semantics did not match any existing groups in the previous step. For example, a field involving sales in a specific region failed to form a semantic connection with the field due to the peculiarity of the vocabulary (such as place names). The field is re-evaluated for semantic matching to see if it can meet the matching criteria with an already formed group. The evaluation is based on the semantic distance calculation formula used in the previous step. If the matching degree is lower than the set threshold, the field is retained as an independent field. This ensures that each field is properly classified, whether it is added to an existing group or exists as a new independent entity. The final field semantic attribution volume is obtained, providing an efficient data organization method for data lake management and optimization.
[0084] See also Figure 3 ,The specific steps for obtaining the field semantic hierarchical structure are:
[0085] S211: Based on the field semantic clustering amount, the semantic center vector of each field group is extracted, the cosine angle value of the semantic center vectors between the field groups is calculated, and the field groups whose cosine angle values are lower than the preset field semantic level threshold are selected to obtain the preliminary semantic similarity between the field groups;
[0086] First, based on the metadata information in the existing data lake, the data features of each field group are extracted to form an initial semantic center vector. The process involves evaluating the data type, usage frequency, and correlation of each field. After a series of data processing steps, such as cleaning and standardization, vectorization technology is used to convert field characteristics into numerical semantic vectors. For example, the fields of a customer data table include name, address, transaction frequency, etc. By evaluating the correlation between the field and the customer classification, a preliminary semantic vector for each field can be established. Then, the cosine angle value between the field vectors is calculated to determine the similarity, which is crucial for subsequent data management and query optimization. Through actual data examples, if the cosine angle value of the customer name and customer address is high, it means that the two fields are semantically similar and belong to the same semantic level. For such fields, the preliminary semantic similarity between field groups is obtained.
[0087] S212: Call the preliminary semantic similarity between field groups to determine whether it is higher than the field semantic hierarchy classification standard value, identify the field groups that meet the standard, and merge them into the same hierarchy group using the formula:
[0088] ;
[0089] Get the semantic level merged value of the field;
[0090] in, Represents the semantic level merged value of the field, Representative field In the The semantic center component of the dimensional vector, Representative field In the The semantic center component of the dimensional vector, Represents the number of vector dimensions;
[0091] Determining whether the similarity is higher than the preset field semantic hierarchy classification standard needs to be set based on the actual business scenario. For example, in user behavior analysis on an e-commerce platform, although user click behavior and purchase behavior may appear different on the surface, they can have a high degree of semantic relevance from the perspective of data analysis. By setting a reasonable judgment threshold, seemingly different but actually related behaviors can be classified into the same hierarchy. For example, on a large e-commerce platform, the fields "number of user clicks" and "number of user purchases" represent the number of times a user clicks on a product and the number of times a user ultimately purchases the product within a certain time window, respectively. These fields can be represented by vectorization as follows: , , when calculating the cosine angle value between field groups;
[0092] First calculate the numerator:
[0093] ;
[0094] Then calculate the two parts of the denominator:
[0095] ;
[0096] ;
[0097] Calculate the cosine angle value:
[0098] ;
[0099] Then calculate the mean deviation of the second part:
[0100] ;
[0101] The final calculated merge value: ;
[0102] Assuming that the semantic level division standard value is set to 1.8, then It is higher than the set threshold, so the two fields, user click count and user purchase count, should belong to the same hierarchical group and be further merged. In this way, the data can be classified and managed more effectively. Finally, the semantic hierarchical merging relationship of the field group is analyzed to obtain the field semantic hierarchical merging value.
[0103] S213: calling the field semantic level merge value, arranging the merged field level group, analyzing the level structure mapping relationship, and obtaining the field semantic level structure;
[0104] The merged field hierarchical groups are sorted and a more refined hierarchical structure mapping relationship is established. The purpose of this step is to better utilize data to support decision-making in practical applications. For example, in financial risk management, by hierarchically structuring different types of transaction data, potential risk events can be quickly identified and responded to. Through actual data cases, such as hierarchical processing of transaction amount and transaction time data, the risk management department can locate abnormal transaction behaviors more quickly, thereby conducting effective risk control. Through such processing, the field semantic hierarchical structure can be obtained.
[0105] See also Figure 4 ,The specific steps for obtaining metadata consistency conflict rate are:
[0106] S311: calling the semantic level field group defined in the field semantic hierarchy structure, extracting the contextual semantic structure of the field in the differentiated data source, parsing the semantic path, and obtaining the field semantic dependency matrix;
[0107] In a data lake environment, performing field semantic analysis on different data sources is a key step in improving data quality management. For example, in e-commerce and medical information, accurately dividing the semantic hierarchy of fields and establishing semantic dependencies can effectively improve the accuracy and efficiency of data integration. The first step in calling based on the semantic hierarchy involves extracting the contextual semantic structure of fields from multiple data sources. This includes analyzing the frequency of use of field data, associated data types, and dependencies between fields. For example, in e-commerce, by extracting the associated usage of product IDs and user IDs, a dependency graph between fields can be constructed. The semantic path of field data can be analyzed to support data consistency between data sources. Through specific case calculations, such as the matching analysis of product data and user behavior data, it can be clearly demonstrated how to optimize the data integration process based on the semantic dependencies of fields. Ultimately, a field semantic dependency matrix is obtained, which will directly affect the efficiency and accuracy of subsequent data integration and analysis. This methodology based on actual data analysis not only improves the transparency of data processing but also provides empirical support for data governance in data lakes.
[0108] S312: Based on the field semantic dependency matrix, determine the overlap of field semantic paths, combine the field context semantic structure, and identify field semantic consistency using the formula:
[0109] ;
[0110] Calculate the field consistency deviation value, filter out the field pairs that exceed the set range, and obtain the field consistency anomaly set;
[0111] in, Represents the field consistency deviation value, represents the contextual semantic features of field k, represents the contextual semantic features of the comparison field k, represents the weight coefficient of field k in the semantic path, and p represents the total number of fields in the field group;
[0112] Calculate consistency deviations between fields to ensure semantic consistency for the same field in different data sources. For example, in financial risk management, different banks may store the credit score and income data of the same user. However, due to different data collection methods, there are certain deviations. It is necessary to quantitatively calculate the consistency of the fields and determine whether there is overlap in the semantic paths of the fields. This can be achieved through correlation analysis between fields. For example, in two data sources, the applicant's monthly income field may be recorded as "monthly_income" and "income_per_month" by different institutions, but essentially represents the same content. Text similarity calculation can determine that they belong to the same field.
[0113] For example, in loan approval, suppose the field group includes monthly income, debt ratio and credit score, and the weights of these three fields are 、 、 , the record of an applicant in Bank A is: monthly income , debt ratio , credit score , while the record in Bank B is: Monthly Income , debt ratio , credit score , substitute the value into the formula to calculate:
[0114] Calculate the numerator:
[0115] ;
[0116] Calculate the denominator:
[0117] ;
[0118] Calculate the field consistency deviation value:
[0119] Obtained by calculation The value indicates that there is a slight deviation in the field data of the user in the two data sources, but the overall consistency is high. If the preset field consistency threshold is , then the deviation value of the field pair does not exceed the set range, so it will not be marked as abnormal. If the value exceeds the set threshold, it is necessary to further analyze whether the field requires data cleaning or supplementary data sources. Based on the calculation results of all fields, the field pairs that exceed the threshold are screened to obtain a field consistency anomaly set. This set contains all fields with large differences between different data sources, providing a basis for subsequent data cleaning, field merging or data standardization.
[0120] S313: Based on the field consistency exception set, call the marked fields, determine the semantic path deviation, filter out the fields that exceed the judgment range, mark the consistency conflicts, and obtain the metadata consistency conflict rate;
[0121] The process of marking and processing fields that exceed consistency deviation thresholds is widely used in practical enterprise resource planning (ERP). For example, in supply chain management, the consistency of product and supplier data directly affects the efficiency of inventory control and order processing. Specific operations include calling a calculated set of field consistency anomalies, analyzing the deviations of each field pair, such as inconsistencies between product and supplier numbers that lead to order misdelivery or delays, and marking the consistency conflicts for each abnormal field pair. This is accomplished by setting specific marking rules, such as setting a marking threshold for abnormal field pairs. This threshold is based on past order error rates and customer complaint data. Through calculations based on actual data, for example, analyzing order issues caused by inconsistencies between product and supplier data in the past month, calculating their frequency and impact, and setting thresholds based on the data. Ultimately, through analysis and marking, the metadata consistency conflict rate is obtained. This can intuitively point out problem areas that need to be addressed first in data management, thereby guiding ERP data maintenance work, ensuring data consistency and accuracy, and improving the response speed and accuracy of the entire supply chain.
[0122] See also Figure 5 ,The steps for obtaining semantic path adjustment data are as follows:
[0123] S411: Based on the metadata consistency conflict rate, extract the data lake metadata change log, identify the replacement field path and the corresponding original path occurrence frequency, combine the number of occurrences of the original path field, analyze the replacement frequency in the change log, filter the paths with replacement frequencies higher than the set threshold, and obtain the high-frequency replacement path data;
[0124] Based on the data lake metadata change log, the frequency of occurrence of alternative field paths and their corresponding original paths is extracted, and the path data is correlated with the corresponding original path data for analysis. For example, in actual data management, an enterprise wants to analyze the most frequently updated data field paths in its data warehouse. By monitoring the daily metadata change log, the change frequency of fields in a specific data table can be identified, and the frequency of occurrence of each field in the log can be calculated through an algorithm. Through the analysis process, not only can the structural adjustment of the data warehouse be optimized, but the data model can also be adjusted more accurately to improve data processing efficiency, thereby obtaining alternative high-frequency path data.
[0125] S412: Based on the high-frequency alternative path data, analyze the degree of adjustment of the alternative position of each path in the data lake metadata to determine the adaptability and position of the alternative path in the structure using the formula:
[0126] ;
[0127] Obtain structure position fitness data;
[0128] in, Represents the structural position fitness data, Representative Path The replacement frequency, Representative Path The frequency of occurrence in the original path, Represents the total number of records in the change log, The log number representing the first appearance of the alternative path in the change log, The log number representing the first appearance of the original path in the change log. Indicates the total number of paths;
[0129] It is necessary to calculate the degree of adjustment of the alternative position of each path in the data lake metadata to determine its adaptability and set the data of the alternative field path. For example, in the data lake management of an enterprise, a field path Due to architectural adjustments, it was replaced by multiple new paths, and the original path Still exists in the data change log, by counting the change log records in the past month, the alternative path The replacement frequency is 30 times, while the original path Frequency of occurrence The total number of log records is 45. Statistics are 1000, alternative paths First appeared in log number 200, original path First appeared in log number 150, through data;
[0130] Substitute the specific values into the formula for calculation: ;
[0131] The calculation shows that the structural position fitness of this path is 0.035, which means that the structural position fitness of this alternative path in the data lake is low. Therefore, in the subsequent adjustment process, path replacement selection needs to be made more cautiously. If the path fitness is higher than a certain set benchmark value, such as 0.05, it can be considered that the path has good adaptability and structural adjustment can be carried out directly. Otherwise, further evaluation of the replacement plan is required to obtain structural position fitness data.
[0132] S413: Filter the field paths in the replacement path according to the structural position fitness data, replace the corresponding field positions of the original path, and obtain semantic path adjustment data;
[0133] Based on the structural position fitness data, those alternative paths with high fitness are screened. For example, in a complex data set, a certain path shows a high structural position fitness due to frequent data changes. The administrator can select the path to replace the field position of the original path. This process involves the readjustment of the data model. By monitoring the performance changes in real time, the data processing efficiency before and after the path replacement is evaluated. In actual applications, regular performance evaluation reports are involved so that administrators can make the most appropriate decisions. This path selection and replacement strategy based on performance evaluation can effectively optimize the data processing process and obtain semantic path adjustment data.
[0134] See also Figure 6 ,The specific steps for obtaining the metadata governance fusion structure are:
[0135] S511: Adjusting data based on the semantic path, extracting fields from the currently collected data, identifying the word vector value of each field, analyzing the semantic relationship of the fields in the differentiated data table, and obtaining the field word vector data;
[0136] Advanced data mining technology is used to extract fields from the current collected data, and the word vector value of each field is calculated through a deep learning model. The process includes preprocessing of field data, model training and vector generation. For example, in an e-commerce data lake, fields such as "user reviews" and "product descriptions" are extracted and converted into specific vector values through a word embedding model. The vector value can reflect the potential semantic correlation between fields. For example, when using the Word2Vec model training process, the context of the field will be considered. By adjusting model parameters such as window size and vector dimension, the calculation accuracy and practicality of each field vector are ensured. Through this method, the field vector can be dynamically adjusted according to user behavior and product changes in different shopping seasons, and finally the field word vector data is obtained. It not only provides similarity information between fields, but also provides a scientific basis for the next step of data structure adjustment.
[0137] S512: Call the field word vector data, compare the fusion similarity of the field word vectors before and after adjustment, analyze the degree of fusion matching between the adjusted field and the original structure, identify the group retention rate, and use the formula:
[0138] ;
[0139] Calculate field fusion similarity data, analyze the similarity data, filter fields that meet the grouping threshold, and obtain field group retention rate data;
[0140] in, Represents the field fusion similarity data, and Respectively represent the field word vector values before and after adjustment, Represents the hierarchical depth of the field in the data table. Represents the number of times a field is referenced in the data lake. The number of times the structure of the representative field has changed;
[0141] Compare the fusion similarity of the field word vectors before and after the adjustment, and calculate the field grouping retention rate. Assume that in the medical data lake, the two fields "patient diagnosis" and "treatment record" have undergone data structure adjustment. The word vector values before the adjustment are The adjusted word vector value is 0.85 The word vector value of "treatment record" before adjustment is 0.55, and after adjustment is 0.58. The depth of these two fields in the database is 3, the number of citations The statistics are 150 times, the number of structural changes The statistics is 2 times, and the values are substituted into the formula for calculation:
[0142] ;
[0143] The calculation shows that the fusion similarity of the field "Patient Diagnosis" before and after adjustment is high, while the fusion similarity of the field "Treatment Record" before and after adjustment is low. Therefore, in the field grouping strategy, the "Patient Diagnosis" field needs to be merged, and the "Treatment Record" field needs to be further adjusted to obtain the field fusion similarity data. This will directly affect the metadata governance strategy in the data lake, ensure that the optimization of the data structure meets business needs, and provide a basis for subsequent field grouping.
[0144] S513: Based on the field grouping retention rate data, the fields that meet the conditions are screened and merged into a unified structure group to obtain a metadata governance fusion structure;
[0145] Qualified fields will be screened out and merged into a unified structure group. For example, if in the data lake of the financial industry, the fields "account balance" and "transaction amount" show a high degree of fusion similarity and an appropriate grouping retention rate, they will be merged into the same data model to optimize data query and analysis efficiency. This merging strategy is determined through a series of precise data analyses, including correlation measurement between fields and grouping efficiency evaluation, to ensure that each merged field group can provide the maximum data integration benefit in the system. The final merging result will form a metadata governance fusion structure, which not only improves the efficiency of data processing, but also reduces the complexity of data management.
[0146] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A data lake metadata governance method based on semantic synthesis and text vectorization, characterized by: The following steps are involved: S1: Based on the field names, data types, and descriptions of the data sources in the data lake, perform vector encoding of field content keywords and adjacent context words, extract semantic co-occurrence groups of fields in differentiated contexts, identify the distance between semantic vectors, determine whether they are within the same field grouping range, and obtain the field semantic aggregation amount; S2: Based on the field semantic aggregation amount, extract the semantic center vector of the field group, compare the cosine angle values of the semantic center vectors between the field groups, determine whether the angle value is higher than the field semantic level, and obtain the field semantic hierarchical structure; S3: Calling the field semantic hierarchical structure, extracting the contextual semantic structure, semantic path, and semantic dependency relationship between fields in each group in the differentiated data source, and determining whether the semantic dependency exceeds the set consistency determination range. If so, marking the field as a conflicting item and obtaining the metadata consistency conflict rate; S4: Based on the metadata consistency conflict rate, collect the replacement field paths in the data lake metadata change log, extract the frequency of occurrence of the replaced fields in the original path, screen the replacement high-frequency paths for structural position replacement, and obtain semantic path adjustment data.
2. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 1 is characterized in that: The field semantic aggregation quantity includes field attribution labels, contextual semantic features, and vector aggregation indicators. The field semantic hierarchical structure includes hierarchical grouping identifiers, inter-semantic layer mappings, and field group hierarchical divisions. The metadata consistency conflict rate includes the conflict field ratio, field conflict distribution, and semantic conflict level. The semantic path adjustment data includes field replacement paths, structural adjustment records, and replacement frequency indicators.
3. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 1 is characterized in that: The specific steps for obtaining the field semantic collection amount are: S111: Based on the field names, data types, and descriptions of the data sources in the data lake, extract keywords and context from the fields. Analyze word frequency and combination structure through vectorized encoding to obtain a local semantic vector for the field. S112: Call the local semantic vector of the field, identify the Euclidean distance and cosine similarity according to the semantic vector of the field in the differentiated context, and perform distance threshold judgment for similar fields using the formula: ; Calculate the field semantic matching value, compare the matching value with the preset semantic attribution benchmark range, group and merge the field vectors that meet the benchmark, and obtain the field semantic attribution grouping quantity; in, Represents the semantic matching value of the field, Representative field With fields The local semantic vector of Representative field With fields Semantic energy value in local semantics, For fields With fields The cosine similarity value between the semantic vectors; S113: Based on the field semantic grouping quantity, for fields that have not yet been assigned, determine whether the semantic matching values with the fields in the assigned group meet the assignment conditions, retain the fields that do not meet the assignment conditions as independent fields, and obtain the field semantic collection quantity.
4. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 3 is characterized in that: The steps for obtaining the field semantic hierarchical structure are specifically as follows: S211: Based on the field semantic aggregation amount, extract the semantic center vector of each field group, calculate the cosine angle value of the semantic center vector between the field groups, select the field groups whose cosine angle value is lower than the preset field semantic level threshold, and obtain the preliminary semantic similarity between the field groups; S212: Call the preliminary semantic similarity between the field groups to determine whether it is higher than the field semantic level division standard value, identify the field groups that meet the standard, and merge them into the same level group using the formula: ; Get the semantic level merged value of the field; in, Represents the semantic level merged value of the field, Representative field In the The semantic center component of the dimensional vector, Representative field In the The semantic center component of the dimensional vector, Represents the number of vector dimensions; S213: calling the field semantic level merge value, arranging the merged field level group, analyzing the level structure mapping relationship, and obtaining the field semantic hierarchical structure.
5. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 4 is characterized in that: The steps for obtaining the metadata consistency conflict rate are specifically as follows: S311: calling the semantic level field group defined in the field semantic hierarchical structure, extracting the contextual semantic structure of the field in the differentiated data source, parsing the semantic path, and obtaining the field semantic dependency matrix; S312: Based on the field semantic dependency matrix, determine the overlap of field semantic paths, combine the field context semantic structure, and identify the field semantic consistency using the formula: ; Calculate the field consistency deviation value, filter out the field pairs that exceed the set range, and obtain the field consistency anomaly set; in, Represents the field consistency deviation value, represents the contextual semantic features of field k, represents the contextual semantic features of the comparison field k, represents the weight coefficient of field k in the semantic path, and p represents the total number of fields in the field group; S313: Based on the field consistency exception set, call the marked fields, determine the semantic path deviation, filter the fields that exceed the judgment range, perform consistency conflict marking, and obtain the metadata consistency conflict rate.
6. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 5 is characterized in that: The steps for obtaining the semantic path adjustment data are specifically as follows: S411: Based on the metadata consistency conflict rate, extract the data lake metadata change log, identify the replacement field path and the corresponding original path occurrence frequency, analyze the replacement frequency in the change log based on the number of occurrences of the original path field, filter the paths with replacement frequencies higher than a set threshold, and obtain high-frequency replacement path data; S412: Based on the high-frequency alternative path data, analyze the degree of adjustment of the alternative position of each path in the data lake metadata, and determine the adaptability and position of the alternative path in the structure using the formula: ; Obtain structure position fitness data; in, Represents the structural position fitness data, Representative Path The replacement frequency, Representative Path The frequency of occurrence in the original path, Represents the total number of records in the change log, The log number representing the first appearance of the alternative path in the change log, The log number representing the first appearance of the original path in the change log. Indicates the total number of paths; S413: According to the structural position fitness data, the field paths in the replacement path are screened, and the corresponding field positions of the original path are replaced to obtain semantic path adjustment data.
7. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 1 is characterized in that: The method further comprises step S5: S5: Based on the semantic path adjustment data, extract the word vector value of the field in the current collected data, compare the fusion similarity and group retention rate of the field word vector before and after the adjustment, merge the qualified fields into a unified structure group, and obtain the metadata governance fusion structure; The metadata governance fusion structure includes a fusion field group, a unified semantic structure, and a word vector similarity matrix.
8. The data lake metadata governance method based on semantic synthesis and text vectorization according to claim 7 is characterized in that: The steps for obtaining the metadata governance fusion structure are specifically as follows: S511: Based on the semantic path adjustment data, extract the fields in the currently collected data, identify the word vector value of each field, analyze the semantic relationship of the fields in the differential data table, and obtain the field word vector data; S512: Call the field word vector data, compare the fusion similarity of the field word vectors before and after the adjustment, analyze the degree of fusion matching between the adjusted field and the original structure, identify the group retention rate, and use the formula: ; Calculate field fusion similarity data, analyze the similarity data, filter fields that meet the grouping threshold, and obtain field group retention rate data; in, Represents the field fusion similarity data, and Respectively represent the field word vector values before and after adjustment, Represents the hierarchical depth of the field in the data table. Represents the number of times a field is referenced in the data lake. The number of times the structure of the representative field has changed; S513: Based on the field grouping retention rate data, filter out qualified fields, merge them into a unified structure group, and obtain a metadata governance fusion structure.
Citation Information
Patent Citations
Semantic-based data lake query system and method
CN114218400A
AI fusion treatment method based on data lake
CN115809235A