Intelligent data governance method and system based on active metadata

By preprocessing active metadata, calculating the difference value of structural complexity and change characteristic metrics, and adaptively selecting the reference distance calculation method in the DTW algorithm, the error problem of the DTW algorithm when processing multi-source heterogeneous active metadata is solved, and the matching effect and data governance efficiency are improved.

CN120011411AActive Publication Date: 2025-05-16BEIJING HUADIAN TIANREN ELECTRIC POWER CONTROL TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510442935.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-16
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

When the DTW algorithm processes multi-source heterogeneous active metadata, there are errors introduced due to different dimension differences, and does not consider the local change characteristics of the sequence, resulting in poor quality of unreasonable path matching and data matching.

Method used

By preprocessing active metadata, calculate the difference value of structural complexity and change characteristic metrics, adaptively select the reference distance calculation method in the DTW algorithm, and use the relative change distance calculation method to improve the matching effect.

Benefits of technology

The matching and mapping effect of active metadata is improved, and the consistency governance of multi-source active metadata is realized, ensuring efficient integration, accurate analysis and effective utilization of active metadata.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011411A_ABST
    Figure CN120011411A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of active metadata management, in particular to an intelligent data management method and system based on active metadata. The method comprises the following steps: calculating the structural complexity of two groups of active metadata according to the number of attributes of the two groups of active metadata, a correlation metric value and a combined entropy value; calculating a stability metric value according to the average value and the maximum value point of the group of active metadata, and obtaining a data feature metric value; further obtaining a change characteristic measurement difference value of the two groups of active metadata; performing weighted summation on the structural complexity and the change characteristic measurement difference value to obtain a complexity difference degree value; and setting a calculation mode of a reference distance of a DTW algorithm, selecting the calculation mode of the reference distance according to the complex difference degree value of the two groups of active metadata, and processing the two groups of active metadata by using the DTW algorithm. According to the method, the matching and mapping effects of the active metadata can be improved, and the consistency treatment of the multi-source active metadata is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of active metadata management, and in particular to an intelligent data governance method and system based on active metadata. Background Art

[0002] Active metadata is a type of active metadata that can automatically update, manage, interpret, and interact with data and operations in the system. It is data that describes data and provides information about data. Currently, various active metadata are usually effectively governed through intelligent means to improve data quality, security, and compliance. In particular, active metadata governance is more necessary in cross-system, cross-business, and cross-data source environments. In such an environment, active metadata involves multiple systems or platforms, and data structures or attributes often differ. Therefore, it is necessary to achieve consistent governance of multi-source data by matching and mapping active metadata.

[0003] In the existing technology, the DTW algorithm (Dynamic Time Warping) is usually used to solve the problem of heterogeneous active metadata alignment and comparison in data matching and mapping in the matching and mapping of heterogeneous active metadata across data sources. Through data matching and mapping, data can be efficiently integrated, information islands can be eliminated, and data quality can be improved, thereby promoting the development of intelligent data governance and improving the accuracy and efficiency of data-driven decision-making.

[0004] However, when the DTW algorithm processes multi-source heterogeneous active metadata, when constructing the distance accumulation matrix for active metadata with different attributes and different dimensions, the difference in different dimensions will introduce errors in calculating the distance. In addition, the DTW algorithm only considers the distance information of a single point pair in the sequence, without considering the local change characteristics of the sequence. The contribution of the local characteristics of the data to the distance calculation is underestimated, which easily leads to unreasonable path matching. Secondly, the purpose of completing the data matching and mapping through the DTW algorithm for active metadata across systems, businesses, and data sources is to achieve cross-data source integration and analysis, which is the main purpose of the active metadata governance method and system. Therefore, it is necessary to perform DTW algorithm calculations between active metadata to complete the integration of multi-source data and discover potential association patterns while retaining the original units and scope of the data. During the operation of the DTW algorithm, the Euclidean distance measurement method is usually selected. However, in this scenario, due to the complexity of multi-source active metadata and the DTW algorithm itself not considering the local change characteristics of the sequence, an optimal curved path with poor effect is ultimately produced, resulting in poor final data matching quality. The system may not be able to provide accurate data mapping, affecting subsequent data governance, analysis, and decision-making. Summary of the invention

[0005] In order to solve the above technical problems, the purpose of the present invention is to provide an intelligent data governance method and system based on active metadata. The technical solutions adopted are as follows: In a first aspect, an embodiment of the present invention provides an intelligent data governance method based on active metadata, the method comprising: Obtain active metadata of different data sources and perform preprocessing to obtain active metadata of different groups; The structural complexity of the two sets of active metadata is calculated based on the number of attributes, correlation measurement values ​​and combined entropy values ​​of the two sets of active metadata; Calculating a stationary measurement value of a set of active metadata according to an average value and a maximum point of the set of active metadata; obtaining a data feature measurement value according to a stationary measurement value of a set of active metadata and a normalized sampling frequency; Based on the difference of the data characteristic measurement values ​​of the two sets of active metadata, the change characteristic measurement difference value of the two sets of active metadata is obtained; the structural complexity and change characteristic measurement difference value of the two sets of active metadata are weightedly summed to obtain the complexity difference degree value; The calculation method of the reference distance of the DTW algorithm is set, and the calculation method of the reference distance is selected according to the complex difference degree value of the two sets of active metadata, and the two sets of active metadata are processed using the DTW algorithm.

[0006] Preferably, obtaining active metadata of different data sources and preprocessing to obtain different groups of active metadata includes: The preprocessing includes data cleaning, data format conversion and data grouping; the active metadata that has undergone data cleaning and format conversion is grouped according to different data sources to obtain different groups of active metadata.

[0007] Preferably, calculating the structural complexity of the two sets of active metadata according to the number of attributes, correlation measurement values ​​and combined entropy values ​​of the two sets of active metadata includes: The number of attributes of the data in the two sets of active metadata is normalized to obtain the attribute characteristic value; the two sets of active metadata are used to perform straight line fitting to obtain two straight lines, and the correlation measurement value of the two sets of active metadata is obtained according to the slopes of the two straight lines; the correlation measurement value is negatively correlated with the exponential function with the natural constant as the base to obtain the correlation characteristic value; the joint entropy value of the two sets of active metadata is normalized to obtain the entropy characteristic value; the average value of the attribute characteristic value, correlation characteristic value and entropy characteristic value of the two sets of active metadata is the structural complexity of the two sets of active metadata.

[0008] Preferably, the calculation formula of the correlation metric value is: , in, represents the correlation metric value between the i-th group of active metadata and the j-th group of active metadata; and They respectively represent the slope of the straight line obtained by straight-line fitting using the i-th group of active metadata and the slope of the straight line obtained by straight-line fitting using the j-th group of active metadata.

[0009] Preferably, calculating the stationarity metric value of a set of active metadata according to the average value and the maximum value point of the set of active metadata includes: The absolute value of the difference between each data value and the average value in a set of active metadata is obtained and summed to obtain the degree of dispersion; a maximum point in the set of active metadata is subtracted from the data adjacent to the left and the data adjacent to the right and summed to obtain the local data change value of the maximum point; the local data change values ​​of all the maximum points in the set of active metadata are summed to obtain the fluctuation change characteristic value; the degree of dispersion of the set of active metadata and the sum of the fluctuation change characteristic values ​​are normalized to obtain the stability measurement value of the set of active metadata.

[0010] Preferably, obtaining the change characteristic measurement difference value of the two sets of active metadata based on the difference of the data characteristic measurement values ​​of the two sets of active metadata includes: The absolute value of the difference between the data feature measurement values ​​of the two sets of active metadata is normalized to obtain the change feature measurement difference value of the two sets of active metadata.

[0011] Preferably, the calculation method of setting the reference distance of the DTW algorithm includes: The reference distance calculation method of the DTW algorithm includes two calculation methods, namely the traditional method and the relative change distance calculation method; the traditional method uses the Euclidean distance as the reference distance; the relative change distance calculation method is: , in, Indicates the relative change distance between the i-th data point and the j-th data point; and Represent the data values ​​of the i-th data point and the i-1-th data point respectively; and Represent the data values ​​of the j-th data point and the j-1-th data point respectively.

[0012] Preferably, the calculation method of the reference distance is selected according to the complexity difference degree values ​​of the two sets of active metadata, and the two sets of active metadata are processed using the DTW algorithm, including: A threshold is set. If the complexity difference value of the two sets of active metadata is less than or equal to the threshold, the traditional method is selected to calculate the reference distance; if the complexity difference value of the two sets of active metadata is greater than the threshold, the relative change distance calculation method is selected.

[0013] Preferably, obtaining the data feature measurement value includes: The data feature metric is obtained by adding the stationarity metric of a set of active metadata and the normalized sampling frequency.

[0014] In a second aspect, the present invention also provides an intelligent data governance system based on active metadata, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of an intelligent data governance method based on active metadata when executed by the processor.

[0015] The embodiments of the present invention have at least the following beneficial effects: the present invention pre-processes active metadata from different sources to obtain active metadata of different groups, thereby improving the quality of active metadata and making subsequent analysis results more accurate; further, by analyzing the complexity of the two groups of active metadata to be subjected to the DTW algorithm based on the number of attributes, correlation measurement values ​​and combined entropy values, the structural complexity is obtained, and the analysis from three dimensions can help analyze whether the calculation of the reference distance using the traditional method during the DTW algorithm will be affected; then, a set of active metadata stationarity measurement values ​​are obtained, and then the data characteristic measurement values ​​are obtained based on the change distribution characteristics of the data in combination with the sampling frequency, thereby measuring the data characteristics of the two groups of active metadata. The difference in characteristics is obtained to obtain the difference value of change feature measurement, which is used to evaluate the comprehensive difference in the stability and sampling frequency of the two sets of active metadata; then the structural complexity and change feature measurement difference value of the two sets of active metadata are combined to obtain the complex difference degree value, which indicates the complexity and difference between the two sets of active metadata. Finally, the calculation method of the reference distance of the DTW algorithm is set, and then the appropriate reference distance calculation method in the DTW algorithm is adaptively selected based on the complex difference degree value to obtain a more appropriate reference distance, and to construct an accurate distance matrix to obtain a more reasonable path matching, improve the matching and mapping effects of active metadata, realize the consistency management of multi-source active metadata, and ensure the efficient integration, accurate analysis and effective utilization of active metadata. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A method flow chart of an intelligent data governance method based on active metadata provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of an intelligent data governance method and system based on active metadata proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form.

[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0020] The following is a detailed description of a specific solution of an intelligent data governance method and system based on active metadata provided by the present invention in conjunction with the accompanying drawings.

[0021] Embodiment 1: The main application scenarios of the present invention are: in view of the characteristics of active metadata across systems, businesses, and data sources, it is necessary to use the DTW algorithm to match and map active metadata to achieve consistency management of active metadata. This is because the DTW algorithm can achieve multi-source heterogeneous data matching and mapping through sequence similarity without the need for unified data processing, even if the timestamps in different data sources (such as multiple sensor data, log data, etc.) have different sampling frequencies, time offsets, or uneven intervals. However, in the process of processing active metadata using the DTW algorithm, the traditional method is to only use the Euclidean distance as the reference distance, and does not consider the data characteristics of the active metadata. Therefore, it is necessary to improve the calculation method of the reference distance when processing active metadata using the DTW algorithm.

[0022] See also Figure 1 , which shows a method flow chart of an intelligent data governance method based on active metadata provided by an embodiment of the present invention, the method comprising the following steps: Step S1, obtaining active metadata of different data sources and performing preprocessing to obtain active metadata of different groups.

[0023] The present invention obtains the complex difference value between two groups of active metadata to be calculated by the DTW algorithm, adaptively selects the calculation method of the reference distance in the DTW algorithm, and obtains a more accurate distance matrix to obtain a more reasonable path matching.

[0024] First, active metadata needs to be obtained, and then preprocessing operations need to be performed on the active metadata. The preprocessing operations include data cleaning, data format conversion, and data grouping.

[0025] Data cleaning includes removing duplicate data, processing missing values, and removing outliers; removing duplicate data: using threshold filtering to check whether there are duplicate records in the data, and removing duplicate data if there are; processing missing values: using linear interpolation to process missing values; removing outliers: using anomaly detection algorithms to detect and remove outliers or outliers in the data (such as KNN). The purpose of data cleaning is to remove noise, duplication, or inconsistency in the data to ensure the quality of active metadata.

[0026] The purpose of data conversion is to facilitate the processing of the DTW algorithm. Metadata from different sources may use different structures (such as JSON, CSV, XML, etc.) and need to be converted into a unified format, such as CSV format, so that the DTW algorithm can process it.

[0027] The purpose of grouping is to group active metadata that has undergone data cleaning and format conversion according to different data sources. The data source of each group of active metadata is the same.

[0028] Step S2: calculating the structural complexity of the two sets of active metadata according to the number of attributes, correlation measurement values ​​and combined entropy values ​​of the two sets of active metadata.

[0029] In the actual operation process of the DTW algorithm, for active metadata with different attributes and different dimensions, when constructing the distance accumulation matrix, the differences in different dimensions will introduce errors in calculating the reference distance. Therefore, we must first consider whether the two sets of active metadata to be subjected to the DTW algorithm contain more attributes and whether their relationship is complex. Then, we obtain the complex difference degree values ​​of the two sets of active metadata to be subjected to the DTW algorithm. According to the complex difference degree values, we adaptively obtain the calculation method of the reference distance in the DTW algorithm to obtain a more accurate distance matrix, find the optimal matching path, and obtain a more accurate correspondence between data.

[0030] In an environment that spans systems, businesses, and data sources, active metadata itself is diverse and complex. Therefore, dimensional information is an important component of active metadata. The increase in dimensionality introduces noise due to excessive redundant features, resulting in a decrease in the accuracy of the reference distance calculation method (Euclidean distance) in the traditional DTW algorithm. Therefore, it is necessary to introduce multi-dimensional feature analysis into the two sets of active metadata to be used for DTW calculation. The three most important dimensions in the multi-dimensional features include the number of attributes, the overall change trend, and the combined entropy value, which comprehensively evaluate the dimensional characteristics of the data from the three perspectives of structure, dynamic characteristics, and uncertainty, respectively, and have a profound impact on the accuracy and stability of the Euclidean distance calculation. Therefore, by comprehensively obtaining the structural complexity of the two sets of active metadata to be used for DTW calculation from these three perspectives, it helps to analyze whether the Euclidean distance calculation may be affected.

[0031] The structural complexity of the two sets of active metadata is calculated based on the number of attributes, correlation measurement values ​​and combined entropy values ​​of the two sets of active metadata. Specifically, the number of attributes of the data in the two sets of active metadata is normalized to obtain the attribute characteristic value; two straight lines are obtained by linear fitting using the two sets of active metadata, and the correlation measurement values ​​of the two sets of active metadata are obtained according to the slopes of the two straight lines; the correlation measurement values ​​are negatively correlated and mapped using an exponential function with a natural constant as the base to obtain the correlation characteristic value; the joint entropy values ​​of the two sets of active metadata are normalized to obtain the entropy characteristic value; the average value of the attribute characteristic value, correlation characteristic value and entropy characteristic value of the two sets of active metadata is the structural complexity of the two sets of active metadata.

[0032] The specific calculation formulas for structural complexity and correlation metrics are: , , in, represents the structural complexity of the i-th group of active metadata and the j-th group of active metadata; Norm() represents the normalization function, which limits the output result to [0,1]; m(∂) represents the number of attributes of the data in the two groups of active metadata (the i-th group of active metadata and the j-th group of active metadata) to be subjected to the DTW algorithm (for example, if one group of active metadata represents time-sales volume and the other group represents time-views, then m(∂) is 3 in these two groups of active metadata); represents the correlation measure between the i-th group of active metadata and the j-th group of active metadata, exp(-) represents the inverse normalization function, exp() represents the exponential function with a natural constant as the base, and represent the slope of the straight line obtained by straight-line fitting using the i-th group of active metadata and the slope of the straight line obtained by straight-line fitting using the j-th group of active metadata, respectively; represents the joint entropy value of the i-th group of active metadata and the j-th group of active metadata. , and They respectively represent the attribute characteristic value, correlation characteristic value and entropy characteristic value of the two sets of active metadata. The larger the three are, the higher the complexity of the two sets of active metadata.

[0033] For the number of attributes of the data in the two sets of active metadata, the DTW algorithm measures their similarity by comparing the "distance" of the two sequences point by point. If the active metadata contains more attributes, the data variation amplitude or scale difference of different attributes may be greater, then the reference distance calculation of the DTW algorithm will be dominated by the high amplitude or larger scale, resulting in unreasonable path matching, and the structural complexity of the two sets of active metadata will be higher.

[0034] The correlation measure of the trends of two sets of active metadata is intended to capture the similarity of the overall trends of the two sets of active metadata. The least squares method is used for straight line fitting, so that the sum of the distances between the fitted straight line and the active metadata is minimized, reflecting the overall trend of a set of active metadata. The absolute value of the difference between the slopes of the two straight lines is The absolute difference between the slopes of the two straight lines. The larger the difference, the more different the trends of the two sets of active metadata are. As a restriction, it ensures that the effect of the difference value will not be magnified when the slope difference is large. , The closer it is to 0, the greater the trend difference between the two sets of active metadata. The DTW algorithm will over-stretch and compress local data points to forcibly align the two sets of active metadata, which may easily lead to unreasonable reference distances.

[0035] The combined entropy value of the two sets of active metadata that will be subjected to the DTW algorithm (the entropy calculation formula is the existing formula) reflects the uncertainty between the two sets of active metadata. The larger the entropy value, the more complex the relationship between the two sets of active metadata and the lower the correlation between the data. In this case, the DTW algorithm will minimize the reference distance as much as possible, making the alignment path more complicated. The calculation result of the reference distance cannot accurately reflect the actual relationship between the data, resulting in a large matching error.

[0036] In summary, The closer it is to 1, the lower the correlation between the two sets of active metadata. The more attributes are included and the less similar the overall change trend is. The more inaccurate the calculation result is when the DTW algorithm calculates the reference distance in the traditional way, the more likely it is to produce an incorrect matching path.

[0037] Step S3, calculating the stationarity measurement value of a group of active metadata according to the average value and the maximum point of the group of active metadata; adding the stationarity measurement value of a group of active metadata and the normalized sampling frequency to obtain the data feature measurement value.

[0038] In step S2, the structural complexity of the two groups of active metadata to be used for DTW calculation is obtained, aiming to analyze the complexity of the two groups of active metadata in terms of attributes, trends and information volume. This is a characteristic of the special scenario problem of the active metadata itself. However, since the DTW algorithm only considers the distance information of a single point pair in the sequence during operation and does not consider the local change characteristics of the active metadata, it is necessary to further analyze the data characteristics of the active metadata in this step to obtain the data feature measurement value of each group of metadata.

[0039] The stationarity measurement value of a group of active metadata is calculated based on the average value and maximum value point of the group of active metadata. Specifically, the absolute value of the difference between each data value and the average value in a group of active metadata is obtained and summed to obtain the degree of dispersion; a maximum value point in the group of active metadata is subtracted from the data adjacent to the left and the data adjacent to the right and summed to obtain the local data change value of the maximum value point; the local data change values ​​of all the maximum value points in the group of active metadata are summed to obtain the fluctuation change characteristic value; the degree of dispersion of the group of active metadata and the sum of the fluctuation change characteristic value are normalized to obtain the stationarity measurement value of the group of active metadata.

[0040] The specific calculation formula is: , in, represents the stability metric of the t-th group of active metadata, which is used to evaluate the data change characteristics of active metadata; tanh( ) represents the hyperbolic tangent function, which limits the output result to (0,1) and is used for normalization; represents the kth active metadata in the tth group of active metadata; represents the average value of the t-th group of active metadata; n represents the number of maximum value points in the t-th group of active metadata; represents the αth maximum point in the tth group of active metadata; and They respectively represent the data points adjacent to the left and right of the αth maximum point in the tth group of active metadata.

[0041] In the above formula, Indicates the degree of dispersion of the t-th group of active metadata. The smaller the result, the closer the value is to the mean and the distribution is more concentrated. On the contrary, the distribution is more dispersed, which may contain more outliers (mutation points). Because these outliers are far away from the mean of the data set, the summation result is amplified. The local data change value of the αth maximum point in the tth group of active metadata is accumulated and summed to obtain the fluctuation change characteristic value of the tth group of active metadata. The purpose is to discover the volatility of the group of data and the local change characteristics of the data by comparing the value difference between the maximum point and the nearest neighbor point. The larger the result, the greater the difference in the value of the maximum point relative to the nearest neighbor point, that is, the data has a higher volatility, and the local peak is more obvious and the fluctuation is violent. On the contrary, it means that the data is relatively stable and the fluctuation is weak. The data feature measurement value reflects the stability of the group of active metadata. When the stability difference between data is large, the Euclidean distance calculation in the traditional way is more likely to be affected by the data with lower stability, and finally produce unreasonable matching paths.

[0042] In addition, the sampling frequency of active metadata determines the time interval between active metadata. If the sampling frequency difference between two sets of active metadata is large, the DTW calculation will be sensitive to the time misalignment, which will cause the calculation of the Euclidean distance (reference distance) to be distorted, resulting in unreasonable matching results and unable to correctly reflect the similarity between the data. Therefore, it is necessary to include the sampling frequency of active metadata in the analysis, combine the stationarity metric with the sampling frequency of active metadata, and add the stationarity metric of a set of active metadata and the normalized sampling frequency to obtain the data feature metric value, which will provide some help for the subsequent judgment of the data metric difference between the two sets of data. The calculation formula of the data feature metric value is: , in, represents the data feature measurement value of the tth group of active metadata, Indicates the sampling frequency of the active metadata set. Thus, the data feature measurement values ​​of the two sets of active metadata to be subjected to the DTW algorithm can be obtained.

[0043] Step S4, obtaining the change feature measurement difference value of the two sets of active metadata based on the difference of the data feature measurement values ​​of the two sets of active metadata; performing weighted summation on the structural complexity and change feature measurement difference value of the two sets of active metadata to obtain the complexity difference value.

[0044] In step S3, the data feature measurement values ​​of the two sets of active metadata are obtained. In order to evaluate the comprehensive difference in the stationarity and sampling frequency of the two sets of active metadata, it is necessary to obtain the absolute value of the difference between the data feature measurement values ​​of the two sets of active metadata and normalize them to obtain the change feature measurement difference value of the two sets of active metadata. The specific calculation formula is: , In the formula, Indicates the difference value of the change feature measurement of the two sets of active metadata (the i-th set of active metadata and the j-th set of active metadata) to be calculated by DTW; The purpose is to evaluate the comprehensive difference in stationarity and sampling frequency between the two sets of data. First, the data feature metric ρ includes the stationarity metric of the active metadata. The DTW algorithm constructs a distance matrix by calculating the distance between each pair of time points. The large stationarity difference between the two sets of active metadata will cause DTW to pay too much attention to changes in the global trend when calculating the reference distance. Some data in the distance matrix will be overestimated, which will affect the subsequent alignment path selection. Secondly, the large sampling frequency difference will also cause the DTW algorithm to try to nonlinearly match these time points during the alignment process due to different time steps, which will mislead the path matching results.

[0045] Therefore, by obtaining the change feature measurement difference value of the two sets of active metadata to be calculated by DTW ,think When is small, the two sets of active metadata have similar stability, similar change trends and volatility, and similar sampling frequencies. Such two sets of active metadata will not produce large errors when using Euclidean distance as the reference distance for DTW algorithm in the traditional way, and will have little effect on the DTW calculation results. On the contrary, When is large, the stationarity differences and sampling frequencies between the data are large, and the traditional method of using Euclidean distance as the reference distance for DTW algorithm may produce large errors and unreasonable path matching.

[0046] Finally, it is necessary to conduct a comprehensive analysis of the structural complexity and change characteristic measurement difference values ​​of the two sets of active metadata. Therefore, the weighted sum of the structural complexity and change characteristic measurement difference values ​​of the two sets of active metadata is used to obtain the complexity difference value. The specific calculation formula is: , in, Indicates the complex difference value between the two sets of active metadata (the i-th set of active metadata and the j-th set of active metadata) to be calculated by DTW; and Respectively represent weight values; It comprehensively reflects the complexity and data difference between the two sets of active metadata. The larger the complexity difference value, the more complex the two sets of active metadata are, the more attributes they contain and the less they have similar overall change trends. At the same time, the stationarity differences and sampling frequencies between the data are large. When DTW uses the Euclidean distance as the reference distance in the traditional way, the calculation results obtained are less accurate, and it is easy to produce wrong matching paths. and The value of needs to be set according to different scenarios and requirements. This solution provides a reference: , .

[0047] Step S5, setting a calculation method of a reference distance of the DTW algorithm, selecting a calculation method of the reference distance according to the complexity difference values ​​of the two sets of active metadata, and processing the two sets of active metadata using the DTW algorithm.

[0048] In step S4, the complexity difference values ​​of the two sets of active metadata that will be subjected to the DTW algorithm are obtained. For the two sets of active metadata with smaller complexity difference values, the method of calculating the reference distance in the traditional DTW algorithm can be used; for the two sets of active metadata with larger and smaller complexity difference values, it is necessary to determine the calculation method of their reference distance.

[0049] Furthermore, the calculation method of the reference distance of the DTW algorithm is set. The calculation method of the reference distance of the DTW algorithm includes two calculation methods, namely the traditional method and the relative change distance calculation method; the traditional method is to use the Euclidean distance as the reference distance; the relative change distance calculation method is: , in, Indicates the relative change distance between the i-th data point and the j-th data point; and Represent the data values ​​of the i-th data point and the i-1-th data point respectively; and Respectively represent the data values ​​of the jth data point and the j-1th data point. It should be noted that the i-th data point and the j-th data point in the embodiment of the present invention are data points in two sequences, that is, the i-th data point of one set of data and the j-th data point of the other set of data in the two sets of active metadata.

[0050] When two sets of data When is larger, since the Euclidean distance is sensitive to the absolute value but ignores the relative change, the relative change distance is introduced to calculate the distance between the current two sets of data, to obtain a more accurate distance matrix and to obtain a more reasonable path matching.

[0051] Using relative change distance as reference distance has the following advantages over traditional Euclidean distance as reference distance: first, Euclidean distance is too dependent on the absolute value of the data; When is large, the dimensional characteristics of active metadata are strong, and the data change characteristics vary greatly. This leads to the fact that large-scale or large-range active metadata dominates the calculation results when obtaining the reference distance through the Euclidean distance calculation method, and the contribution of other dimensions is seriously underestimated. The use of relative change distance eliminates such scale effects because it calculates relative differences, avoids the dimensional disaster caused by multi-dimensional data, and reduces the dependence on absolute numerical differences, so that the influence of each dimension on the distance calculation result is balanced to a certain extent. When larger, it can provide a more reliable reference distance calculation result than the Euclidean distance.

[0052] Specifically, it is necessary to select a calculation method for the reference distance based on the complex difference value of the two sets of active metadata and set a threshold. If the complex difference value of the two sets of active metadata is less than or equal to the threshold, the traditional method is selected to calculate the reference distance; if the complex difference value of the two sets of active metadata is greater than the threshold, the relative change distance calculation method is selected. The setting of the threshold in the embodiment of the present invention needs to be adjusted by the implementer according to the actual situation. The specific environment of the implementation of the scheme is different, and the setting of the threshold is also different. According to the different requirements for matching accuracy or calculation efficiency, if the accuracy requirement is high, a higher threshold can be considered to ensure a more detailed match. If efficiency is prioritized, a smaller threshold can be appropriately selected. In this scheme, a compromise is taken into account, taking into account both accuracy and efficiency, and a reference threshold of 0.7 is given.

[0053] After selecting the calculation method of the reference distance, the DTW algorithm is performed. The present invention adaptively obtains the DTW distance calculation method, constructs a more accurate distance matrix, finds the optimal matching path, obtains a more accurate correspondence between data, and matches accordingly, thereby completing data mapping, improving the accuracy of data matching and mapping, and improving data quality. The intelligent data governance method and system can better automatically perform tasks such as data fusion and cleaning, thereby improving the efficiency and quality of data governance.

[0054] Embodiment 2: This embodiment provides an intelligent data governance system based on active metadata, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of an intelligent data governance method based on active metadata are implemented. Since an intelligent data governance method based on active metadata has been described in detail in Embodiment 1, it will not be described in detail here.

[0055] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The above is a description of a specific embodiment of this specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0056] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0057] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An intelligent data governance method based on active metadata, characterized in that: The method includes: Obtain active metadata of different data sources and perform preprocessing to obtain active metadata of different groups; The structural complexity of the two sets of active metadata is calculated based on the number of attributes, correlation measurement values ​​and combined entropy values ​​of the two sets of active metadata; Calculating a stationary measurement value of a set of active metadata according to an average value and a maximum point of the set of active metadata; obtaining a data feature measurement value according to a stationary measurement value of a set of active metadata and a normalized sampling frequency; Based on the difference of the data characteristic measurement values ​​of the two sets of active metadata, the change characteristic measurement difference value of the two sets of active metadata is obtained; the structural complexity and change characteristic measurement difference value of the two sets of active metadata are weightedly summed to obtain the complexity difference degree value; The calculation method of the reference distance of the DTW algorithm is set, and the calculation method of the reference distance is selected according to the complex difference degree value of the two sets of active metadata, and the two sets of active metadata are processed using the DTW algorithm.

2. The intelligent data governance method based on active metadata according to claim 1, characterized in that: The obtaining of active metadata from different data sources and preprocessing to obtain different groups of active metadata includes: The preprocessing includes data cleaning, data format conversion and data grouping; the active metadata that has undergone data cleaning and format conversion is grouped according to different data sources to obtain different groups of active metadata.

3. The intelligent data governance method based on active metadata according to claim 1 is characterized in that: The calculating the structural complexity of the two sets of active metadata according to the number of attributes, correlation measurement values ​​and combined entropy values ​​of the two sets of active metadata includes: The number of attributes of the data in the two sets of active metadata is normalized to obtain the attribute characteristic value; the two sets of active metadata are used to perform straight line fitting to obtain two straight lines, and the correlation measurement value of the two sets of active metadata is obtained according to the slopes of the two straight lines; the correlation measurement value is negatively correlated with the exponential function with the natural constant as the base to obtain the correlation characteristic value; the joint entropy value of the two sets of active metadata is normalized to obtain the entropy characteristic value; the average value of the attribute characteristic value, correlation characteristic value and entropy characteristic value of the two sets of active metadata is the structural complexity of the two sets of active metadata.

4. The method for intelligent data governance based on active metadata according to claim 3, characterized in that: The calculation formula of the correlation metric value is: , in, represents the correlation metric value between the i-th group of active metadata and the j-th group of active metadata; and They respectively represent the slope of the straight line obtained by straight-line fitting using the i-th group of active metadata and the slope of the straight line obtained by straight-line fitting using the j-th group of active metadata.

5. The intelligent data governance method based on active metadata according to claim 1 is characterized in that: The step of calculating the stationarity measurement value of a group of active metadata according to an average value and a maximum value point of the group of active metadata includes: The absolute value of the difference between each data value and the average value in a set of active metadata is obtained and summed to obtain the degree of dispersion; a maximum point in the set of active metadata is subtracted from the data adjacent to the left and the data adjacent to the right and summed to obtain the local data change value of the maximum point; the local data change values ​​of all the maximum points in the set of active metadata are summed to obtain the fluctuation change characteristic value; the degree of dispersion of the set of active metadata and the sum of the fluctuation change characteristic values ​​are normalized to obtain the stability measurement value of the set of active metadata.

6. The method for intelligent data governance based on active metadata according to claim 1, characterized in that: The difference between the data feature measurement values ​​of the two sets of active metadata is used to obtain the change feature measurement difference value of the two sets of active metadata, including: The absolute value of the difference between the data feature measurement values ​​of the two sets of active metadata is normalized to obtain the change feature measurement difference value of the two sets of active metadata.

7. The method for intelligent data governance based on active metadata according to claim 1, characterized in that: The calculation method of setting the reference distance of the DTW algorithm includes: The reference distance calculation method of the DTW algorithm includes two calculation methods, namely the traditional method and the relative change distance calculation method; the traditional method uses the Euclidean distance as the reference distance; the relative change distance calculation method is: , in, Indicates the relative change distance between the i-th data point and the j-th data point; and Represent the data values ​​of the i-th data point and the i-1-th data point respectively; and Represent the data values ​​of the j-th data point and the j-1-th data point respectively.

8. The method for intelligent data governance based on active metadata according to claim 1, characterized in that: The method of selecting the calculation method of the reference distance according to the complex difference degree values ​​of the two sets of active metadata and processing the two sets of active metadata using the DTW algorithm includes: A threshold is set. If the complexity difference value of the two sets of active metadata is less than or equal to the threshold, the traditional method is selected to calculate the reference distance; if the complexity difference value of the two sets of active metadata is greater than the threshold, the relative change distance calculation method is selected.

9. The method of intelligent data governance based on active metadata according to claim 1, characterized in that: The obtaining of the data characteristic measurement value includes: The data feature metric is obtained by adding the stationarity metric of a set of active metadata and the normalized sampling frequency.

10. An intelligent data governance system based on active metadata, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of an intelligent data governance method based on active metadata as described in any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Data management method based on data weaving architecture

    CN116303336A

  • Multi-platform metadata standardization processing method and device

    CN117851387A

  • Metering parameter intelligent cleaning method and system based on carrier chip with MCU

    CN118035660A

  • Systems and methods of detecting and responding to malware on a file system

    US20180048658A1

  • Consilence of data-mining

    WO2007147166A2