Metadata-driven data-in-data management method and system

By introducing an adaptive metadata evolution mechanism and a multi-level data quality warning and repair mechanism, the shortcomings of the existing data middle platform in metadata management, data quality control and data processing process optimization are solved, and efficient, flexible and intelligent data middle platform management is achieved.

CN120030076AActive Publication Date: 2025-05-23GUANGZHOU TAIXIN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510150017.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-23
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing data middle platform has shortcomings in metadata management, data quality control and data processing process optimization, resulting in low data processing efficiency and quality. Especially when facing changing business needs and diversified data sources, the system seems to be inflexible, precise and efficient enough.

Method used

Adaptive metadata evolution mechanism and multi-level data quality early warning and repair mechanism are introduced, and the data field change patterns are identified through machine learning, the metadata model is automatically adjusted, data quality is monitored in real time, early warning and repair mechanisms are promptly triggered, and data processing flow is optimized.

Benefits of technology

Real-time update and adaptability enhancement of metadata models are realized, the monitoring and repair efficiency of data quality is improved, the automation and efficiency of data processing is improved, and the flexibility and intelligence level of the data middle platform are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030076A_ABST
    Figure CN120030076A_ABST
Patent Text Reader

Abstract

The invention provides a data-in-data management method and system based on metadata driving. The method comprises the following steps: connecting to a plurality of heterogeneous data sources through a plurality of access modes; automatically adjusting and optimizing the metadata model; using a machine learning method to identify the rule of each field change in the historical data, and generating an evolution rule of metadata evolution; automatically feeding back the difference between the real-time monitoring data source and the adaptive metadata model; monitoring the quality of a data source in real time by using the adaptive metadata model, designing an early warning rule according to the quality of the data source, and generating an early warning rule set; when the data quality problem triggers early warning; a machine learning-based repair model for a data quality problem is designed, and a proper repair method is selected for optimization. According to the invention, by introducing an adaptive metadata evolution mechanism and a multi-level data quality early warning and repairing mechanism, the defects in the prior art are overcome, and an efficient, flexible and intelligent data center management solution is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of metadata drive, and in particular, to a data management method and system for a data middle platform based on metadata drive. Background Art

[0002] With the rapid development of informatization and digitalization, data has become a core asset for modern enterprise decision-making and operations. In order to manage and utilize data more efficiently, more and more companies have begun to adopt data middle-end architecture to integrate and manage massive data scattered in different business systems. The data middle-end brings together various data resources of the enterprise through a unified data management platform, providing one-stop data processing, analysis and service capabilities to support business decision-making and innovation. However, although the data middle-end has great potential in data integration and application, the construction and application of most data middle-ends currently still face many challenges, especially in metadata management, data quality control and processing flow optimization, resulting in serious deficiencies in data processing efficiency and quality of existing technologies.

[0003] First, most of the existing data middle platforms rely on static metadata models. Metadata defines the structure, type, and relationship of data and is the core of data management. However, traditional metadata models are usually manually formulated during the initial construction, and it is often difficult to keep up with changes in demand in subsequent business changes and data source updates, resulting in lags and inconsistencies in metadata. For example, when new business needs emerge or new data sources are added, the existing metadata model may not be updated effectively in a timely manner, which not only leads to a disconnect between the data source and the metadata model, but may also affect subsequent data integration, cleaning, and analysis work, thereby reducing the overall processing efficiency of the data middle platform.

[0004] Secondly, data quality issues have always been a major pain point in data center management. Data quality involves multiple aspects such as data accuracy, completeness, consistency, and timeliness. In traditional data processing processes, data quality management mostly relies on manually set static rules or predefined quality inspection standards. Although some technical means can achieve basic monitoring of data quality, such as data cleaning and verification, these means often appear to be inadequate when faced with large-scale, multi-source heterogeneous data. In particular, when data quality is abnormal, traditional methods often lack a real-time automatic repair mechanism, resulting in a backlog of problems or delayed repairs, which not only affects the reliability of data quality, but also increases the cost and risk of manual intervention.

[0005] Finally, the existing data middle platform system also has deficiencies in the automation and intelligence of data processing. At present, most systems rely on fixed ETL (extraction, transformation, loading) processes for data integration, which are generally pre-configured and lack flexible adaptive capabilities. As the business changes, the traditional data processing process may not be adjusted in time, resulting in bottlenecks, repetitive operations and inefficient processing in the process, which in turn affects the real-time and efficiency of data processing. In addition, the real-time monitoring and intelligent decision-making mechanism of data quality have not yet been fully realized, and the lack of automatic adjustment and optimization capabilities based on dynamic feedback has further increased the complexity and difficulty of data middle platform management.

[0006] In summary, existing technologies have varying degrees of deficiencies in metadata management, data quality control, and intelligent optimization of data processing processes. In particular, when faced with ever-changing business needs and diversified data sources, traditional systems are often not flexible, accurate, and efficient enough. Therefore, a new technical solution is urgently needed to solve these problems, so that the data center can operate more intelligently and efficiently to meet rapidly changing business needs and high-quality data management requirements. Summary of the invention

[0007] The purpose of the present invention is to propose a metadata-driven data middle-office data management method and system, which overcomes the shortcomings of the prior art by introducing an adaptive metadata evolution mechanism and a multi-level data quality warning and repair mechanism, and provides an efficient, flexible and intelligent data middle-office management solution.

[0008] In order to achieve the above object, a first aspect of the present invention provides a metadata-driven data management method for a data center, the method comprising the following steps: Connect to various heterogeneous data sources through various access methods, obtain raw data, and use the preset preliminary metadata model to perform preliminary analysis on the raw data; Automatically adjust and optimize the metadata model according to the historical flow pattern, change rules and business needs of the original data to obtain an adaptive metadata model after adaptive evolution; Using machine learning methods to identify the rules of changes in various fields in historical data and generate evolution rules for metadata evolution; wherein the optimization goal of the evolution rules is to maximize the prediction accuracy and adaptability of the adaptive metadata model; The differences between the data source and the adaptive metadata model will be monitored in real time. When deviations in the adaptive metadata model are detected, automatic feedback will be provided and the adaptive metadata model will be adjusted. The adaptive metadata model is used to monitor the quality of the data source in real time, and warning rules are designed according to the quality of the data source to generate a warning rule set; the warning rules include: When the data quality of a field exceeds the preset tolerance range, an alarm is automatically triggered. When the warning rule is generated, the field weights are assigned different weight values ​​according to the importance of the field through the business weights in the metadata model. At the same time, according to the quality assessment results of each field, the most serious data problems will be monitored first. If the missing value of a field exceeds 30% and it is a core business field, an alarm will be triggered immediately. When a data quality problem triggers an early warning, the data will be repaired through the feedback mechanism and a repair feedback data set will be output; the specific repair method is determined by the type of field and the degree of abnormality of the data: If a field is missing, its historical data trend will be checked first, and the missing data will be filled based on the trend or average value; If the value of a field is outside the predetermined reasonable range, an attempt will be made to correct it; For time series data, time windows and historical data will be used for prediction to ensure data consistency; Design a machine learning-based repair model for data quality issues based on the repair feedback dataset, select appropriate repair methods for different field types and specific data repair methods, and generate the final repaired dataset; The machine learning-based repair model designed for data quality issues includes: If a field is missing, a multi-level adaptive repair mechanism is designed to fill it in, including: If the data of a certain field is missing, an appropriate regression model or neural network will be selected for prediction and filling based on the correlation between fields and the distribution of data; If outliers occur, a variational autoencoder model based on deep learning is introduced to repair the anomaly according to the deep structure of the data; For time series data, time series modeling is used to correct the time field to ensure the consistency of the data in time series.

[0009] Furthermore, the multiple access methods include API, database connection, and file import; the type of the data source is one of a relational database, a non-relational database, a file system, and a data lake.

[0010] Furthermore, the preliminary metadata model includes structural information, field names, field types, and field relationships of the data source.

[0011] Furthermore, during the evolution of the preliminary metadata model, the change rules of each field are captured in real time, and the metadata model is dynamically adjusted using the following formula: ; in, Represents the evolved metadata model, recording changes in fields and their relationships; is the initial metadata model; Indicates a field at a point in time Changes in is the change weight coefficient, which indicates the influence of changes in different fields; n is the total number of fields.

[0012] Furthermore, the evolution rule set of the metadata evolution includes rules for adjusting and optimizing the metadata model, which is expressed as ,in, It is Evolution rules that describe how to handle changes in field types or adjustments to relationships; Indicates the time when the rule is generated. The evolutionary rules will be updated in time according to the actual data changes; When the adaptive metadata model is detected to have a deviation, automatic feedback is performed and the adaptive metadata model is adjusted, which is expressed as: Assume that in the feedback mechanism, based on changes in business requirements or data quality issues, the final adjustment of the metadata model is as follows: ; in, It is the final adjusted metadata model; It is a metadata model obtained through adaptive evolution; It is an adjustment item based on real-time data feedback.

[0013] Furthermore, in the process of using the adaptive metadata model to monitor the quality of the data source in real time, an adaptive quality assessment algorithm is introduced to dynamically adjust the quality assessment standard according to different data sources and business requirements, which is expressed as: ; in, Indicates the quality of the data source, including quality information of each field, including indicators such as completeness, inconsistency, and accuracy. The format is ,in Representative Quality assessment results of fields; is an evaluation function that combines field data , metadata model and quality standards To calculate the quality of each field; It is an adaptive metadata model; is the quality standard for each field.

[0014] Furthermore, the generation process of the warning rule is as follows: ; in, It is Warning rules; It is the warning generation function, combined with field data , Field Weight and threshold To generate warning rules; Is Field The importance weight of is the quality threshold of the field; The data repair through the feedback mechanism is an intelligent repair based on historical trends and machine learning, which is expressed as: ; in, is the repair feedback dataset; It is the data after quality assessment; is the increment of repair data.

[0015] Furthermore, further repair is performed on the repair feedback dataset, which is expressed as: ; in, It is the final repaired dataset, which contains all the repaired data; It is a repair feedback dataset, which contains the results of quality assessment and preliminary repair; It indicates the repair increment after optimization, that is, the part that is further repaired.

[0016] Furthermore, a regularization term is introduced into the adaptive weighted regression model. , to prevent overfitting, expressed as: ; in, is a regularized loss function used to constrain model parameters; is the regularization coefficient, which controls the strength of regularization; The first The weight of the parameter.

[0017] Combined with regularization term , the goal of the adaptive weighted regression model is to minimize the sum of the prediction error and the regularization loss, the loss function It is expressed as: ; in, is the missing value of the prediction; is the actual missing value; The loss function of the variational autoencoder model is expressed as: ; in, is the loss function of the variational autoencoder; is the distribution of latent variables generated by the encoder; is the distribution of sample data generated by the decoder; is the KL divergence, which measures the difference between the encoder distribution and the prior distribution.

[0018] In a second aspect of the present invention, a metadata-driven data middle platform data management system is provided, the system comprising: A data source access unit is used to connect to various heterogeneous data sources through various access methods, obtain raw data, and perform preliminary analysis on the raw data using a preset preliminary metadata model; The meta-model preliminary construction unit is used to automatically adjust and optimize the metadata model according to the historical flow pattern, change rules and business needs of the original data, and obtain an adaptive metadata model after adaptive evolution; A metadata evolution unit, used to use a machine learning method to identify the law of changes in various fields in historical data and generate evolution rules for metadata evolution; wherein the optimization goal of the evolution rules is to maximize the prediction accuracy and adaptability of the adaptive metadata model; The meta-model adjustment unit is used to monitor the difference between the data source and the adaptive metadata model in real time, and when a deviation of the adaptive metadata model is detected, automatically provide feedback and adjust the adaptive metadata model; The metadata warning unit is used to use the adaptive metadata model to monitor the quality of the data source in real time, and to design warning rules according to the quality of the data source to generate a warning rule set; the warning rules include: When the data quality of a field exceeds the preset tolerance range, an alarm is automatically triggered. When the warning rule is generated, the field weights are assigned different weight values ​​according to the importance of the field through the business weights in the metadata model. At the same time, according to the quality assessment results of each field, the most serious data problems will be monitored first. If the missing value of a field exceeds 30% and it is a core business field, an alarm will be triggered immediately. The metadata repair unit is used to repair the data through the feedback mechanism and output the repair feedback data set when the data quality problem triggers an early warning. The specific repair method is determined according to the type of field and the degree of abnormality of the data: If a field is missing, its historical data trend will be checked first, and the missing data will be filled based on the trend or average value; If the value of a field is outside the predetermined reasonable range, an attempt will be made to correct it; For time series data, time windows and historical data will be used for prediction to ensure data consistency; The metadata repair unit is used to design a machine learning-based repair model for data quality issues based on the repair feedback data set, select appropriate repair methods for different field types and specific data repair methods, and generate the final repaired data set; The machine learning-based repair model designed for data quality issues includes: If a field is missing, a multi-level adaptive repair mechanism is designed to fill it in, including: If the data of a certain field is missing, an appropriate regression model or neural network will be selected for prediction and filling based on the correlation between fields and the distribution of data; If outliers occur, a variational autoencoder model based on deep learning is introduced to repair the anomaly according to the deep structure of the data; The data correction unit is used to correct the time field of time series data using time series modeling to ensure the consistency of data in time series.

[0019] The beneficial technical effects of the present invention are at least as follows: The present invention solves the problem of traditional metadata lag by introducing an adaptive metadata evolution mechanism. The mechanism can automatically derive and update the metadata model according to changes in data flow and business requirements, so that it always keeps in sync with the data source and business requirements. This mechanism not only reduces manual intervention, but also can timely adjust the metadata structure and rules when new data sources are added or business requirements change, avoiding the disconnection between data and metadata models and improving the efficiency and accuracy of data processing.

[0020] The present invention solves the limitations of traditional static data quality management through a multi-level data quality management system. Under this mechanism, the system can monitor data quality in real time at the field level, record level, and data set level, and detect problems such as missing values, outliers, and duplicate data in the data. At the same time, when data quality problems are detected, the system can automatically trigger a repair mechanism to automatically fill in missing values, correct outliers, remove duplicate data, etc. according to preset rules to ensure data accuracy and consistency. This mechanism greatly reduces manual intervention and improves the efficiency and automation level of data quality repair.

[0021] Through the intelligent decision-making module, the present invention can automatically optimize the data flow path and data cleaning rules according to the real-time feedback in the data processing process, thereby improving the automation and efficiency of data processing. The intelligent decision-making module can flexibly adjust the data integration and processing flow according to the real-time feedback, optimize the ETL process, reduce unnecessary repeated operations and bottlenecks in data processing, thereby improving the overall data processing speed and accuracy.

[0022] To sum up, the present invention effectively solves the problems of metadata lag, insufficient data quality management and inefficient data processing automation in the prior art through an adaptive metadata evolution mechanism and a real-time multi-level data quality warning and repair mechanism, improves the flexibility, intelligence level and processing efficiency of the data middle platform, and thus provides enterprises with more efficient and reliable data management capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0024] Figure 1 This is a flow chart of the data management method of the intelligent data middle platform driven by metadata of the present invention. DETAILED DESCRIPTION

[0025] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0026] like Figure 1 As shown, the metadata-driven intelligent data middle platform data management method provided by the embodiment of the present invention includes: 1011. Connect to various heterogeneous data sources through various access methods, obtain raw data, and use the preset preliminary metadata model to perform preliminary analysis on the raw data.

[0027] Specifically, in this step, we first connect to various heterogeneous data sources through various access methods (such as API, database connection, file import, etc.) to obtain the raw data. The data source type can be a relational database, a non-relational database, a file system, a data lake, etc.

[0028] Perform a preliminary analysis of the connected data source, focusing on the basic attributes of the data, such as field name, data type, data structure, etc. This information will be converted into a preliminary metadata model. Through the data metadata structure extraction process, it is possible to clarify the field type, the relationship between fields, and whether there are redundant or missing fields.

[0029] Output is a preliminary metadata model , which records the structural information, field names, field types, field relationships, etc. of the original data source.

[0030] For example, suppose that when accessing a data source, the system recognizes that it contains the following fields: Field 1: user_id (type: integer) Field 2: timestamp (type: timestamp) Field 3: order_amount (type: float) This field information will build a preliminary metadata model, including field names, data types, and relationships between fields.

[0031] 1012. Automatically adjust and optimize the metadata model according to the historical flow pattern, change rules and business needs of the original data to obtain an adaptive metadata model after adaptive evolution; Specifically, because data may change at different points in time as the external environment and business needs change, the initial metadata model will usually have certain deviations or incompleteness.

[0032] Dynamic data flow analysis: Based on historical data flow, the system can identify the pattern of field changes. For example, if a field has different names or data type changes in multiple data sources, the system can determine the dynamic mapping relationship between fields based on pattern recognition.

[0033] Model evolution rule generation: By continuously tracking historical data and change trends, the system will automatically generate rules for metadata evolution. For example, if a field order_amount often changes its data type from integer to floating value, the system will generate a rule for dynamic data type adjustment for the field.

[0034] The evolution mechanism not only focuses on changes in data structure, but also infers and adapts future data structures based on changes in business needs and the cycle of data updates. For example, if a business scenario will add new fields or change the data model in the future, the system will predict the evolution trend of the data structure based on past data flow patterns and business needs.

[0035] Metadata model after adaptive evolution of output variables , which has been adjusted according to the changing patterns of data sources and business needs.

[0036] Furthermore, during the evolution of the metadata model, the system will capture the changing rules of each field in real time and use the following formula to dynamically adjust the metadata model: ; in, Represents the evolved metadata model, recording changes in fields and their relationships; is the initial metadata model; Indicates a field at a point in time The changes may be changes in field types, addition of fields, deletion of fields, etc. is the change weight coefficient, which indicates the influence of changes in different fields and is dynamically adjusted based on factors such as the frequency of historical data and the importance of changes. Part of the innovation is the introduction of weighted variation coefficient , which can reflect the impact of changes in different fields on the final model adjustment. This weighting coefficient is not only based on the frequency of field changes, but also takes into account the priority of business needs. For example, a field may be a core business field, which will be given a higher weight when it changes, thus affecting the focus of model evolution.

[0037] 1013. Use machine learning methods to identify the patterns of changes in various fields in historical data and generate evolution rules for metadata evolution; wherein the optimization goal of the evolution rules is to maximize the prediction accuracy and adaptability of the adaptive metadata model.

[0038] Specifically, the evolution rules are not static, but are optimized according to new data changes during the continuous operation of the system. Each data change will trigger the system to feedback and adjust the existing rules. For example, if a field does not achieve the expected effect under the past rule adjustment, the system will automatically learn such mismatches and update the evolution rules.

[0039] The optimization goal of the evolutionary rules is to maximize the prediction accuracy and adaptability of the data model. By building a global optimization algorithm, the system can adjust the rules based on feedback from a large amount of historical data to improve the adaptability of the model.

[0040] Output variable: Evolution rule set , contains rules for tuning and optimizing metadata models. Evolutionary rule sets can be represented as a set of dynamically updated rules:

[0041] in, It is Evolution rules that describe how to handle changes in field types or adjustments to relationships; Indicates the time when the rule is generated. The evolutionary rule will be updated in time according to the actual data changes. In the formula, the evolutionary rule Continuous updates in the time dimension reflect the dynamic adaptability of the rules. This is an innovative mechanism that adjusts the rules in real time based on changes in data flow and business needs. The continuous learning system can effectively avoid the problem of traditional rules becoming outdated or not meeting actual needs.

[0042] In addition, the dynamic weights and timing optimization of rules (such as weighted timing loss function) make the generation and optimization of rules more in line with business goals and data change laws.

[0043] 1014. The difference between the data source and the adaptive metadata model will be monitored in real time. When a deviation of the adaptive metadata model is detected, automatic feedback will be provided and the adaptive metadata model will be adjusted.

[0044] Specifically, a feedback and repair algorithm is introduced in this step. Based on the changes in the new data, feedback update items are automatically generated, and the metadata model is adjusted in real time to reduce the difference between the model and the actual data.

[0045] The output variable is the final metadata model after real-time feedback adjustment , ensuring a high degree of match with the actual data environment.

[0046] Assume that in the feedback mechanism, based on changes in business requirements or data quality issues, the final adjustment of the metadata model is as follows:

[0047] in, It is the final adjusted metadata model; It is a metadata model obtained through adaptive evolution; It is an adjustment item based on real-time data feedback.

[0048] The innovation of the feedback and repair algorithm lies in the introduction of adaptive adjustment and feedback items driven by business needs. It is not just a simple supplement based on data changes, but also combines the dynamic changes of business needs to ensure that the data model is not only technically accurate but also meets the needs of business goals.

[0049] 1015. Use the adaptive metadata model to monitor the quality of the data source in real time, and design warning rules according to the quality of the data source to generate a warning rule set.

[0050] Specifically, data quality assessment and monitoring is the basis for ensuring the stable operation of the data processing system. In this step, the system will use the adaptive metadata model generated in the previous step ( ) to monitor the quality of the data, focusing on the following dimensions: Data integrity: Based on ,The system first evaluates the integrity of each data field, checking whether there are missing values, duplicate values, etc. In the metadata model, each field has a predefined "completeness standard", for example, the missing value ratio of a field cannot exceed 30%.

[0051] Data consistency: Based on business logic, the system determines whether the change in field value is in line with expectations. For example, the data of the time field timestamp should be strictly increased in chronological order. If the time value is found to be reversed or jumped, it is considered that there is a consistency problem in the field.

[0052] Data accuracy: Accuracy is a criterion for evaluating whether data matches its actual object or historical record. If a field value deviates from the normal range, it will be marked as abnormal by comparing historical data.

[0053] The assessment of these quality checks is based on the comparison of data with the metadata model, using the field information, data relationships and business logic in the model to determine whether the quality meets the standards.

[0054] The output variable is , contains quality information for each field, including indicators such as completeness, inconsistency, and accuracy, in the format of ,in Representative The quality assessment results of the fields.

[0055] Furthermore, in this step, an adaptive quality assessment algorithm is introduced, which can dynamically adjust the quality assessment criteria according to different data sources and business needs, greatly enhancing the adaptability of the system in complex and dynamic environments. Traditional methods usually have fixed quality standards, which makes them difficult to adapt to rapidly changing business environments, while the system of the present invention can intelligently adjust the quality assessment rules. The quality assessment process based on the metadata model can be expressed by the following formula: ; in, is an evaluation function that combines field data , metadata model and quality standards To calculate the quality of each field. It is an adaptive metadata model that contains information such as field definitions, data types, and standard values. It is the quality standard for each field, which may include missing value threshold, outlier range, time series standard, etc.

[0056] Furthermore, based on the assessed quality indicators, the system will intelligently generate a set of early warning rules. When the data quality of a field exceeds the preset tolerance range, the system will automatically trigger an alarm. When generating early warning rules, the system will combine the following factors: Field importance: Through the business weight in the metadata model, the importance of the field is assigned different weight values. For example, order_amount is a key field, and its abnormal impact is greater, so it has a higher weight.

[0057] Severity of quality issues: Based on the quality assessment results of each field, data issues with higher severity will be monitored first. If the missing value of a field exceeds 30% and it is a core business field, the system should immediately trigger an alert.

[0058] In addition, the generation of early warning rules not only depends on the current data quality assessment, but also dynamically adjusts based on the trend changes of historical data. Through historical data analysis, the system can determine which fields have had quality problems in the past and generate corresponding early warning mechanisms.

[0059] Furthermore, the output variable represents the generated warning rule set , which contains multiple rules, each of which can define the threshold, impact weight and trigger condition of a quality issue.

[0060] This solution innovatively introduces a dynamic warning generation mechanism based on field weights. Compared with the traditional monitoring system based on fixed rules, this system is more flexible and can automatically adjust the warning strategy according to the importance of the business, reducing unnecessary warnings and improving the accuracy and effectiveness of monitoring. The generation process of warning rules is as follows: ; in, It is Warning rules; It is the warning generation function, combined with field data , Field Weight and threshold To generate warning rules; Is Field The importance weight of comes from the metadata model; It is the quality threshold of the field. For example, an alert is triggered when the missing value ratio exceeds 30%.

[0061] 1016. When data quality issues trigger an early warning, data repair will be performed through the feedback mechanism and a repair feedback data set will be output.

[0062] Specifically, once a data quality problem triggers an early warning, the system will repair the data through a feedback mechanism. The specific repair method depends on the type of field and the degree of data abnormality: Missing value repair: If a field is missing, the system will first check its historical data trend and fill in the missing data based on the trend or average value. For example, if the average value of order_amount in the past month is , then missing data may be filled with .

[0063] Outlier repair: If the value of a field exceeds the predetermined reasonable range, the system will try to correct it. For example, if an order_amount exceeds the maximum preset value (such as ), the system will use the trained model to predict the correction value.

[0064] Consistency repair: For time series data, the system will use time windows and historical data for prediction to ensure data consistency. For example, if the timestamp field has a reverse order problem, the system will re-sort it and use the trends of the previous and next time points to fill in the data.

[0065] Furthermore, the repaired data will re-enter the quality assessment process to ensure that the data quality is restored to a normal range, and will eventually be provided to the data processing module for continued use.

[0066] The output variable is : Represents the repaired dataset, which includes all data quality issues repaired through the early warning and feedback mechanism.

[0067] The innovation of this step is that the repair strategy does not rely solely on simple interpolation or mean filling, but introduces an intelligent repair mechanism based on historical trends and machine learning, which is particularly suitable for high-quality data repair in a dynamic business environment. The data repair process can be expressed by the following formula: ; in, It is the repaired data; It is the data after quality assessment; It is the increment of repair data, including filling missing values, correcting outliers, etc.

[0068] 1017. Design a machine learning-based repair model for data quality issues based on the repair feedback dataset, select appropriate repair methods for optimization according to different field types and specific data repair methods, and generate the final repaired dataset.

[0069] Specifically, the core task of data repair is to select appropriate repair methods for different types of data quality issues through intelligent mechanisms. These quality issues usually include but are not limited to: missing values, outliers, duplicate data, inconsistent data, etc.

[0070] Missing value repair: A multi-level adaptive repair mechanism is designed for missing data. Specifically, if the data of a field is missing, the system will select an appropriate regression model or neural network for prediction and filling based on the correlation between fields and the distribution of data.

[0071] Outlier repair: In addition to traditional anomaly detection methods (such as isolation forest, K nearest neighbor, etc.), this paper also introduces a variational autoencoder (VAE) model based on deep learning. This model can not only effectively detect outliers, but also repair anomalies based on the deep structure of the data.

[0072] Consistency repair: If the data has time series inconsistencies (for example, the timestamp field order is abnormal), the system will use time series modeling (such as LSTM, ARIMA, etc.) to correct the time field to ensure the consistency of the data in time series.

[0073] Furthermore, based on these repair strategies, the goal of the present invention is to ultimately output a high-quality dataset that has been fully repaired and optimized. , this data set will be able to provide a reliable basis for subsequent data analysis, prediction and decision-making.

[0074] Further, Formula 1: Formula representation of the data repair process, the repaired dataset can be expressed as: ; in, Represents the final repaired dataset, which contains all the repaired data. Represents the preliminary restoration dataset, which contains the results of quality assessment and preliminary restoration. It indicates the repair increment after optimization, that is, the part that is further repaired, such as filling missing values, correcting outliers, etc.

[0075] Furthermore, in order to achieve efficient repair, the present invention adopts a series of machine learning algorithms for data quality issues, which can automatically select the most appropriate repair model according to the data type and quality issues: Missing value repair model: The present invention adopts an adaptive weighted regression model. The model calculates the weight coefficient based on the relationship between the missing field and other related fields, and performs weighted prediction on the missing value. On this basis, the present invention introduces a regularization term to prevent overfitting and ensure the generalization ability of the model.

[0076] Outlier detection and repair: In terms of outlier repair, the traditional isolation forest method may ignore the deep features of the data, so this paper introduces variational autoencoder (VAE) as an outlier repair model. VAE can learn the distribution from the latent space of the data and repair outliers based on the learned latent structure.

[0077] Consistency repair: To repair the inconsistency problem of time series, the present invention adopts a method based on long short-term memory network (LSTM), which can learn the long-term dependency of time fields, thereby ensuring the sequential consistency of time fields.

[0078] Furthermore, adaptive weighted regression and regularization repair model: In order to improve the accuracy of missing value repair, the present invention designs a weighted regression model, which dynamically calculates the weight according to the correlation between each data point and other points, and uses regularization terms to prevent overfitting. The regularization terms designed by the present invention are as follows: ; in, Represents the regularized loss function, which is used to constrain model parameters. Represents the regularization coefficient, which controls the strength of regularization. Indicates the model The weight of the parameter.

[0079] Objective function: Combined with the regularization term, the ultimate goal is to minimize the sum of the prediction error and the regularization loss: ; in, Represents missing values ​​for the prediction. Represents the actual missing value (used to calculate the error during training). This repair method can dynamically adjust the weights according to the diversity of the data, thereby effectively handling different data quality issues.

[0080] In terms of outlier repair, a variational autoencoder (VAE) is used as a repair model. VAE can effectively detect and repair outliers by learning the latent space of data. The present invention uses the following model: ; in, Represents the loss function of the variational autoencoder. Represents the distribution of latent variables generated by the encoder. Represents the distribution of sample data generated by the decoder. represents the KL divergence, which measures the difference between the encoder distribution and the prior distribution.

[0081] 1018. For time series data, time series modeling is used to correct the time field to ensure the consistency of the data in time series.

[0082] Specifically, once the data is repaired and an optimized dataset is generated , the system will enter the stage of feedback and self-optimization. In order to ensure the long-term effectiveness of the repair process, the system will automatically perform quality feedback and model updates to form a continuous optimization mechanism. The repaired data set will enter the automatic quality assessment system, which will automatically assess whether the data quality meets the expected standards. If the data quality still does not meet the standards, the system will automatically select other repair models or update the existing model.

[0083] Furthermore, the repair model Continuous training and adjustment will be performed based on new data feedback. Whenever new data is repaired and fed back, the repair model will be updated to adapt to the new data characteristics.

[0084] The iterative update of the repair model can be expressed as: ; in, Represents the updated repair model. Indicates the current repair model. Represents the model update increment, which adjusts the model parameters according to the newly repaired feedback data.

[0085] The embodiment of the present invention further provides a metadata-driven data middle platform data management system, characterized in that the system includes: A data source access unit is used to connect to various heterogeneous data sources through various access methods, obtain raw data, and perform preliminary analysis on the raw data using a preset preliminary metadata model; The meta-model preliminary construction unit is used to automatically adjust and optimize the metadata model according to the historical flow pattern, change rules and business needs of the original data to obtain an adaptive metadata model after adaptive evolution; A metadata evolution unit, used to use a machine learning method to identify the law of changes in various fields in historical data and generate evolution rules for metadata evolution; wherein the optimization goal of the evolution rules is to maximize the prediction accuracy and adaptability of the adaptive metadata model; The meta-model adjustment unit is used to monitor the difference between the data source and the adaptive metadata model in real time, and when a deviation of the adaptive metadata model is detected, automatically provide feedback and adjust the adaptive metadata model; The metadata warning unit is used to use the adaptive metadata model to monitor the quality of the data source in real time, and to design warning rules according to the quality of the data source to generate a warning rule set; the warning rules include: When the data quality of a field exceeds the preset tolerance range, an alarm is automatically triggered. When the warning rule is generated, the field weights are assigned different weight values ​​according to the importance of the field through the business weights in the metadata model. At the same time, according to the quality assessment results of each field, the most serious data problems will be monitored first. If the missing value of a field exceeds 30% and it is a core business field, an alarm will be triggered immediately. The metadata repair unit is used to repair the data through the feedback mechanism and output the repair feedback data set when the data quality problem triggers an early warning. The specific repair method is determined according to the type of field and the degree of abnormality of the data: If a field is missing, its historical data trend will be checked first, and the missing data will be filled based on the trend or average value; If the value of a field is outside the predetermined reasonable range, an attempt will be made to correct it; For time series data, time windows and historical data will be used for prediction to ensure data consistency; The metadata repair unit is used to design a machine learning-based repair model for data quality issues based on the repair feedback data set, select appropriate repair methods for different field types and specific data repair methods, and generate the final repaired data set; The machine learning-based repair model designed for data quality issues includes: If a field is missing, a multi-level adaptive repair mechanism is designed to fill it in, including: If the data of a certain field is missing, an appropriate regression model or neural network will be selected for prediction and filling based on the correlation between fields and the distribution of data; If outliers occur, a variational autoencoder model based on deep learning is introduced to repair the anomaly according to the deep structure of the data; The data correction unit is used to correct the time field of time series data using time series modeling to ensure the consistency of data in time series.

[0086] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0087] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a division of logical functions. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0088] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0089] Although embodiments of the present invention have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A data management method based on metadata-driven data middle platform, characterized in that: The method comprises the following steps: Connect to various heterogeneous data sources through various access methods, obtain raw data, and use the preset preliminary metadata model to perform preliminary analysis on the raw data; Automatically adjust and optimize the metadata model according to the historical flow pattern, change rules and business needs of the original data to obtain an adaptive metadata model after adaptive evolution; Using machine learning methods to identify the rules of changes in various fields in historical data and generate evolution rules for metadata evolution; wherein the optimization goal of the evolution rules is to maximize the prediction accuracy and adaptability of the adaptive metadata model; The differences between the data source and the adaptive metadata model will be monitored in real time. When deviations in the adaptive metadata model are detected, automatic feedback will be provided and the adaptive metadata model will be adjusted. The adaptive metadata model is used to monitor the quality of the data source in real time, and warning rules are designed according to the quality of the data source to generate a warning rule set; the warning rules include: When the data quality of a field exceeds the preset tolerance range, an alarm is automatically triggered. When the warning rule is generated, the field weights are assigned different weight values ​​according to the importance of the field through the business weights in the metadata model. At the same time, according to the quality assessment results of each field, the most serious data problems will be monitored first. If the missing value of a field exceeds 30% and it is a core business field, an alarm will be triggered immediately. When a data quality problem triggers an early warning, the data will be repaired through the feedback mechanism and a repair feedback data set will be output; the specific repair method is determined by the type of field and the degree of abnormality of the data: If a field is missing, its historical data trend will be checked first, and the missing data will be filled based on the trend or average value; If the value of a field is outside the predetermined reasonable range, an attempt will be made to correct it; For time series data, time windows and historical data will be used for prediction to ensure data consistency; Design a machine learning-based repair model for data quality issues based on the repair feedback dataset, select appropriate repair methods for different field types and specific data repair methods, and generate the final repaired dataset; The machine learning-based repair model designed for data quality issues includes: If a field is missing, a multi-level adaptive repair mechanism is designed to fill it in, including: If the data of a certain field is missing, an appropriate regression model or neural network will be selected for prediction and filling based on the correlation between fields and the distribution of data; If outliers occur, a variational autoencoder model based on deep learning is introduced to repair the anomaly according to the deep structure of the data; For time series data, time series modeling is used to correct the time field to ensure the consistency of the data in time series.

2. According to the metadata-driven data management method of claim 1, it is characterized in that: The multiple access methods include API, database connection, and file import; the type of the data source is one of a relational database, a non-relational database, a file system, and a data lake.

3. The metadata-driven data management method according to claim 1 is characterized in that: The preliminary metadata model includes structural information of the data source, field names, field types, and field relationships.

4. The metadata-driven data management method according to claim 1 is characterized in that: During the evolution of the preliminary metadata model, the change pattern of each field is captured in real time, and the metadata model is dynamically adjusted using the following formula: ; in, Represents the evolved metadata model, recording changes in fields and their relationships; is the initial metadata model; Indicates a field at a point in time Changes in is the change weight coefficient, which indicates the influence of changes in different fields; n is the total number of fields.

5. The metadata-driven data management method according to claim 4 is characterized in that: The evolution rule set of the metadata evolution includes rules for adjusting and optimizing the metadata model, which is expressed as ,in, It is Evolution rules that describe how to handle changes in field types or adjustments to relationships; Indicates the time when the rule is generated. The evolutionary rules will be updated in time according to the actual data changes; When the adaptive metadata model is detected to have a deviation, automatic feedback is performed and the adaptive metadata model is adjusted, which is expressed as: Assume that in the feedback mechanism, based on changes in business requirements or data quality issues, the final adjustment of the metadata model is as follows: ; in, It is the final adjusted metadata model; It is a metadata model obtained through adaptive evolution; It is an adjustment item based on real-time data feedback.

6. The metadata-driven data management method according to claim 1 is characterized in that: In the process of using the adaptive metadata model to monitor the quality of the data source in real time, an adaptive quality assessment algorithm is introduced to dynamically adjust the quality assessment standard according to different data sources and business requirements, which is expressed as: ; in, Indicates the quality of the data source, including quality information of each field, including indicators such as completeness, inconsistency, and accuracy. The format is ,in Representative Quality assessment results of fields; is an evaluation function that combines field data , metadata model and quality standards To calculate the quality of each field; It is an adaptive metadata model; is the quality standard for each field.

7. The metadata-driven data management method according to claim 6 is characterized in that: The generation process of the warning rule is as follows: ; in, It is Warning rules; It is the warning generation function, combined with field data , Field Weight and threshold To generate warning rules; Is Field The importance weight of is the quality threshold of the field; The data repair through the feedback mechanism is an intelligent repair based on historical trends and machine learning, which is expressed as: ; in, is the repair feedback dataset; It is the data after quality assessment; is the increment of repair data.

8. The metadata-driven data management method according to claim 7 is characterized in that: Further repair is performed on the repair feedback dataset, which is expressed as: ; in, It is the final repaired dataset, which contains all the repaired data; It is a repair feedback dataset, which contains the results of quality assessment and preliminary repair; It indicates the repair increment after optimization, that is, the part that is further repaired.

9. The metadata-driven data management method according to claim 1 is characterized in that: Introducing a regularization term on the adaptive weighted regression model , to prevent overfitting, expressed as: ; in, is a regularized loss function used to constrain model parameters; is the regularization coefficient, which controls the strength of regularization; The first The weight of the parameters; Combined with regularization term , the goal of the adaptive weighted regression model is to minimize the sum of the prediction error and the regularization loss, the loss function It is expressed as: ; in, is the missing value of the prediction; is the actual missing value; The loss function of the variational autoencoder model is expressed as: ; in, is the loss function of the variational autoencoder; is the distribution of latent variables generated by the encoder; is the distribution of sample data generated by the decoder; is the KL divergence, which measures the difference between the encoder distribution and the prior distribution.

10. The metadata-driven data management system is characterized by: The system comprises: A data source access unit is used to connect to various heterogeneous data sources through various access methods, obtain raw data, and perform preliminary analysis on the raw data using a preset preliminary metadata model; The meta-model preliminary construction unit is used to automatically adjust and optimize the metadata model according to the historical flow pattern, change rules and business needs of the original data to obtain an adaptive metadata model after adaptive evolution; A metadata evolution unit, used to use a machine learning method to identify the law of changes in various fields in historical data and generate evolution rules for metadata evolution; wherein the optimization goal of the evolution rules is to maximize the prediction accuracy and adaptability of the adaptive metadata model; The meta-model adjustment unit is used to monitor the difference between the data source and the adaptive metadata model in real time, and when a deviation of the adaptive metadata model is detected, automatically provide feedback and adjust the adaptive metadata model; The metadata warning unit is used to use the adaptive metadata model to monitor the quality of the data source in real time, and to design warning rules according to the quality of the data source to generate a warning rule set; the warning rules include: When the data quality of a field exceeds the preset tolerance range, an alarm is automatically triggered. When the warning rule is generated, the field weights are assigned different weight values ​​according to the importance of the field through the business weights in the metadata model. At the same time, according to the quality assessment results of each field, the most serious data problems will be monitored first. If the missing value of a field exceeds 30% and it is a core business field, an alarm will be triggered immediately. The metadata repair unit is used to repair the data through the feedback mechanism and output the repair feedback data set when the data quality problem triggers an early warning. The specific repair method is determined according to the type of field and the degree of abnormality of the data: If a field is missing, its historical data trend will be checked first, and the missing data will be filled based on the trend or average value; If the value of a field is outside the predetermined reasonable range, an attempt will be made to correct it; For time series data, time windows and historical data will be used for prediction to ensure data consistency; The metadata repair unit is used to design a machine learning-based repair model for data quality issues based on the repair feedback data set, select appropriate repair methods for different field types and specific data repair methods, and generate the final repaired data set; The machine learning-based repair model designed for data quality issues includes: If a field is missing, a multi-level adaptive repair mechanism is designed to fill it in, including: If the data of a certain field is missing, an appropriate regression model or neural network will be selected for prediction and filling based on the correlation between fields and the distribution of data; If outliers occur, a variational autoencoder model based on deep learning is introduced to repair the anomaly according to the deep structure of the data; The data correction unit is used to correct the time field of time series data using time series modeling to ensure the consistency of data in time series.

Citation Information

Patent Citations

  • Method for improving data quality of heterogeneous system

    CN117312290A

  • Intelligent data access and integration system based on dynamic expansion architecture

    CN119311754A

  • Data Visibility and Quality Management Platform

    US20240281419A1

  • Adaptive outlier detection and correction

    US20250005001A1

Cited By

  • Data source fault processing system based on AI server

    CN120315968A

  • Recall activation method for failure clues

    CN120578921A

  • A method for recalling and reactivating expired clues

    CN120578921B