Data quality monitoring method based on big data analysis

By dividing data into different life cycle stages and adopting dynamic scheduling strategies, the full life cycle coverage problem of data quality monitoring in the big data environment is solved, and the comprehensive and dynamic nature of data quality monitoring is achieved, and the adaptability and efficiency of the system are improved.

CN120011174APending Publication Date: 2025-05-16ZHEJIANG NINGYIN CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510089874.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology is difficult to dynamically monitor and manage the entire life cycle of data in a big data environment, resulting in difficult to effectively solve data quality problems.

Method used

By dividing data into different life cycle stages, using data life cycle state and scheduling strategy, combining weighted similarity calculation and custom optimization functions, the comprehensive and dynamic nature of data quality monitoring is achieved.

Benefits of technology

It improves the comprehensiveness and accuracy of data quality monitoring, improves the adaptability and efficiency of the system, and ensures the continuous optimization and security of data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011174A_ABST
    Figure CN120011174A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data quality monitoring based on big data analysis, and discloses a data quality monitoring method based on big data analysis. According to the scheme, the data is divided into life cycle stages such as acquisition, cleaning, storage, use, analysis, archiving and destruction, so that the problem of full-process coverage of data quality monitoring is solved, and the monitoring comprehensiveness is improved. By formulating the scheduling strategy for the data item, the priority and resource allocation are optimized, the efficiency is improved, and the delay is reduced. A dynamic scheduling and log mechanism enhances system adaptability and automatically adjusts data states to cope with changes. And the knowledge graph is adopted to construct the data asset association graph, so that the monitoring pertinence and effect are improved. Automatic scheduling and state adjustment reduce personal errors and ensure timely repair of data quality problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data quality monitoring based on big data analysis, and specifically to a data quality monitoring method based on big data analysis. Background Art

[0002] With the rapid development of big data technology, data has become an indispensable asset in corporate decision-making and operations. Especially in the consumer finance industry, companies rely on a large amount of data to guide business decisions, risk management, customer analysis, etc. However, with the rapid increase in data volume, how to ensure data quality has become a huge challenge. The quality of data directly affects the accuracy of business analysis results and the effectiveness of decision-making. Therefore, data quality monitoring and governance have become the most critical part of data management.

[0003] Many current data quality monitoring methods focus more on the quality assessment of a single data item or tabular data, and fail to fully consider the changes and evolution of data in its life cycle. For example, in different stages such as data collection, cleaning, storage, and use, the quality standards and quality issues of data may be very different. Existing technologies usually lack dynamic monitoring and management of the entire data life cycle, and it is difficult to deal with data quality issues in a big data environment. Many existing technical methods are based on static rules, which means that they cannot dynamically adjust and optimize the data quality monitoring process. For example, rule-based data quality monitoring methods cannot handle the quality issues of real-time data streams, and the setting of rules usually relies on manual experience and lacks the ability to adjust automatically. Traditional data quality assessments mostly focus on single dimensions such as accuracy, completeness, and timeliness, while ignoring the correlation between data and multi-dimensional quality issues. For example, in the financial industry, the interdependence between data items is very complex, and only evaluating from a single quality perspective cannot fully reflect the quality status of the data. In addition, existing quality assessment methods often do not take into account the quality changes of different data items at different times and in different scenarios.

[0004] This solution can better address the limitations of traditional methods in a big data environment by combining data lifecycle management, multi-dimensional data quality assessment, and automated scheduling mechanisms. Specifically, this solution defines data lifecycle states and scheduling strategies to make data quality monitoring more comprehensive and dynamic; introduces weighted similarity calculations and custom optimization functions to improve the accuracy and flexibility of quality assessment; and introduces encryption algorithms to ensure data security, further enhancing the practicality and security of the system. Summary of the invention

[0005] The present invention provides a data quality monitoring method based on big data analysis, which promotes solving the problems mentioned in the above background technology.

[0006] The present invention provides the following technical solution: a data quality monitoring method based on big data analysis, comprising:

[0007] Divide data into different life cycle stages;

[0008] The different life cycle stages include collection, cleaning, storage, use, analysis, archiving and destruction;

[0009] The data life cycle state is recorded as L i (t), where i∈{1,2,...,n} represents different life cycle stages of a data item;

[0010] Among them, L i (t) represents the life cycle state of data item i at time t;

[0011] The priority of obtaining data, denoted as P i ;

[0012] Develop a data scheduling strategy for each data item based on the data priority;

[0013] The specific data scheduling strategy is:

[0014] S i (t) = αP i +βC i (t);

[0015] Among them, S i (t) represents the scheduling status of data item i at time t; C i (t) is the computing resource requirement of data item i at time t; α and β are weight coefficients;

[0016] Record data status changes through logs and automatically adjust data status according to dynamic scheduling formulas;

[0017] The dynamic scheduling formula is as follows:

[0018] S i (t+1)=S i (t)+ΔS i ;

[0019] Among them, ΔS i Indicates the amount of change in the state of data item i;

[0020] Knowledge graph technology is used to model metadata, build a relationship map between data assets, and monitor data flow and data quality through dynamic mapping.

[0021] Optionally, the C i (t), including:

[0022] Let the i-th data item be D i ;

[0023] Get data item D i The storage space required at time t is denoted as Storage(D i );

[0024] Get data item D i The network bandwidth required at time t is denoted as DataTraffic (D i );

[0025] Then C i (t) = δ1·Storage(D i )+δ2·DataTraffic(D i );

[0026] Among them, δ1 and δ2 are coefficients for adjusting the weights of storage space and network bandwidth.

[0027] Optionally, the metadata is modeled using knowledge graph technology to construct a relationship graph between data assets, and the data flow and data quality are monitored through dynamic mapping, specifically including:

[0028] Metadata about the collected data items, including the data’s unique identifier, source, type, and collection time;

[0029] The metadata is recorded as: M i =(ID i ,type i ,source i ,timestamp i );

[0030] Among them, M i is the metadata of data item i; ID i is the unique identifier of data item i; type i is the type of data item i; source i The source of data item i; timestamp i is the collection time of data item i;

[0031] Real-time access to data and metadata;

[0032] By acquiring data and metadata in real time, the mapping relationship between data items is dynamically updated;

[0033] Set a mapping matrix R i,j , specifically:

[0034] R i,j=γ1·Similarity(M i ,M j )+γ2·Correlation(M i ,M j );

[0035] Among them, R i,j is the mapping relationship between data items i and j; Similarity(M i ,M j ) is used to calculate the similarity between data item i and data item j in the metadata space; Correlation(M i ,M j ) is used to measure the correlation between data item i and data item j; γ1 and γ2 are adjustment coefficients, which respectively control the weights of similarity and correlation in the mapping relationship;

[0036] Build a knowledge graph based on the mapping relationship between data items;

[0037] The icon structure is represented as: G = (V, E);

[0038] Among them, V is a data node, representing each data item; E is an edge between data, representing the relationship between data items;

[0039] Optimize the knowledge graph through the graph neural network model: E'=E+ΔE;

[0040] Among them, ΔE represents the incremental correction of the relationship edge;

[0041] Establish intelligent data modeling based on adaptive algorithms according to data lifecycle and metadata mapping.

[0042] Optionally, the Similarity(M i ,M j ), including:

[0043] Get data item M i The kth feature of k (M i );

[0044] Get data item M j The kth feature of k (M j );

[0045] Get the total number of features in the metadata, denoted as m;

[0046] The similarity calculation is as follows:

[0047]

[0048] Optionally, the Correlation(M i ,M j ), including:

[0049] Get data item M i The mean of the kth feature is denoted by

[0050] Get data item M j The mean of the kth feature is denoted by

[0051] The correlation is measured by calculating the degree of association between data items, specifically:

[0052]

[0053] Optionally, establishing intelligent data modeling based on an adaptive algorithm according to the data life cycle and metadata mapping specifically includes:

[0054] Let the i-th data item be D i ;

[0055] Get data item D i The kth feature of k (D i );

[0056] Set the data modeling function to f(D i );

[0057] The modeling function is as follows:

[0058] Among them, f(D i ) is the data item D i Modeling function of ω k is the weight coefficient of the kth feature;

[0059] Get data item D i The kth true eigenvalue of k (D i );

[0060] Calculate data item D i The error:

[0061]

[0062] Set the optimization function to f opt (D i ), the optimization function is as follows:

[0063]

[0064] Among them, err(D i ) is the data item D i The error of opt (D i ) is the optimized objective function;

[0065] The optimization function is iteratively optimized by the gradient descent method, and the model parameters are updated as follows:

[0066]

[0067] Among them, θ k is the parameter of the kth iteration; η is the learning rate; To optimize the gradient of the function with respect to the parameters;

[0068] Automatically correct data quality through intelligent data quality management and predictive strategies.

[0069] Optionally, automatically correcting data quality through intelligent data quality management and prediction strategies specifically includes:

[0070] The data item D at time r i Recorded as

[0071] Get data items The effectiveness of

[0072] Get data items The consistency of

[0073] Get data items The completeness of

[0074] Set the quality measure function to qualitymeasure(D i ), specifically:

[0075]

[0076] Set the prediction data quality function to Q pred (t), specifically:

[0077]

[0078] Among them, ρ i For data item D i The weight of the data reflects its impact on the overall data quality; is the data item D at time t i Quality measures;

[0079] Set validity, consistency and completeness thresholds respectively;

[0080] Based on the prediction quality function Q pred (t), when the data quality is lower than any threshold, the automatic correction mechanism is activated, specifically:

[0081]

[0082] Where, ΔD i Represents data item D i The quality correction amount;

[0083] Protect data through data security policies.

[0084] Optionally, protecting data through a data security policy specifically includes:

[0085] Get the public key, denoted as K pub ;

[0086] Encrypting data items is as follows: i )=E(D i ,K pub );

[0087] Among them, E is the encryption operation; D i is the data item to be encrypted;

[0088] Store the encrypted data hash value into the blockchain;

[0089] The data storage formula is: Store(H(D i ))=B(H(D i ),T);

[0090] Among them, H(D i ) is data D i The hash value of; B is the blockchain storage operation; T is the timestamp;

[0091] Set up an audit interface to record each data change;

[0092] The traceback formula is:

[0093] Among them, T trace H is the result of change traceability; i A hash value for each data change.

[0094] The present invention has the following beneficial effects:

[0095] 1. By dividing the data into different life cycle stages, the problem of full process coverage of data quality monitoring is solved and the comprehensiveness of monitoring is improved. By dividing the data into different life cycle stages such as collection, cleaning, storage, use, analysis, archiving and destruction, this method can cover the entire life cycle of the data, thereby realizing comprehensive monitoring of data quality. At each stage of the data, the quality requirements of the data may be different. Phased management can more accurately monitor and repair data quality problems, avoid the overall data quality degradation caused by neglect of a certain stage, and thus improve the accuracy of decision-making. By formulating a data scheduling strategy for each data item, the flexibility and priority issues of data scheduling are solved, and the efficiency of the system is improved. By formulating a data scheduling strategy based on the priority of the data, this solution can flexibly adjust the priority of data processing according to the urgency of the data item and the computing resource requirements, ensuring that resources are used most effectively. This can reduce data processing delays caused by improper resource allocation, thereby improving the efficiency and response speed of data quality monitoring. By dynamically scheduling formulas and logging data state changes, the dynamic changes in data traffic and processing requirements are solved, and the adaptability of the system is enhanced. Through the dynamic scheduling formula and log recording mechanism, this solution can dynamically calculate the amount of change in data status according to the data flow and processing requirements monitored by the system, and automatically adjust the data status. This dynamic adjustment mechanism can optimize the data scheduling strategy according to the actual situation, improve the adaptability and flexibility of the system, and effectively deal with the real-time and complexity problems in the big data environment. By using the knowledge graph technology to build the association relationship map between data assets, the visualization and correlation problems of data asset management are solved, and the monitoring effect is improved. In data quality monitoring, the association relationship map between data assets is built through the knowledge graph, which can make the source of data quality problems clearer, and then monitor and repair them more targeted. Through dynamic mapping, this solution can reflect the quality problems in the data flow process in real time, help managers better understand the factors affecting data quality and the changing trend of data quality status, and ultimately improve the monitoring and governance effect of data quality. Through the dynamic scheduling formula and the mechanism of automatically adjusting the data status in this solution, data quality problems can be automatically dealt with, manual operations can be reduced, and monitoring efficiency and accuracy can be improved. The automated data scheduling and status adjustment mechanism can respond immediately when quality problems occur, ensuring that data quality problems are repaired in a timely manner and reducing errors and delays in manual operations.

[0096] 2. By obtaining the storage space and network bandwidth required by the data item at time t, the problem of opaque resource demand is solved and the accuracy of resource allocation is improved. By obtaining the storage space and network bandwidth required by the data item at time t in this solution, the specific resource requirements of each data item at different life cycle stages can be accurately understood, avoiding ineffective consumption of resources. In addition, the dynamic monitoring and calculation of storage space and network bandwidth can provide the system with a clearer resource usage situation, provide a basis for further optimizing resource scheduling and management, and thus improve the overall system efficiency. By defining the storage space and network bandwidth weight adjustment coefficients, the flexibility and priority problems in resource scheduling are solved, and the intelligent level of resource allocation is improved. By introducing the weight coefficients of storage space and network bandwidth, the system can flexibly adjust the allocation of resources according to the importance and priority of data items. For example, when some data items need to be processed urgently, the system can adjust the weight coefficients to give priority to allocating more storage space and bandwidth resources to ensure timely processing of data items. This intelligent scheduling method not only improves the response speed of the system, but also avoids the neglect of important data items when resources are limited, thereby effectively improving the overall efficiency of the system and the accuracy of data quality monitoring. By combining the calculation of storage space and network bandwidth, the resource bottleneck problem in the big data system is solved, and the data processing bottleneck caused by insufficient resources is avoided. In the big data environment, how to reasonably allocate and schedule resources in massive data processing is a common challenge. Without accurate prediction of storage space and bandwidth requirements, the system may encounter resource bottlenecks in the processing of certain data items, which will affect the efficiency of the overall system and the real-time performance of data quality monitoring. By calculating the storage space and network bandwidth required for the data items, this solution can predict the resource requirements of each data item before data processing, avoiding the sudden shortage of resources during the processing process. At the same time, by dynamically adjusting resource allocation, it can ensure that data items can be processed smoothly, avoid delays caused by resource bottlenecks, and further improve data processing efficiency. By accurately controlling storage space and network bandwidth, the reliability and stability of the data quality monitoring system are improved.

[0097] 3. By collecting and updating metadata in real time, the problem of untimely tracking of data sources and data status is solved, and the accuracy of data flow monitoring is improved. By collecting metadata of data items in real time, including the unique identifier, source, type and collection time of the data items, a comprehensive metadata description is provided for each data item. This method can track the source, type and life cycle stage of data in real time to ensure full monitoring of the data flow process. Real-time acquisition of data and metadata, and dynamic update of the mapping relationship between data items, not only makes data monitoring more accurate, but also can timely discover and correct data quality problems, and improve the transparency and reliability of data management. Through the dynamic mapping relationship matrix, the problem of unclear association between data items is solved, and the correlation analysis between data is enhanced. By setting the mapping matrix, the mapping relationship between data items is defined, and the relationship between data items is measured using similarity and correlation indicators. The mapping matrix can not only clearly present the mutual connection between data items, but also dynamically optimize and update the mapping relationship according to the adjustment coefficient of similarity and correlation. Through this dynamic adjustment, the system can more accurately establish the association between data items, ensuring more efficient and intelligent data flow and quality monitoring. By constructing a knowledge graph, the problem of lack of global perspective in traditional data relationship analysis is solved, and the relevance and structured degree between data assets are improved. By constructing a knowledge graph, this solution can intuitively represent the relationship between data items. The knowledge graph provides a global perspective between data assets in the form of nodes and edges, which can not only reveal the direct relationship between data items, but also explore potential relevance. In this way, managers can clearly see the position and interdependence of data items in the overall system, which is convenient for data quality monitoring and management decisions, and improves the visualization and management efficiency of data assets. By optimizing the knowledge graph through the graph neural network, the problems of unstable graph structure and insufficient data relevance in large-scale data management are solved, and the accuracy and efficiency of graph analysis are improved. By introducing the graph neural network model, this solution can incrementally correct the relationship edges in the graph, further improving the accuracy and stability of the graph. The graph neural network can continuously optimize the relationship between nodes and edges, so that the knowledge graph can still maintain high analysis efficiency and low error rate when facing a large amount of data. This dynamic optimization of the graph can provide more accurate relationship analysis in data flow, data processing and data quality monitoring, and improve the intelligence level of the data monitoring system. Through intelligent data modeling based on adaptive algorithms, the static and inefficient problems of the data modeling process are solved, and the flexibility and accuracy of modeling and monitoring are improved. By introducing intelligent data modeling based on adaptive algorithms, this solution enables the data modeling process to be dynamically adjusted according to changes in the data life cycle and metadata mapping. The adaptive algorithm can automatically adjust the modeling parameters and strategies according to the real-time situation of data flow, so as to more efficiently adapt to the changes in the needs of different data items at different stages.This method not only improves the efficiency of data modeling, but also ensures the flexibility of data quality monitoring and analysis, and can respond to changing business needs and data environments in real time. Through dynamic mapping and optimization technology, the delay and inaccuracy problems in data flow monitoring are solved, and the real-time and accuracy of data quality monitoring are improved. Through dynamic mapping relationships and graph neural network optimization technology, changes in data flow can be reflected in real time, and the relationship between data items can be optimized and adjusted through the graph model. This dynamic mapping can not only respond quickly to changes in data flow, but also effectively improve the real-time and accuracy of data quality monitoring. By continuously optimizing and adjusting the mapping matrix and graph, the system can identify anomalies or quality problems in the data flow more quickly, so as to provide early warnings and adjustments before problems occur, greatly improving the response speed and accuracy of data quality monitoring.

[0098] 4. By obtaining the kth feature of the data item and performing similarity calculation, the problem of inaccurate similarity measurement between data is solved and the accuracy of data matching is improved. By obtaining the kth feature of each data item and performing feature-based similarity calculation, the similarity calculation can be more refined and accurate. In this way, the similarity between data items can be measured from multiple dimensions, the accuracy of data matching is improved, and the deviation caused by a single feature is avoided. By obtaining the total number of features in the metadata and optimizing the similarity calculation, the problem of high computational complexity in high-dimensional data is solved and the processing efficiency is improved. By obtaining the total number of features in the metadata and combining feature selection technology to optimize the similarity calculation process, unnecessary calculations are reduced. In this way, the calculation focus can be focused on features with high influence, thereby avoiding the problem of high computational complexity caused by high-dimensional data. At the same time, the use of feature selection technology effectively reduces the amount of calculation, improves the efficiency of similarity calculation, and ensures the data processing speed and responsiveness in a big data environment. By optimizing data matching and correlation analysis through similarity calculation, the problem of insufficient understanding of multidimensional data relationships in data management is solved and the multi-dimensional accuracy of data quality monitoring is improved. By calculating the similarity of multiple features of each data item, the relationship between data items can be comprehensively evaluated from multiple dimensions. This method not only improves the accuracy of data quality monitoring, but also makes data correlation analysis more in-depth, and can discover potential data patterns or anomalies, thereby improving the comprehensiveness and accuracy of data quality management. By calculating the similarity at the feature level, the problem that traditional methods cannot flexibly respond to changes in different data feature structures is solved, and the adaptability and flexibility of the data monitoring system are enhanced. By calculating the similarity of the kth feature of each data item and combining it with the dynamic adjustment of the total number of features, the system can flexibly adapt to changes in the data feature structure. Whether it is a new data feature or a change in the feature, the system can adaptively adjust in the similarity calculation, so that the data monitoring system can continuously adapt to new data features and business needs. Through this flexible and adaptive mechanism, the real-time and accuracy of data monitoring and quality analysis are improved.

[0099] 5. By obtaining the mean of the kth feature of the data item and calculating the correlation, the problem that the traditional correlation measurement method cannot accurately capture the subtle correlation between data is solved, and the depth of data analysis is improved. By obtaining the mean of the kth feature of the data item, the similarity and correlation between data items can be measured from a refined feature perspective. This method can not only capture the global correlation between data items, but also go deep into the local features of each data item, thereby improving the accuracy of correlation calculation. In this way, potential subtle data relationships can be discovered, effectively improving the depth and accuracy of data quality monitoring. By calculating the correlation between data items to measure the correlation, the problem of difficulty in calculating the correlation relationship between multiple features in high-dimensional data is solved, and the calculation efficiency is improved. By calculating the mean of the kth feature of the data item and further calculating the correlation between the data items to measure the correlation between the data, the calculation complexity can be effectively reduced. By focusing on the mean of each feature and calculating the correlation at the feature level, the calculation burden when processing high-dimensional data is avoided, and the efficiency and scalability of the system are improved. By introducing the measurement method of correlation, the problem of difficult feature selection in multi-dimensional data quality analysis is solved, and the flexibility and adaptability of the model are improved. By introducing the correlation between data items, the importance of different features can be measured. Through correlation measurement, the system can dynamically adjust the contribution of each feature in data quality monitoring, thereby improving the flexibility and adaptability of the system. Regardless of how the number and dimensions of data features change, the system can quickly adjust and optimize the analysis strategy to provide more accurate monitoring results. By accurately calculating the correlation between data items, the problem of misjudgment in data monitoring is solved, and the accuracy and reliability of data quality detection are improved. By calculating the mean of the kth feature of the data item and measuring the correlation between the data items, the true relationship between the data items can be accurately reflected. This refined correlation calculation method effectively reduces the probability of misjudgment and ensures the high accuracy and reliability of data quality monitoring results. By accurately modeling the feature association between data, the system can more keenly discover data quality problems, take timely countermeasures, and improve the efficiency and effectiveness of data management.

[0100] 6. By setting the data modeling function and optimizing the error, the problem of the inability to accurately capture data features and optimization in traditional data modeling is solved, and the accuracy and prediction ability of the model are improved. By setting the data modeling function and dynamically optimizing the weight coefficient of each feature, the characteristics of the data item can be more accurately described, thereby improving the accuracy of modeling. By calculating and optimizing the model error, it is ensured that the modeling function can adapt to the changes of different data features. The optimization goal is to minimize the error, thereby improving the accuracy of the prediction results and solving the problem that the traditional modeling method cannot dynamically adapt to data changes. Iterative optimization through optimization function and gradient descent method solves the slow convergence speed and local optimal problems in data modeling, and improves the model training efficiency and global optimality. By iteratively optimizing the optimization function, the model parameters are gradually updated. This method can not only improve the training efficiency, but also avoid the model from falling into the local optimal solution, ensuring that the global optimal model parameters are finally obtained. Through this optimization method, the efficiency and effect of the data modeling process can be greatly improved. Through intelligent data quality management and prediction strategy, data quality is automatically corrected, which solves the dilemma of insufficient manual intervention and difficulty in solving data quality problems in real time, and improves the automation and real-time nature of data quality management. Through the intelligent data quality management and prediction strategy in this solution, the system can automatically identify and correct data quality problems, thereby improving the automation of data quality management. The system dynamically adjusts and corrects data quality according to the mapping relationship between data life cycle and metadata, making the data quality management process more efficient and accurate, and avoiding the lag and limitations of manual intervention. Through the application of adaptive algorithms, the problem of inconsistent quality assessment of data items in different life cycle stages is solved, and the flexibility and accuracy of data quality monitoring are improved. Through intelligent data modeling based on adaptive algorithms, the modeling function and optimization strategy can be dynamically adjusted according to the life cycle stage of the data and different metadata characteristics, thereby improving the flexibility and accuracy of data quality assessment. Data at each stage has its own specific quality requirements. By automatically optimizing and correcting data quality, it can ensure that data at different stages can be efficiently and accurately monitored. Through the weight adjustment of data item features and error minimization optimization, the problem of feature importance evaluation in multi-feature modeling is solved, and the scientific nature of feature selection is improved. Through the weight adjustment of data item features, this solution can dynamically evaluate the contribution of each feature and automatically adjust the feature weight during the modeling process. By optimizing the model error, the influence of different features can be accurately reflected, which improves the scientificity and accuracy of feature selection and further optimizes the data modeling process.

[0101] 7. By setting the quality measurement function and the predicted data quality function, the problem of data quality assessment and prediction is solved, and the accuracy and predictability of data quality management are improved. By setting the quality measurement function and the predicted data quality function, not only can the quality of data items be evaluated in real time, but also the quality of data at future moments can be predicted. This method provides more accurate quality assessment results by dynamically calculating the validity, consistency and integrity of data items. By predicting data quality, potential quality problems can be identified in time, measures can be taken in advance, and risks caused by data quality problems can be avoided. By defining the validity, consistency and integrity thresholds, the dilemma of lack of reasonable threshold judgment in data quality management is solved, and the accuracy of data quality control is improved. By setting the validity, consistency and integrity thresholds, this solution defines specific quality standards for each data item, which can accurately determine whether the data meets the expected quality requirements. This threshold-based judgment mechanism effectively solves the ambiguity problem of data quality assessment and enhances the accuracy and operability of data quality control. By starting the automatic correction mechanism based on the predicted quality function, the problem of insufficient manual intervention and untimely repair of data quality problems is solved, and the correction efficiency and data quality assurance capabilities are improved. Through the quality function based on prediction, the correction mechanism is automatically started when the data quality is found to be lower than the threshold. This mechanism corrects the data through machine learning methods, which can quickly and accurately solve data quality problems, improve the automation level of data quality repair, and ensure that data quality management is more efficient and real-time. Correcting data through machine learning methods solves the limitations of traditional manual correction methods and improves the intelligence and accuracy of correction. Automatic correction of data quality through machine learning methods can be dynamically adjusted according to data characteristics and historical data. Machine learning algorithms can automatically adjust correction strategies according to data changes in big data environments, avoiding the limitations of traditional methods. This intelligent data correction method not only improves the correction efficiency, but also improves the accuracy of data correction, ensuring continuous optimization of data quality. Protecting data through data security policies solves the security risks that may be caused during data quality correction and enhances the security and compliance of data processing. Through data security policies, data security is guaranteed while correcting data to prevent data from being improperly processed or leaked during the correction process. This strategy not only ensures the improvement of data quality, but also guarantees data compliance and privacy protection, improving the security of data governance.

[0102] 8. Protecting data security through encryption operations and public key mechanisms solves the problem of unauthorized access or tampering of data during transmission and storage, ensuring the confidentiality and integrity of data. By obtaining the public key and using encryption operations to encrypt the data items, the data items are encrypted and then stored, ensuring that the data can be effectively protected during transmission and storage. The encryption process prevents the data from being decrypted or tampered with without authorization, thereby ensuring the confidentiality and integrity of the data. In addition, data encryption also enhances data privacy protection and effectively prevents external attackers from maliciously acquiring or tampering with the data content. Storing data hash values ​​through blockchain technology solves the problem of data tampering and loss, and improves the immutability and traceability of data storage. By storing the encrypted data hash value in the blockchain, the hash value of each data item is recorded on the blockchain. Blockchain has the characteristics of immutability and distributed storage, ensuring that data cannot be changed or deleted once written. This measure greatly improves the security and credibility of data storage and prevents data from being tampered with or lost. The data storage formula ensures that data records cannot be tampered with, solves the problem of data loss or tampering that may occur in traditional storage methods, and improves the security and transparency of data. Traditional database storage methods cannot effectively prevent data modification or deletion, and lack records and audits of data changes. This solution uses data storage formulas to ensure that all changes in the data storage process will be recorded as data hash values ​​and stored on the blockchain. This not only enhances the security of data storage, but also improves the transparency of data operations. In the blockchain, each data change will be tracked and recorded to form an audit chain that cannot be tampered with, ensuring the historical authenticity of data storage. The audit interface and change traceability function solve the problem of insufficient transparency of data operations and improve the traceability and auditability of data changes. By setting up an audit interface and a traceability formula, each data change is recorded and the hash value of each change can be traced. The audit interface can not only realize real-time recording of data changes, but also effectively track historical changes, providing a complete audit trail for data operations. This mechanism greatly enhances the transparency and traceability of data management, improves the ability to supervise data changes, and ensures the compliance of data during management and use. By combining blockchain and encryption technology, the dual problems of data security and privacy protection are solved, and the security and compliance of the overall system are improved. By combining blockchain technology and encryption technology, not only the security and privacy of data are guaranteed, but also the compliance of data during storage and processing is ensured. Encryption technology ensures that data is not leaked during transmission and storage, while blockchain technology provides an unalterable data storage solution and ensures that every change in data can be tracked through the audit interface. This combined method effectively solves the problems of data security and privacy protection and provides comprehensive protection for data management. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0104] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0105] Example, see Figure 1 , a data quality monitoring method based on big data analysis, comprising:

[0106] Divide data into different life cycle stages;

[0107] The different life cycle stages include collection, cleaning, storage, use, analysis, archiving and destruction;

[0108] The data life cycle state is recorded as L i (t), where I∈{1,2,...,n} represents different life cycle stages of data items;

[0109] Among them, L i (t) represents the life cycle state of data item i at time t;

[0110] The priority of obtaining data, denoted as P i ;

[0111] Develop a data scheduling strategy for each data item based on the data priority;

[0112] The specific data scheduling strategy is:

[0113] S i (t) = αP i +βC i (t);

[0114] Among them, S i (t) represents the scheduling status of data item i at time t; C i (t) is the computing resource requirement of data item i at time t; α and β are weight coefficients that adjust the impact of data priority and resource requirements on the scheduling state;

[0115] Record data status changes through logs and automatically adjust data status according to dynamic scheduling formulas;

[0116] The dynamic scheduling formula is as follows:

[0117] S i (t+1)=S i (t)+ΔS i ;

[0118] Among them, ΔS i It represents the change in the state of data item i, which is dynamically calculated based on the data flow and processing requirements monitored by the system;

[0119] Knowledge graph technology is used to model metadata, build a relationship map between data assets, and monitor data flow and data quality through dynamic mapping.

[0120] By dividing data into different life cycle stages, the problem of full process coverage of data quality monitoring is solved and the comprehensiveness of monitoring is improved. By dividing data into different life cycle stages such as collection, cleaning, storage, use, analysis, archiving and destruction, this method can cover the entire life cycle of data, thereby realizing comprehensive monitoring of data quality. At each stage of data, the quality requirements of data may be different. Phased management can more accurately monitor and repair data quality problems, avoid the overall data quality degradation caused by neglect of a certain stage, and thus improve the accuracy of decision-making. By formulating a data scheduling strategy for each data item, the flexibility and priority issues of data scheduling are solved, and the efficiency of the system is improved. By formulating a data scheduling strategy based on the priority of the data, this solution can flexibly adjust the priority of data processing according to the urgency of the data item and the computing resource requirements, ensuring that resources are used most effectively. This can reduce data processing delays caused by improper resource allocation, thereby improving the efficiency and response speed of data quality monitoring. By dynamically scheduling formulas and logging data state changes, the dynamic changes in data traffic and processing requirements are solved, and the adaptability of the system is enhanced. Through the dynamic scheduling formula and log recording mechanism, this solution can dynamically calculate the amount of change in data status according to the data flow and processing requirements monitored by the system, and automatically adjust the data status. This dynamic adjustment mechanism can optimize the data scheduling strategy according to the actual situation, improve the adaptability and flexibility of the system, and effectively deal with the real-time and complexity problems in the big data environment. By using the knowledge graph technology to build the association relationship map between data assets, the visualization and correlation problems of data asset management are solved, and the monitoring effect is improved. In data quality monitoring, the association relationship map between data assets is built through the knowledge graph, which can make the source of data quality problems clearer, and then monitor and repair them more targeted. Through dynamic mapping, this solution can reflect the quality problems in the data flow process in real time, help managers better understand the factors affecting data quality and the changing trend of data quality status, and ultimately improve the monitoring and governance effect of data quality. Through the dynamic scheduling formula and the mechanism of automatically adjusting the data status in this solution, data quality problems can be automatically dealt with, manual operations can be reduced, and monitoring efficiency and accuracy can be improved. The automated data scheduling and status adjustment mechanism can respond immediately when quality problems occur, ensuring that data quality problems are repaired in a timely manner and reducing errors and delays in manual operations.

[0121] The C i (t), including:

[0122] Let the i-th data item be D i ;

[0123] Get data item D i The storage space required at time t is denoted as Storage(Di );

[0124] Get the network bandwidth required for data item Di at time t, denoted as DataTraffic(D i );

[0125] Then C i (t) = δ1·Storage(D i )+δ2·DataTraffic(D i );

[0126] Among them, δ1 and δ2 are coefficients for adjusting the weights of storage space and network bandwidth.

[0127] By obtaining the storage space and network bandwidth required by the data item at time t, the problem of opaque resource demand is solved and the accuracy of resource allocation is improved. By obtaining the storage space and network bandwidth required by the data item at time t in this solution, the specific resource requirements of each data item at different life cycle stages can be accurately understood, avoiding ineffective consumption of resources. In addition, the dynamic monitoring and calculation of storage space and network bandwidth can provide the system with a clearer resource usage, provide a basis for further optimizing resource scheduling and management, and thus improve the overall system efficiency. By defining the storage space and network bandwidth weight adjustment coefficients, the flexibility and priority problems in resource scheduling are solved, and the intelligent level of resource allocation is improved. By introducing the weight coefficients of storage space and network bandwidth, the system can flexibly adjust the allocation of resources according to the importance and priority of the data items. For example, when some data items need to be processed urgently, the system can adjust the weight coefficients to give priority to allocating more storage space and bandwidth resources to ensure timely processing of the data items. This intelligent scheduling method not only improves the response speed of the system, but also avoids the neglect of important data items when resources are limited, thereby effectively improving the overall efficiency of the system and the accuracy of data quality monitoring. By combining the calculation of storage space and network bandwidth, the resource bottleneck problem in the big data system is solved, and the data processing bottleneck caused by insufficient resources is avoided. In the big data environment, how to reasonably allocate and schedule resources in massive data processing is a common challenge. Without accurate prediction of storage space and bandwidth requirements, the system may encounter resource bottlenecks in the processing of certain data items, which will affect the efficiency of the overall system and the real-time performance of data quality monitoring. By calculating the storage space and network bandwidth required for the data items, this solution can predict the resource requirements of each data item before data processing, avoiding the sudden shortage of resources during the processing process. At the same time, by dynamically adjusting resource allocation, it can ensure that data items can be processed smoothly, avoid delays caused by resource bottlenecks, and further improve data processing efficiency. By accurately controlling storage space and network bandwidth, the reliability and stability of the data quality monitoring system are improved.

[0128] The use of knowledge graph technology to model metadata, construct a graph of associations between data assets, and monitor data flow and data quality through dynamic mapping, specifically includes:

[0129] Metadata about the collected data items, including the data’s unique identifier, source, type, and collection time;

[0130] The metadata is recorded as: M i =(ID i ,type i ,source i ,timestamp i );

[0131] Among them, M i is the metadata of data item i; ID i is the unique identifier of data item i; type i is the type of data item i; source i The source of data item i; timestamp i is the collection time of data item i;

[0132] Real-time access to data and metadata;

[0133] By acquiring data and metadata in real time, the mapping relationship between data items is dynamically updated;

[0134] Set a mapping matrix R i,j , specifically:

[0135] R i,j =γ1·Similarity(M i ,M j )+γ2·Correlation(M i ,M j );

[0136] Among them, R i,j is the mapping relationship between data items i and j; Similarity(M i ,M j ) is used to calculate the similarity between data item i and data item j in the metadata space; Correlation(M i ,M j ) is used to measure the correlation between data item i and data item j; γ1 and γ2 are adjustment coefficients, which respectively control the weights of similarity and correlation in the mapping relationship;

[0137] Build a knowledge graph based on the mapping relationship between data items;

[0138] The icon structure is represented as: G = (V, E);

[0139] Among them, V is a data node, representing each data item; E is an edge between data, representing the relationship between data items;

[0140] Optimize the knowledge graph through the graph neural network model: E'=E+ΔE;

[0141] Among them, ΔE represents the incremental correction of the relationship edge;

[0142] Establish intelligent data modeling based on adaptive algorithms according to data lifecycle and metadata mapping.

[0143] By collecting and updating metadata in real time, the problem of untimely tracking of data sources and data status is solved, and the accuracy of data flow monitoring is improved. By collecting metadata of data items in real time, including the unique identifier, source, type and collection time of the data items, a comprehensive metadata description is provided for each data item. This method can track the source, type and life cycle stage of data in real time to ensure full monitoring of the data flow process. Real-time acquisition of data and metadata and dynamic update of the mapping relationship between data items not only makes data monitoring more accurate, but also can timely discover and correct data quality problems, improving the transparency and reliability of data management. Through the dynamic mapping relationship matrix, the problem of unclear association between data items is solved and the correlation analysis between data is enhanced. By setting the mapping matrix, the mapping relationship between data items is defined, and the relationship between data items is measured using similarity and correlation indicators. The mapping matrix can not only clearly present the mutual connection between data items, but also dynamically optimize and update the mapping relationship according to the adjustment coefficient of similarity and correlation. Through this dynamic adjustment, the system can more accurately establish the association between data items, ensuring more efficient and intelligent data flow and quality monitoring. By constructing a knowledge graph, the problem of lack of global perspective in traditional data relationship analysis is solved, and the relevance and structured degree between data assets are improved. By constructing a knowledge graph, this solution can intuitively represent the relationship between data items. The knowledge graph provides a global perspective between data assets in the form of nodes and edges, which can not only reveal the direct relationship between data items, but also explore potential relevance. In this way, managers can clearly see the position and interdependence of data items in the overall system, which is convenient for data quality monitoring and management decisions, and improves the visualization and management efficiency of data assets. By optimizing the knowledge graph through the graph neural network, the problems of unstable graph structure and insufficient data relevance in large-scale data management are solved, and the accuracy and efficiency of graph analysis are improved. By introducing the graph neural network model, this solution can incrementally correct the relationship edges in the graph, further improving the accuracy and stability of the graph. The graph neural network can continuously optimize the relationship between nodes and edges, so that the knowledge graph can still maintain high analysis efficiency and low error rate when facing a large amount of data. This dynamic optimization of the graph can provide more accurate relationship analysis in data flow, data processing and data quality monitoring, and improve the intelligence level of the data monitoring system. Through intelligent data modeling based on adaptive algorithms, the static and inefficient problems of the data modeling process are solved, and the flexibility and accuracy of modeling and monitoring are improved. By introducing intelligent data modeling based on adaptive algorithms, this solution enables the data modeling process to be dynamically adjusted according to changes in the data life cycle and metadata mapping. The adaptive algorithm can automatically adjust the modeling parameters and strategies according to the real-time situation of data flow, so as to more efficiently adapt to the changes in the needs of different data items at different stages.This method not only improves the efficiency of data modeling, but also ensures the flexibility of data quality monitoring and analysis, and can respond to changing business needs and data environments in real time. Through dynamic mapping and optimization technology, the delay and inaccuracy problems in data flow monitoring are solved, and the real-time and accuracy of data quality monitoring are improved. Through dynamic mapping relationships and graph neural network optimization technology, changes in data flow can be reflected in real time, and the relationship between data items can be optimized and adjusted through the graph model. This dynamic mapping can not only respond quickly to changes in data flow, but also effectively improve the real-time and accuracy of data quality monitoring. By continuously optimizing and adjusting the mapping matrix and graph, the system can identify anomalies or quality problems in the data flow more quickly, so as to provide early warnings and adjustments before problems occur, greatly improving the response speed and accuracy of data quality monitoring.

[0144] Similarity(M i ,M j ), including:

[0145] Get data item M i The kth feature of k (M i );

[0146] Get data item M j The kth feature of k (M j );

[0147] Get the total number of features in the metadata, denoted as m;

[0148] The similarity calculation is as follows:

[0149]

[0150] By obtaining the kth feature of the data item and performing similarity calculation, the problem of inaccurate similarity measurement between data is solved and the accuracy of data matching is improved. By obtaining the kth feature of each data item and performing feature-based similarity calculation, the similarity calculation can be more refined and accurate. In this way, the similarity between data items can be measured from multiple dimensions, the accuracy of data matching is improved, and the deviation caused by a single feature is avoided. By obtaining the total number of features in the metadata and optimizing the similarity calculation, the problem of high computational complexity in high-dimensional data is solved and the processing efficiency is improved. By obtaining the total number of features in the metadata and combining feature selection technology to optimize the similarity calculation process, unnecessary calculations are reduced. In this way, the calculation focus can be focused on features with high influence, thereby avoiding the problem of high computational complexity caused by high-dimensional data. At the same time, the use of feature selection technology effectively reduces the amount of calculation, improves the efficiency of similarity calculation, and ensures the data processing speed and responsiveness in a big data environment. By optimizing data matching and correlation analysis through similarity calculation, the problem of insufficient understanding of multidimensional data relationships in data management is solved and the multi-dimensional accuracy of data quality monitoring is improved. By calculating the similarity of multiple features of each data item, the relationship between data items can be comprehensively evaluated from multiple dimensions. This method not only improves the accuracy of data quality monitoring, but also makes data correlation analysis more in-depth, and can discover potential data patterns or anomalies, thereby improving the comprehensiveness and accuracy of data quality management. By calculating the similarity at the feature level, the problem that traditional methods cannot flexibly respond to changes in different data feature structures is solved, and the adaptability and flexibility of the data monitoring system are enhanced. By calculating the similarity of the kth feature of each data item and combining it with the dynamic adjustment of the total number of features, the system can flexibly adapt to changes in the data feature structure. Whether it is a new data feature or a change in the feature, the system can adaptively adjust in the similarity calculation, so that the data monitoring system can continuously adapt to new data features and business needs. Through this flexible and adaptive mechanism, the real-time and accuracy of data monitoring and quality analysis are improved.

[0151] The Correlation (M i ,M j ), including:

[0152] Get data item M i The mean of the kth feature is denoted by

[0153] Get data item M j The mean of the kth feature is denoted by

[0154] The correlation is measured by calculating the degree of association between data items, specifically:

[0155]

[0156] By obtaining the mean of the kth feature of the data item and calculating the correlation, the problem that the traditional correlation measurement method cannot accurately capture the subtle correlation between data is solved, and the depth of data analysis is improved. By obtaining the mean of the kth feature of the data item, the similarity and correlation between data items can be measured from a refined feature perspective. This method can not only capture the global correlation between data items, but also go deep into the local features of each data item, thereby improving the accuracy of correlation calculation. In this way, potential subtle data relationships can be discovered, effectively improving the depth and accuracy of data quality monitoring. By calculating the correlation between data items to measure the correlation, the problem of difficulty in calculating the correlation relationship between multiple features in high-dimensional data is solved, and the calculation efficiency is improved. By calculating the mean of the kth feature of the data item and further calculating the correlation between data items to measure the correlation between data, the calculation complexity can be effectively reduced. By focusing on the mean of each feature and calculating the correlation at the feature level, the calculation burden when processing high-dimensional data is avoided, and the efficiency and scalability of the system are improved. By introducing the measurement method of correlation, the problem of difficult feature selection in multi-dimensional data quality analysis is solved, and the flexibility and adaptability of the model are improved. By introducing the correlation between data items, the importance of different features can be measured. Through correlation measurement, the system can dynamically adjust the contribution of each feature in data quality monitoring, thereby improving the flexibility and adaptability of the system. Regardless of how the number and dimensions of data features change, the system can quickly adjust and optimize the analysis strategy to provide more accurate monitoring results. By accurately calculating the correlation between data items, the problem of misjudgment in data monitoring is solved, and the accuracy and reliability of data quality detection are improved. By calculating the mean of the kth feature of the data item and measuring the correlation between the data items, the true relationship between the data items can be accurately reflected. This refined correlation calculation method effectively reduces the probability of misjudgment and ensures the high accuracy and reliability of data quality monitoring results. By accurately modeling the feature association between data, the system can more keenly discover data quality problems, take timely countermeasures, and improve the efficiency and effectiveness of data management.

[0157] According to the data life cycle and metadata mapping, intelligent data modeling based on adaptive algorithms is established, specifically including:

[0158] Let the i-th data item be D i ;

[0159] Get data item D i The kth feature ofk (D i );

[0160] Set the data modeling function to f(D i ), used to describe data item D i Features;

[0161] The modeling function is as follows:

[0162] Among them, f(D i ) is the data item D i Modeling function of ω k is the weight coefficient of the kth feature;

[0163] Get data item D i The kth true eigenvalue of k (D i );

[0164] Calculate data item D i To optimize the model:

[0165]

[0166] Set the optimization function to f opt (D i ), the optimization goal is to minimize the error function err(D i ), the optimization function is as follows:

[0167]

[0168] Among them, err(D i ) is the data item D i The error of opt (D i ) is the optimized objective function;

[0169] The optimization function is iteratively optimized by the gradient descent method, and the model parameters are updated as follows:

[0170]

[0171] Among them, θ k is the parameter of the kth iteration; η is the learning rate; To optimize the gradient of the function with respect to the parameters;

[0172] Automatically correct data quality through intelligent data quality management and predictive strategies.

[0173] By setting the data modeling function and optimizing the error, the problem of the inability to accurately capture data features and optimization in traditional data modeling is solved, and the accuracy and prediction ability of the model are improved. By setting the data modeling function and dynamically optimizing the weight coefficient of each feature, the characteristics of the data item can be more accurately described, thereby improving the accuracy of modeling. By calculating and optimizing the model error, it is ensured that the modeling function can adapt to the changes of different data features. The optimization goal is to minimize the error, thereby improving the accuracy of the prediction results and solving the problem that the traditional modeling method cannot dynamically adapt to data changes. By iteratively optimizing the optimization function and the gradient descent method, the slow convergence speed and local optimal problems in data modeling are solved, and the model training efficiency and global optimality are improved. By iteratively optimizing the optimization function, the model parameters are gradually updated. This method can not only improve the training efficiency, but also avoid the model from falling into the local optimal solution, ensuring that the global optimal model parameters are finally obtained. Through this optimization method, the efficiency and effect of the data modeling process can be greatly improved. Through intelligent data quality management and prediction strategies, data quality is automatically corrected, which solves the dilemma of insufficient manual intervention and difficulty in solving data quality problems in real time, and improves the automation and real-time performance of data quality management. Through the intelligent data quality management and prediction strategy in this solution, the system can automatically identify and correct data quality problems, thereby improving the automation of data quality management. The system dynamically adjusts and corrects data quality according to the mapping relationship between data life cycle and metadata, making the data quality management process more efficient and accurate, and avoiding the lag and limitations of manual intervention. Through the application of adaptive algorithms, the problem of inconsistent quality assessment of data items in different life cycle stages is solved, and the flexibility and accuracy of data quality monitoring are improved. Through intelligent data modeling based on adaptive algorithms, the modeling function and optimization strategy can be dynamically adjusted according to the life cycle stage of the data and different metadata characteristics, thereby improving the flexibility and accuracy of data quality assessment. Data at each stage has its own specific quality requirements. By automatically optimizing and correcting data quality, it can ensure that data at different stages can be efficiently and accurately monitored. Through the weight adjustment of data item features and error minimization optimization, the problem of feature importance evaluation in multi-feature modeling is solved, and the scientific nature of feature selection is improved. Through the weight adjustment of data item features, this solution can dynamically evaluate the contribution of each feature and automatically adjust the feature weight during the modeling process. By optimizing the model error, the influence of different features can be accurately reflected, which improves the scientificity and accuracy of feature selection and further optimizes the data modeling process.

[0174] The automatic correction of data quality through intelligent data quality management and prediction strategy specifically includes:

[0175] The data item D at time r i Recorded as

[0176] Get data items The effectiveness of For example, whether the data conforms to the expected format;

[0177] Get data items The consistency of For example, whether the data is consistent with data from other sources;

[0178] Get data items The completeness of For example, whether key fields are missing;

[0179] Set the quality measure function to qualitymeasure(D i ), specifically:

[0180]

[0181] Set the prediction data quality function to Q pred (t), Q pred (t) Predict the quality of data at time t in the future, specifically:

[0182]

[0183] Among them, ρ i For data item D i The weight of the data reflects its impact on the overall data quality; is the data item D at time t i Quality measures;

[0184] Set validity, consistency and completeness thresholds respectively;

[0185] Based on the prediction quality function Q pred (t), when the data quality is lower than any threshold, the automatic correction mechanism is activated, specifically:

[0186]

[0187] Where, ΔD i Represents data item D i The quality correction amount is corrected by machine learning method;

[0188] Protect data through data security policies.

[0189] By setting up quality measurement functions and predictive data quality functions, the problems of data quality assessment and prediction are solved, and the accuracy and predictability of data quality management are improved. By setting up quality measurement functions and predictive data quality functions, not only can the quality of data items be evaluated in real time, but also the quality of data at future moments can be predicted. This method provides more accurate quality assessment results by dynamically calculating the validity, consistency and integrity of data items. By predicting data quality, potential quality problems can be identified in a timely manner, measures can be taken in advance, and risks caused by data quality problems can be avoided. By defining validity, consistency and integrity thresholds, the dilemma of lack of reasonable threshold judgment in data quality management is solved, and the accuracy of data quality control is improved. By setting validity, consistency and integrity thresholds, this solution defines specific quality standards for each data item, which can accurately determine whether the data meets the expected quality requirements. This threshold-based judgment mechanism effectively solves the ambiguity problem of data quality assessment and enhances the accuracy and operability of data quality control. By starting the automatic correction mechanism based on the predicted quality function, the problems of insufficient manual intervention and untimely repair of data quality problems are solved, and the correction efficiency and data quality assurance capabilities are improved. Through the quality function based on prediction, the correction mechanism is automatically started when the data quality is found to be lower than the threshold. This mechanism corrects the data through machine learning methods, which can quickly and accurately solve data quality problems, improve the automation level of data quality repair, and ensure that data quality management is more efficient and real-time. Correcting data through machine learning methods solves the limitations of traditional manual correction methods and improves the intelligence and accuracy of correction. Automatic correction of data quality through machine learning methods can be dynamically adjusted according to data characteristics and historical data. Machine learning algorithms can automatically adjust correction strategies according to data changes in big data environments, avoiding the limitations of traditional methods. This intelligent data correction method not only improves the correction efficiency, but also improves the accuracy of data correction, ensuring continuous optimization of data quality. Protecting data through data security policies solves the security risks that may be caused during data quality correction and enhances the security and compliance of data processing. Through data security policies, data security is guaranteed while correcting data to prevent data from being improperly processed or leaked during the correction process. This strategy not only ensures the improvement of data quality, but also guarantees data compliance and privacy protection, improving the security of data governance.

[0190] The data security policy is used to protect the data, specifically including:

[0191] Get the public key, denoted as K pub ;

[0192] Encrypting data items is as follows: i )=E(Di ,K pub );

[0193] Among them, E is the encryption operation; D i is the data item to be encrypted;

[0194] The encrypted data hash value is stored in the blockchain to ensure that the data is not tampered with;

[0195] The data storage formula is: Store(H(D i ))=B(H(D i ),T);

[0196] Among them, H(D i ) is data D i The hash value of; B is the blockchain storage operation; T is the timestamp;

[0197] Set up an audit interface to record each data change and ensure that the changes are traceable;

[0198] The traceback formula is:

[0199] Among them, T trace H is the result of change traceability; i A hash value for each data change.

[0200] By protecting data security through encryption operations and public key mechanisms, the problem of unauthorized access or tampering of data during transmission and storage is solved, ensuring the confidentiality and integrity of data. By obtaining the public key and using encryption operations to encrypt the data items, the data items are encrypted and then stored, ensuring that the data can be effectively protected during transmission and storage. The encryption process prevents the data from being decrypted or tampered with without authorization, thereby ensuring the confidentiality and integrity of the data. In addition, data encryption also enhances data privacy protection and effectively prevents external attackers from maliciously acquiring or tampering with the data content. By storing data hash values ​​through blockchain technology, the problem of data tampering and loss is solved, and the immutability and traceability of data storage are improved. By storing the encrypted data hash value in the blockchain, the hash value of each data item is recorded on the blockchain. Blockchain has the characteristics of immutability and distributed storage, ensuring that data cannot be changed or deleted once written. This measure greatly improves the security and credibility of data storage and prevents data from being tampered with or lost. The data storage formula ensures that data records cannot be tampered with, solves the problem of data loss or tampering that may occur in traditional storage methods, and improves the security and transparency of data. Traditional database storage methods cannot effectively prevent data modification or deletion, and lack the record and audit of data changes. This solution uses a data storage formula to ensure that all changes in the data storage process will be recorded as the hash value of the data and stored on the blockchain. This not only enhances the security of data storage, but also improves the transparency of data operations. In the blockchain, each data change will be tracked and recorded, forming an audit chain that cannot be tampered with, ensuring the historical authenticity of data storage. Through the audit interface and change traceability function, the problem of insufficient transparency of data operations is solved, and the traceability and auditability of data changes are improved. By setting up an audit interface and a traceability formula, each data change is recorded, and the hash value of each change can be traced. The audit interface can not only realize real-time recording of data changes, but also effectively track historical changes, providing a complete audit trail for data operations. This mechanism greatly enhances the transparency and traceability of data management, improves the ability to supervise data changes, and ensures the compliance of data during management and use. By combining blockchain and encryption technology, the dual problems of data security and privacy protection are solved, and the security and compliance of the overall system are improved. By combining blockchain technology and encryption technology, not only the security and privacy of data are guaranteed, but also the compliance of data during storage and processing is ensured. Encryption technology ensures that data is not leaked during transmission and storage, while blockchain technology provides an unalterable data storage solution and ensures that every change in data can be tracked through the audit interface. This combined approach effectively solves the problems of data security and privacy protection and provides comprehensive protection for data management.

[0201] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0202] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A data quality monitoring method based on big data analysis, characterized in that: include: Divide data into different life cycle stages; The different life cycle stages include collection, cleaning, storage, use, analysis, archiving and destruction; The data life cycle state is recorded as L i (t), where i∈{1,2,...,n} represents different life cycle stages of a data item; Among them, L i (t) represents the life cycle state of data item i at time t; The priority of obtaining data, denoted as P i ; Develop a data scheduling strategy for each data item based on the data priority; The specific data scheduling strategy is: S i (t)=αP i +βC i (t); Among them, S i (t) represents the scheduling status of data item i at time t; C i (t) is the computing resource requirement of data item i at time t; α and β are weight coefficients; Record data status changes through logs and automatically adjust data status according to dynamic scheduling formulas; The dynamic scheduling formula is as follows: S i (t+1)=S i (t)+ΔS i ; Among them, ΔS i Indicates the amount of change in the state of data item i; Knowledge graph technology is used to model metadata, build a relationship map between data assets, and monitor data flow and data quality through dynamic mapping.

2. According to the data quality monitoring method based on big data analysis according to claim 1, it is characterized in that: The C i (t), including: Let the i-th data item be D i ; Get data item D i The storage space required at time t is denoted as Storage(D i ); Get data item D i The network bandwidth required at time t is denoted as DataTraffic (D i ); Then C i (t) = δ1·Storage(D i ) + δ2·DataTraffic(D i ); Among them, δ1 and δ2 are coefficients for adjusting the weights of storage space and network bandwidth.

3. The data quality monitoring method based on big data analysis according to claim 1 is characterized in that: The use of knowledge graph technology to model metadata, construct a graph of associations between data assets, and monitor data flow and data quality through dynamic mapping, specifically includes: Metadata about the collected data items, including the data’s unique identifier, source, type, and collection time; The metadata is recorded as: M i =(ID i ,type i ,source i ,timestamp i ); Among them, M i is the metadata of data item i; ID i is the unique identifier of data item i; type i is the type of data item i; source i The source of data item i; timestamp i is the collection time of data item i; Real-time access to data and metadata; By acquiring data and metadata in real time, the mapping relationship between data items is dynamically updated; Set a mapping matrix R i,j , specifically: R i,j =γ1·Similarity(M i ,M j )+γ2·Correlation(M i ,M j ); Among them, R i,j is the mapping relationship between data items i and j; Similarity(M i ,M j ) is used to calculate the similarity between data item i and data item j in the metadata space; Correlation(M i ,M j ) is used to measure the correlation between data item i and data item j; γ1 and γ2 are adjustment coefficients, which respectively control the weights of similarity and correlation in the mapping relationship; Build a knowledge graph based on the mapping relationship between data items; The icon structure is represented as: G = (V, E); Among them, V is a data node, representing each data item; E is an edge between data, representing the relationship between data items; Optimize the knowledge graph through the graph neural network model: E'=E+ΔE; Among them, ΔE represents the incremental correction of the relationship edge; Establish intelligent data modeling based on adaptive algorithms according to data lifecycle and metadata mapping.

4. The data quality monitoring method based on big data analysis according to claim 3 is characterized in that: Similarity(M i ,M j ), including: Get data item M i The kth feature of k (M i ); Get data item M j The kth feature of k (M j ); Get the total number of features in the metadata, denoted as m; The similarity calculation is as follows:

5. The data quality monitoring method based on big data analysis according to claim 4 is characterized in that: The Correlation (M i ,M j ), including: Get data item M i The mean of the kth feature is denoted by Get data item M j The mean of the kth feature is denoted by The correlation is measured by calculating the degree of association between data items, specifically:

6. The data quality monitoring method based on big data analysis according to claim 5 is characterized in that: According to the data life cycle and metadata mapping, intelligent data modeling based on adaptive algorithms is established, specifically including: Let the i-th data item be D i ; Get data item D i The kth feature of k (D i ); Set the data modeling function to f(D i ); The modeling function is as follows: Among them, f(D i ) is the data item D i Modeling function of ω k is the weight coefficient of the kth feature; Get data item D i The kth true eigenvalue of k (D i ); Calculate data item D i The error: Set the optimization function to f opt (D i ), the optimization function is as follows: Among them, err(D i ) is the data item D i The error of opt (D i ) is the optimized objective function; The optimization function is iteratively optimized by the gradient descent method, and the model parameters are updated as follows: Among them, θ k is the parameter of the kth iteration; η is the learning rate; To optimize the gradient of the function with respect to the parameters; Automatically correct data quality through intelligent data quality management and predictive strategies.

7. The data quality monitoring method based on big data analysis according to claim 6 is characterized in that: The automatic correction of data quality through intelligent data quality management and prediction strategy specifically includes: The data item D at time r i Recorded as Get data items The effectiveness of Get data items The consistency of Get data items The completeness of Set the quality measure function to qualitymeasure(D i ), specifically: Set the prediction data quality function to Q pred (t), specifically: Among them, ρ i For data item D i The weight of the data reflects its impact on the overall data quality; is the data item D at time t i Quality measures; Set validity, consistency and completeness thresholds respectively; Based on the prediction quality function Q pred (t), when the data quality is lower than any threshold, the automatic correction mechanism is activated, specifically: Where, ΔD i Represents data item D i The quality correction amount; Protect data through data security policies.

8. The data quality monitoring method based on big data analysis according to claim 7 is characterized in that: The data security policy is used to protect the data, specifically including: Get the public key, denoted as K pub ; Encrypting data items is as follows: i )=E(D i ,K pub ); Among them, E is the encryption operation; D i is the data item to be encrypted; Store the encrypted data hash value into the blockchain; The data storage formula is: Store(H(D i ))=B(H(D i ),T); Among them, H(D i ) is data D i The hash value of; B is the blockchain storage operation; T is the timestamp; Set up an audit interface to record each data change; The traceback formula is: Among them, T trace It is the result of change traceability; i A hash value for each data change.