An intelligent data quality monitoring method and system
By building an intelligent data quality monitoring method of unified data model and adaptive threshold module, the problem of not being able to identify complex data abnormalities in the existing technology is solved, real-time monitoring and efficient governance of multi-dimensional data quality are realized, and the accuracy and response speed of data quality detection are improved.
Patent Information
- Application Number
- CN202510703750.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing data quality monitoring methods and systems are unable to effectively identify complex data abnormal patterns, and lack the comprehensive evaluation and real-time monitoring capabilities of multi-dimensional data quality, resulting in data quality problems being ignored and unable to be processed in a timely manner, which may cause losses to the business.
Using intelligent data quality monitoring method, a unified data model is built based on Schema mapping rules, a sliding window statistical mechanism and a threshold dynamic adjustment algorithm is used to build an adaptive threshold module, a pre-trained model is used to perform multi-model fusion detection, a threshold judgment is used to combine the adaptive threshold module, a context-aware and dynamic priority management module is built for deduplication optimization, a root cause analysis is used for root cause analysis, a multi-dimensional evaluation module is generated for comprehensive scoring, and a cross-platform monitoring interface is built for real-time monitoring and early warning through the microservice architecture.
It realizes efficient processing and real-time abnormality detection of multi-source heterogeneous data, improves the accuracy and response speed of abnormality detection, reduces the false alarm rate, enhances the comprehensiveness and real-time nature of data quality governance, and improves data governance efficiency and troubleshooting speed.
Smart Images

Figure CN120234211B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data quality management, and particularly to an intelligent data quality monitoring method and system. Background Art
[0002] With the rapid development of information technology, data plays a crucial role in various fields. However, the quality of data directly affects data-based decision-making, analysis, and the normal operation of business processes. Existing data quality monitoring methods and systems have many deficiencies and are difficult to meet the growing complex data quality monitoring requirements.
[0003] Currently, common data quality monitoring methods mainly rely on preset fixed rules to simply verify data, such as checking whether the data format is correct and whether it is within a specific value range. For some simple data errors, such as inconsistent formats and values significantly exceeding the reasonable range, these methods can detect them to a certain extent. However, in the face of complex data anomaly patterns, especially potential problems hidden in a large amount of data, these traditional methods are inadequate.
[0004] Most existing data quality monitoring systems have relatively single functions. They often can only monitor specific types of data sources or specific data quality problems, lacking the comprehensive evaluation of multi-dimensional data quality and the ability of real-time monitoring and early warning.
[0005] Due to mainly relying on fixed rules, existing technologies cannot effectively identify new and complex anomaly patterns in data because these anomalies may not violate the preset simple rules, resulting in the neglect of data quality problems. Moreover, existing systems lack real-time performance. In the case of continuously increasing data volume and higher requirements for data timeliness in business, they cannot timely detect and handle data anomalies, which may cause serious losses to the business. In addition, existing systems are not comprehensive enough in data quality evaluation. They can only measure data quality from individual dimensions and cannot provide users with a comprehensive and accurate data quality status report. Summary of the Invention
[0006] In view of the above situation, the main purpose of the present invention is to propose an intelligent data quality monitoring method and system to solve the above technical problems.
[0007] The present invention proposes an intelligent data quality monitoring method, and the method includes the following steps:
[0008] Step 1: Construct a unified data model based on Schema mapping rules, and use the unified data model to process multi-source heterogeneous data to obtain a standardized distributed data stream;
[0009] Step 2: Construct an adaptive threshold module based on the sliding window statistical mechanism and the threshold dynamic adjustment algorithm. The adaptive threshold module and the pre-trained model form an anomaly detection engine, and the pre-trained model includes a supervised model and an unsupervised model;
[0010] Based on the random forest, SVM, and neural network algorithms, use the historical data as the training set and input it into the pre-trained model for training and optimization to obtain an optimized pre-trained model;
[0011] Perform multi-model fusion detection on the standardized distributed data stream using the optimized pre-trained model, and combine it with the adaptive threshold module for threshold determination processing to obtain a processed list of abnormal events;
[0012] Step 3: Construct a context awareness and dynamic priority management module based on the multi-source heterogeneous data fusion algorithm, the adaptive weight optimization algorithm, and the locality-sensitive hashing optimization algorithm. Input the processed list of abnormal events into the context awareness and dynamic priority management module for deduplication, optimization, and transmission to obtain a deduplicated and optimized list of abnormal events;
[0013] Step 4: Conduct root cause analysis on the deduplicated and optimized list of abnormal events using a root cause analysis tool to obtain the root cause analysis result;
[0014] Step 5: Construct a rule library based on general rules and preset rules. The rule library and the classification model form a repair recommendation engine, and use the repair recommendation engine to assist in generating a list of repair strategies after risk assessment based on the root cause analysis result;
[0015] Step 6: Construct a multi-dimensional evaluation module based on the weight allocation algorithm. Use the multi-dimensional evaluation module to perform refined index calculation and dynamic weight allocation processing on the standardized distributed data stream and the list of abnormal events, generate a comprehensive data quality score, and output a visualization report according to the comprehensive data quality score;
[0016] Step 7: Construct a unified console based on the microservice architecture. The unified console and the encapsulated adapter form a cross-platform monitoring interface. Input the list of repair strategies after risk assessment and the visualization report into the cross-platform monitoring interface to generate a warning notice.
[0017] The present invention also proposes an intelligent data quality monitoring system, and the system includes:
[0018] A data access layer for:
[0019] Construct a unified data model based on the Schema mapping rule, and use the unified data model to process multi-source heterogeneous data to obtain a standardized distributed data stream;
[0020] The first core processing layer for:
[0021] An adaptive threshold module is constructed based on a sliding window statistical mechanism and a threshold dynamic adjustment algorithm. The adaptive threshold module and a pre-trained model form an anomaly detection engine. The pre-trained model includes a supervised model and an unsupervised model. Based on algorithms such as random forest, SVM, and neural network, historical data is used as a training set and input into the pre-trained model for training and optimization to obtain an optimized pre-trained model.
[0022] The optimized pre-trained model is used to perform multi-model fusion detection on the standardized distributed data stream, and combined with the adaptive threshold module for threshold determination processing to obtain a processed list of abnormal events.
[0023] The intelligent analysis layer is used for:
[0024] A context-aware and dynamic priority management module is constructed based on multi-source heterogeneous data fusion algorithm, adaptive weight optimization algorithm, and locality-sensitive hashing optimization algorithm. The processed list of abnormal events is input into the context-aware and dynamic priority management module for deduplication optimization transmission to obtain a deduplicated and optimized list of abnormal events;
[0025] The root cause analysis tool is used to perform root cause analysis on the deduplicated and optimized list of abnormal events to obtain the root cause analysis result;
[0026] A rule library is constructed based on general rules and preset rules. The rule library and a classification model form a repair recommendation engine. The root cause analysis result is used to assist in generating a list of repair strategies after risk assessment by the repair recommendation engine;
[0027] The second core processing layer is used for:
[0028] A multi-dimensional evaluation module is constructed based on a weight allocation algorithm. The multi-dimensional evaluation module is used to perform refined index calculation and dynamic weight allocation processing on the standardized distributed data stream and the list of abnormal events to generate a comprehensive data quality score, and a visualization report is output according to the comprehensive data quality score;
[0029] The cross-platform management layer is used for:
[0030] A unified console is constructed based on a microservices architecture. The unified console and a packaging adapter form a cross-platform monitoring interface. The list of repair strategies after risk assessment and the visualization report are input into the cross-platform monitoring interface to generate a warning notice.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] 1. The present invention supports the access of multi-source heterogeneous data through a unified adapter architecture, and the deployment cycle is shortened by 60%;
[0033] 2. By adopting a pre-trained model and combining a neural network to capture temporal dependencies, the anomaly detection accuracy of the present invention is increased to over 95%. By introducing sliding window statistics and KS tests and dynamically adjusting the threshold interval, the false alarm rate is reduced from 15% of traditional methods to below 2%, adapting to scenarios of sudden changes in data distribution.
[0034] 3. By dynamically integrating multi-dimensional contexts, the present invention realizes intelligent classification and efficient transmission of anomaly events, significantly improving the accuracy and response speed of root cause analysis, reducing redundant processing, enhancing cross-scenario adaptability, and ensuring the comprehensiveness and real-time nature of data quality governance.
[0035] 4. By designing a comprehensive evaluation system covering accuracy (MAE ≤ 0.05), integrity (missing rate ≤ 5%), consistency (primary-foreign key conflict rate ≤ 1%), and timeliness (delay ≤ 1 second), the present invention supports dynamically allocating weights according to business scenarios, generating interactive visual reports, providing field-level quality drilling and root cause association diagrams, and increasing data governance efficiency by 40%.
[0036] 5. By constructing a data lineage graph based on a graph database, the present invention realizes minute-level tracing of anomaly events and shortens the troubleshooting time by 50%.
[0037] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flowchart of the steps of an intelligent data quality monitoring method proposed by the present invention.
[0039] Figure 2 is an overall architecture diagram of an intelligent data quality monitoring method proposed by the present invention.
[0040] Figure 3 is an overall framework diagram of an intelligent data quality monitoring system proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals are the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0042] These and other aspects of the embodiments of the present invention will be apparent from the following description and the accompanying drawings. In these descriptions and drawings, specific embodiments of some embodiments of the present invention are specifically disclosed as some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0043] Please refer to Figure 1 , an embodiment of the present invention provides an intelligent data quality monitoring method, which includes the following steps:
[0044] Step 1: Construct a unified data model based on the Schema mapping rule, and use the unified data model to process multi-source heterogeneous data to obtain a standardized distribution data stream.
[0045] In step 1, using the unified data model to process multi-source heterogeneous data to obtain a standardized distribution data stream specifically includes the following steps:
[0046] Parse the format of the multi-source heterogeneous data to obtain the parsed multi-source heterogeneous data;
[0047] Successively perform format conversion, redundancy cleaning, and missing value processing on the parsed multi-source heterogeneous data to obtain the processed data;
[0048] Partition the processed data using the hash partitioning strategy to obtain the partitioned data, and distribute and back up the partitioned data using the dynamic scaling strategy to obtain a standardized distribution data stream.
[0049] In the step of parsing the format of the multi-source heterogeneous data to obtain the parsed multi-source heterogeneous data, the following relational expressions exist for the corresponding process:
[0050] ;
[0051] Among them, represents a logical quantifier, represents the th heterogeneous data source, represents the set of heterogeneous data sources, represents the th parsed multi-source heterogeneous data, represents the set of heterogeneous data parsing function sets;
[0052] In the step of successively performing format conversion, redundancy cleaning, and missing value processing on the parsed multi-source heterogeneous data to obtain the processed data, the following relational expressions exist for the corresponding process:
[0053] ;
[0054] Among them, Represents the processed data, Represents the metadata mapping function, Represents the redundancy elimination function, Represents the timestamp alignment function, Represents the data stream merging operation, Represents the schema alignment operation, Represents the total number of data streams participating in the processing;
[0055] In the step of partitioning the processed data using the hash partitioning strategy to obtain the partitioned data, and then distributing and backing up the partitioned data using the dynamic scaling strategy to obtain the standardized distributed data stream, the following relational expressions exist for the corresponding process:
[0056] ;
[0057] Among them, Represents the processed data The th data item in, Represents the allocated data set, Represents the data set allocated to partition , Represents the consistent hashing function, Represents the modulo operation, Represents the partition key, Represents the constraint condition, Represents finding the parameter value that reaches the minimum value, Represents the dynamic number of partitions, Represents the routing policy function, Represents routing to node 's data set, Represents the set of available nodes, Represents node At time 's comprehensive load index, Represents the average load, Represents the standardized distributed data stream, Represents at time 's number of available computing nodes, Represents the node-level standardized operation chain, Represents the global data stream merging operation, Represents the encryption process, Represents the compression process, Represents the serialization process, Represents the global data stream merging operation.
[0058] Step 2: Construct an adaptive threshold module based on the sliding window statistical mechanism and the threshold dynamic adjustment algorithm. The adaptive threshold module and the pre-trained model form an anomaly detection engine. The pre-trained model includes a supervised model and an unsupervised model;
[0059] Based on the random forest, SVM, and neural network algorithms, use the historical data as the training set and input it into the pre-trained model for training and optimization to obtain an optimized pre-trained model;
[0060] Perform multi-model fusion detection on the standardized distributed data stream using the optimized pre-trained model, and combine it with the adaptive threshold module for threshold determination processing to obtain a processed list of anomaly events.
[0061] Please refer to Figure 2 , in Step 2, perform multi-model fusion detection on the standardized distributed data stream using the optimized pre-trained model, and combine it with the adaptive threshold module for threshold determination processing to obtain a processed list of anomaly events, which specifically includes the following steps:
[0062] Perform feature extraction and text encoding on the standardized distributed data stream in sequence to obtain a set of feature vectors;
[0063] Input the set of feature vectors into the supervised model for prediction to obtain an anomaly probability;
[0064] Input the set of feature vectors into the unsupervised model for calculation to obtain an anomaly score;
[0065] Perform multi-model fusion detection on the anomaly probability and the anomaly score to obtain a preliminary list of anomaly marks;
[0066] Perform sliding window statistics, threshold dynamic adjustment processing, and KS test processing on the preliminary list of anomaly marks in sequence using the adaptive threshold module to obtain a list of anomaly events after fine screening;
[0067] Perform data source merging processing and additional metadata processing on the list of anomaly events after fine screening in sequence to obtain a processed list of anomaly events.
[0068] In the step of performing feature extraction and text encoding on the standardized distributed data stream in sequence to obtain a set of feature vectors, the following relational expressions exist for the corresponding process:
[0069] ;
[0070] Among them, represents the set of feature vectors, represents the total data volume, represents the data item index, represents the TF-IDF feature mapping function, represents the BERT encoder output mapping, represents the feature cross - fusion operator, represents the data unit in the standardized distribution data stream set, represents the standardized distribution data stream set.
[0071] In the step of inputting the feature vector set into the supervised model for prediction, the relational expressions in the corresponding process are as follows:
[0072] ;
[0073] Among them, represents the anomaly probability output by the supervised model, represents the Sigmoid activation function, represents the weight matrix of the supervised model, represents the bias term of the supervised model;
[0074] In the step of inputting the feature vector set into the unsupervised model for calculation to obtain the anomaly score, the relational expressions in the corresponding process are as follows:
[0075] ;
[0076] Among them, represents the anomaly score of the unsupervised model, represents the number of unsupervised models, represents the feature vector index, represents the auto - encoder's reconstruction error for the th feature vector;
[0077] In the step of performing multi - model fusion detection on the anomaly probability and anomaly score, the relational expressions in the corresponding process are as follows:
[0078] ;
[0079] Among them, represents the comprehensive anomaly score, and respectively represent two weight coefficients, represents the anomaly probability output by the supervised model, represents the anomaly score output by the unsupervised model;
[0080] In the step of successively performing sliding window statistics, threshold dynamic adjustment processing, and KS - test processing on the preliminary anomaly marking list using the adaptive threshold module, the relational expressions in the corresponding process are as follows:
[0081] ;
[0082] Among them, Represents the average score within the window, Represents the standard deviation within the window, Represents the window size, Represents the threshold value after dynamic adjustment, Represents the sensitivity coefficient, Represents the KS test statistic, Represents the supremum, Represents the cumulative distribution function of the current window, Represents the cumulative distribution function of the historical window, Represents a variable.
[0083] Step 3: Construct a context-aware and dynamic priority management module based on the multi-source heterogeneous data fusion algorithm, the adaptive weight optimization algorithm, and the locality-sensitive hashing optimization algorithm, and input the processed list of abnormal events into the context-aware and dynamic priority management module for deduplication optimization transmission to obtain a deduplicated and optimized list of abnormal events.
[0084] In step 3, inputting the processed list of abnormal events into the context-aware and dynamic priority management module for deduplication optimization transmission to obtain a deduplicated and optimized list of abnormal events specifically includes the following steps:
[0085] Perform real-time collection of context information on the processed list of abnormal events to obtain a list of abnormal events with context information;
[0086] Perform comprehensive severity score calculation, dynamic weight optimization, and priority classification processing on the list of abnormal events with context information in sequence to obtain a list of abnormal events with priority labels and scores;
[0087] Perform similarity calculation, event repetition merging processing, and deduplication transmission optimization processing on the list of abnormal events with priority labels and scores in sequence to obtain a deduplicated and optimized list of abnormal events.
[0088] In the step of performing real-time collection of context information on the processed list of abnormal events, the following relational expressions exist for the corresponding process:
[0089] ;
[0090] Among them, Represents the quantitative evaluation value of the context environment of the th abnormal event, Represents the time correlation score, Represents the space correlation score, Represents the third weight coefficient, Represents the resource correlation score;
[0091] In the steps of successively performing comprehensive severity score calculation, dynamic weight optimization, and priority classification processing on the list of abnormal events with context information, the relational expressions in the corresponding process are as follows:
[0092] ;
[0093] Among them, represents the comprehensive severity score of the th abnormal event, , , and respectively represent four dynamic weights, represents the event weight type, represents the impact scope score, represents the duration score, represents the dynamic value of the th dynamic weight at time , represents the initial weight, represents the decay coefficient, represents the natural constant, represents the exponential decay function, represents the summation index, represents the normalized score, represents the minimum score in the current abnormal event list, represents the maximum score in the current abnormal event list;
[0094] It should be noted that, .
[0095] In the steps of successively performing similarity calculation, event duplicate merging processing, and duplicate removal and transmission optimization processing on the list of abnormal events with priority labels and scores to obtain the list of abnormal events after duplicate removal and optimization, the relational expressions in the corresponding process are as follows:
[0096] ;
[0097] Among them, represents the cosine similarity between the th abnormal event and the th abnormal event in the abnormal event list, represents the feature vector of the th abnormal event, represents the feature vector of the th abnormal event, represents the vector norm, represents the event merging processing of the th abnormal event and the th abnormal event, Represents the new merged event, Represents the similarity threshold, Represents the maximum value operator, Represents the new merged event The comprehensive severity score of, Represents the Comprehensive severity score of the Represents the new merged event Priority label of, Represents the highest priority level operator, Represents the Priority label of the Represents the Priority label of the Represents the list of anomaly events after deduplication optimization The comprehensive transmission cost of, Represents the list of anomaly events after deduplication optimization, Or the list of anomaly events The th event instance of, Represents the single-event transmission bandwidth requirement, Represents the penalty coefficient, Represents the counting operator, Represents the difference set operator, Represents the bandwidth resources saved after deduplication.
[0098] Step 4: Use a root cause analysis tool to perform root cause analysis on the list of anomaly events after deduplication optimization to obtain the root cause analysis result.
[0099] In Step ⑷, use a root cause analysis tool to perform root cause analysis on the list of anomaly events after deduplication optimization to obtain the root cause analysis result, which specifically includes the following steps:
[0100] Perform field parsing processing and data item association processing on the list of anomaly events after deduplication optimization in sequence to obtain the list of associated data items for anomaly events;
[0101] Use a graph database to perform reverse tracing processing and path screening on the list of associated data items for anomaly events and the list of anomaly events after deduplication optimization in sequence to obtain the anomaly field bloodline path graph;
[0102] Based on historical anomaly event data, construct a Bayesian network model, input the anomaly field bloodline path graph into the Bayesian network model for root cause ranking, and obtain the list of root cause candidates;
[0103] Perform root cause verification processing on the list of root cause candidates to obtain the verified result;
[0104] Confirm the root cause label for the verified result to obtain the root cause analysis result.
[0105] In the steps of sequentially performing field parsing processing and data item association processing on the deduplicated and optimized abnormal event list, the relational expressions in the corresponding process are as follows:
[0106] ;
[0107] Among them, represents the field association function, represents the th field of the deduplicated and optimized abnormal event list, represents the th field of the deduplicated and optimized abnormal event list, represents the field and the field 's co-occurrence times in the historical data, represents the total number of times the field appears alone, represents the total number of times the field appears alone, represents the set of strongly associated field pairs, represents the association degree threshold;
[0108] In the steps of sequentially performing reverse tracing processing and path screening on the abnormal event associated data item list and the deduplicated and optimized abnormal event list using a graph database to obtain the abnormal field blood relationship path graph, the relational expressions in the corresponding process are as follows:
[0109] ;
[0110] Among them, represents the path weight, represents the path from the source node to the target node , represents the confidence of the dependency edge from the source node to the target node , represents the path weight threshold, represents the set of abnormal field blood relationship path graphs;
[0111] In the step of inputting the abnormal field blood relationship path graph into the Bayesian network model for root cause ranking, the relational expressions in the corresponding process are as follows:
[0112] ;
[0113] Among them, represents the Candidate root causes The root cause score of represents the prior probability of the candidate root cause ; represents the conditional probability of observing the blood relationship path under the candidate root cause ; represents the marginal probability of observing the set of blood relationship path graphs of abnormal fields ; represents the list of root causes sorted by score represents the upper limit of the number of root cause candidates represents selecting the result with the highest score from the candidate set;
[0114] In the step of performing root cause verification processing on the root cause candidate list, the following relational expressions exist in the corresponding process:
[0115] ;
[0116] Among them, represents the verification confidence of the candidate root cause ; represents the number of times the candidate root cause truly causes an abnormality in the verification set ; represents the number of times the predicted candidate root cause
[0117] is the root cause;
[0118] ;
[0119] Among them, represents the final root cause label of the event ; represents the set of root causes that pass the verification represents the set intersection ; represents the set of blood relationship paths of the event in the graph
[0120] Step 5: Build a rule base based on general rules and preset rules. The rule base and the classification model constitute a repair recommendation engine, and use the repair recommendation engine to assist in generating a list of repair strategies after risk assessment for the root cause analysis result.
[0121] In Step 5, using the repair recommendation engine to assist in generating a list of repair strategies after risk assessment for the root cause analysis result specifically includes the following steps:
[0122] Match the preset repair rules from the rule base based on the root cause analysis result to generate a basic repair plan;
[0123] The root cause feature extraction and historical repair record feature extraction are sequentially performed on the root cause analysis result and the basic repair plan, and the extracted root cause features and the extracted historical repair record features are obtained respectively;
[0124] A feature vector is constructed based on the extracted root cause features and the extracted historical repair record features to obtain a comprehensive feature vector;
[0125] The comprehensive feature vector is input into a pre-trained classification model for prediction to obtain a repair strategy score, and the repair strategy scores are sorted to obtain a machine learning recommended strategy list;
[0126] The basic repair plan and the machine learning recommended strategy list are subjected to strategy merging processing, and weighted score calculation and screening processing are sequentially performed to obtain a fused repair strategy list;
[0127] A matching risk control rule is generated according to the rule base, and the fused repair strategy list is used to calculate the risk score using the matching risk control rule to obtain the calculated risk score;
[0128] The calculated risk score is subjected to strategy correction processing to obtain a risk-assessed repair strategy list.
[0129] In the step of matching a preset repair rule from the rule base based on the root cause analysis result to generate a basic repair plan, the relational expressions existing in the corresponding process are as follows:
[0130] ;
[0131] Among them, represents the th matched basic repair plan, represents the rule matching algorithm, represents the predefined repair rule base, represents the root cause analysis result, represents the index set of candidate repair plans;
[0132] In the step of sequentially performing root cause feature extraction and historical repair record feature extraction on the root cause analysis result and the basic repair plan, the relational expressions existing in the corresponding process are as follows:
[0133] ;
[0134] Among them, represents the numerical feature vector of the root cause, represents the feature encoder, represents the historical repair record feature, represents the historical data feature extractor, Indicates the historical repair dataset;
[0135] In the step of constructing a feature vector based on the extracted root cause features and the extracted historical repair record features to obtain a comprehensive feature vector, the relational expressions existing in the corresponding process are as follows:
[0136] ;
[0137] Among them, Indicates the comprehensive feature vector, Indicates the concatenation operation;
[0138] In the step of performing strategy merging processing on the basic repair plan and the machine learning recommendation strategy list, and sequentially performing weighted scoring calculation and screening processing to obtain the merged repair strategy list, the relational expressions existing in the corresponding process are as follows:
[0139] ;
[0140] Among them, Indicates the merged comprehensive score, Indicates the rule priority score of the basic repair plan, Indicates the merged repair strategy list, Indicates the screening function, Indicates the maximum number of retained strategies;
[0141] In the step of generating matching risk control rules according to the rule library, and performing risk scoring calculation on the merged repair strategy list using the matching risk control rules to obtain the calculated risk score, the relational expressions existing in the corresponding process are as follows:
[0142] ;
[0143] Among them, Indicates the subset of matching risk rules, Indicates the risk rule library, Indicates the risk score calculated for the th strategy, Indicates the weight of rule <00> Indicates the rule violation check function, Indicates the th strategy in the merged strategy list;
[0144] In the step of performing strategy correction processing on the calculated risk score to obtain the risk-assessed repair strategy list, the relational expressions existing in the corresponding process are as follows:
[0145] ;
[0146] Among them, Represents a list of repair strategies after risk assessment, Represents a filtering function, Represents the risk tolerance threshold.
[0147] It should be noted that the classification model belongs to the prior art, such as Random Forest, Support Vector Machine (SVM), or Neural Network, etc. These models have been widely applied in data classification and prediction scenarios. They can be directly invoked through publicly available machine learning frameworks, and their implementation principles and training methods are well-known in the prior art.
[0148] Step 6: Construct a multi-dimensional evaluation module based on the weight assignment algorithm, and use the multi-dimensional evaluation module to perform refined index calculation and dynamic weight assignment processing on the standardized distributed data stream and the abnormal event list, generate a comprehensive data quality score, and output a visualization report according to the comprehensive data quality score.
[0149] In Step 6, using the multi-dimensional evaluation module to perform refined index calculation and dynamic weight assignment processing on the standardized distributed data stream and the abnormal event list, generate a comprehensive data quality score, and output a visualization report according to the comprehensive data quality score, which specifically includes the following steps:
[0150] Use the multi-dimensional evaluation module to perform index calculations on the standardized distributed data stream and the abnormal event list for accuracy, integrity, consistency, timeliness, uniqueness, and compliance in sequence to obtain a real-time data quality score;
[0151] Generate business scenario labels through clustering data and manual presetting, map the business scenario labels to business priorities in combination with a preset weight template library to obtain an initial weight vector;
[0152] Based on historical data quality scores, construct a business loss, and use the gradient descent optimization algorithm in combination with the initial weight vector to perform dynamic weight updates on the real-time data quality score and the business loss to obtain an optimized weight vector;
[0153] Perform a comprehensive score calculation on the real-time data quality score in combination with the optimized weight vector to generate a comprehensive data quality score;
[0154] Perform comprehensive score display, multi-dimensional score comparison, and construction of an abnormal event heat map on the comprehensive data quality score, the real-time data quality score, and the abnormal event list in sequence to obtain an interactive dashboard;
[0155] Perform causal chain diagram construction and visualization rendering processing on the abnormal event list and the abnormal field lineage path graph in sequence to obtain a root cause analysis report;
[0156] Predict the historical data quality score and historical business losses using the LSTM model to obtain the trend prediction result;
[0157] Convert the formats of the interactive dashboard, root cause analysis report, and trend prediction result to obtain a visualization report.
[0158] In the step of calculating the indicators of accuracy, integrity, consistency, timeliness, uniqueness, and compliance for the standardized distribution data stream and the abnormal event list using the multi-dimensional evaluation module in sequence to obtain the real-time data quality score, the relational expressions existing in the corresponding process are as follows:
[0159] ;
[0160] Among them, represents the real-time data quality score, represents the accuracy score, represents the integrity score, represents the consistency score, represents the timeliness score, represents the uniqueness score, represents the compliance score;
[0161] In the step of generating business scenario labels through clustering data and manual presetting, and performing business priority mapping on the business scenario labels in combination with the preset weight template library to obtain the initial weight vector, the relational expressions existing in the corresponding process are as follows:
[0162] ;
[0163] Among them, represents the initial weight vector, represents the normalization function, represents the preset weight template library, represents the clustering analysis algorithm, represents the clustering feature matrix, represents the union operator, represents the manually preset business scenario labels;
[0164] In the step of constructing business losses based on the historical data quality score, and performing dynamic weight update on the real-time data quality score and business losses using the gradient descent optimization algorithm in combination with the initial weight vector to obtain the optimized weight vector, the relational expressions existing in the corresponding process are as follows:
[0165] ;
[0166] Among them, represents the optimized weight vector, represents the 6-dimensional probability space Weight vector Obtain the optimal parameter values, Indicates the business loss, Indicates the historical scoring matrix, Indicates the weight vector The transpose of, Indicates the regularization coefficient, Indicates the regularization term, Indicates the L2 norm;
[0167] In the step of calculating the comprehensive score of data quality by combining the real-time data quality score with the optimized weight vector to generate the comprehensive data quality score, the relational expressions in the corresponding process are as follows:
[0168] ;
[0169] Among them, Indicates the comprehensive data quality score, Indicates the transpose of the optimized weight vector, Indicates the Optimal weight vector optimized in the Indicates the Real-time quality score in the
[0170] In the step of successively performing comprehensive score display, multi-dimensional score comparison, and abnormal event heat map construction on the comprehensive data quality score, real-time data quality score, and abnormal event list to obtain an interactive dashboard, the relational expressions in the corresponding process are as follows:
[0171] ;
[0172] Among them, Indicates the interactive dashboard, Indicates the visualization function;
[0173] It should be noted that, Includes instrument analysis panel, radar comparison chart, and heat map rendering.
[0174] In the step of successively performing causal chain diagram construction and visualization rendering on the abnormal event list and the abnormal field blood relationship path map to obtain a root cause analysis report, the relational expressions in the corresponding process are as follows:
[0175] ;
[0176] Among them, Indicates the root cause analysis report, Indicates the causal reasoning algorithm, Indicates the field blood relationship map, Indicates the rendering visualization process, Indicates generating a report document, Indicates saving an image file;
[0177] In the step of predicting the historical data quality score and historical business losses using the LSTM model, the relational expressions in the corresponding process are as follows:
[0178] ;
[0179] Among them, Indicates the future prediction score, Indicates the future loss prediction, Indicates the LSTM model, Indicates the time window, Indicates the LSTM model parameters;
[0180] In the step of performing format conversion on the interactive dashboard, root cause analysis report, and trend prediction results to obtain a visualization report, the relational expressions in the corresponding process are as follows:
[0181] ;
[0182] Among them, Indicates the visualization report, Indicates the format conversion function, Indicates outputting the optimization results and visualization content in a web format.
[0183] Step 7: Build a unified console based on the microservice architecture. The unified console and the encapsulated adapter form a cross-platform monitoring interface. Input the list of repair strategies after risk assessment and the visualization report into the cross-platform monitoring interface to generate a warning notification.
[0184] Furthermore, the warning notification includes the exception type, the affected scope, and the repair suggestions. The repair suggestions are executed by the repair executor after being confirmed by the user.
[0185] Please refer to Figure 3 , the present invention also proposes an intelligent data quality monitoring system, and the system includes:
[0186] The data access layer is used for:
[0187] Build a unified data model based on the Schema mapping rule, and use the unified data model to process multi-source heterogeneous data to obtain a standardized distribution data stream;
[0188] The first core processing layer is used for:
[0189] Build an adaptive threshold module based on the sliding window statistical mechanism and the threshold dynamic adjustment algorithm. The adaptive threshold module and the pre-trained model form an anomaly detection engine. The pre-trained model includes a supervised model and an unsupervised model;
[0190] Based on the random forest, SVM, and neural network algorithms, the historical data is used as the training set and input into the pre-trained model for training and optimization to obtain an optimized pre-trained model;
[0191] The optimized pre-trained model is used to perform multi-model fusion detection on the standardized distributed data stream, and combined with the adaptive threshold module for threshold determination processing to obtain a processed list of abnormal events;
[0192] The intelligent analysis layer is used for:
[0193] Based on the multi-source heterogeneous data fusion algorithm, the adaptive weight optimization algorithm, and the locality-sensitive hashing optimization algorithm, a context awareness and dynamic priority management module is constructed. The processed list of abnormal events is input into the context awareness and dynamic priority management module for deduplication and optimized transmission to obtain a deduplicated and optimized list of abnormal events;
[0194] The root cause analysis tool is used to perform root cause analysis on the deduplicated and optimized list of abnormal events to obtain the root cause analysis result;
[0195] Based on the general rules and preset rules, a rule library is constructed. The rule library and the classification model constitute a repair recommendation engine, and the root cause analysis result is used to assist in generating a list of repair strategies after risk assessment using the repair recommendation engine;
[0196] The second core processing layer is used for:
[0197] Based on the weight allocation algorithm, a multi-dimensional evaluation module is constructed. The multi-dimensional evaluation module is used to perform refined index calculation and dynamic weight allocation processing on the standardized distributed data stream and the list of abnormal events, generate a comprehensive data quality score, and output a visual report according to the comprehensive data quality score;
[0198] The cross-platform management layer is used for:
[0199] Based on the microservices architecture, a unified console is constructed. The unified console and the encapsulated adapter constitute a cross-platform monitoring interface. The list of repair strategies after risk assessment and the visual report are input into the cross-platform monitoring interface to generate a warning notice.
[0200] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0201] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0202] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. An intelligent data quality monitoring method, characterized in that The method includes the following steps: Step 1: Construct a unified data model based on the Schema mapping rule, and use the unified data model to process multi-source heterogeneous data to obtain a standardized distribution data stream; Step 2: Construct an adaptive threshold module based on the sliding window statistical mechanism and the threshold dynamic adjustment algorithm. The adaptive threshold module and the pre-trained model form an anomaly detection engine. The pre-trained model includes a supervised model and an unsupervised model; Based on the random forest, SVM, and neural network algorithms, use the historical data as the training set to input into the pre-trained model for training and optimization to obtain an optimized pre-trained model; Use the optimized pre-trained model to perform multi-model fusion detection on the standardized distribution data stream, and combine with the adaptive threshold module to perform threshold determination processing to obtain a processed list of anomaly events; Step 3: Construct a context awareness and dynamic priority management module based on the multi-source heterogeneous data fusion algorithm, the adaptive weight optimization algorithm, and the locality-sensitive hashing optimization algorithm. Input the processed list of anomaly events into the context awareness and dynamic priority management module for deduplication optimization transmission to obtain a deduplicated and optimized list of anomaly events; Step 4: Use a root cause analysis tool to perform root cause analysis on the deduplicated and optimized list of anomaly events to obtain the root cause analysis result; Step 5: Construct a rule library based on general rules and preset rules. The rule library and the classification model form a repair suggestion engine. Use the repair suggestion engine to assist in generating a list of repair strategies after risk assessment for the root cause analysis result; Step 6: Construct a multi-dimensional evaluation module based on the weight allocation algorithm. Use the multi-dimensional evaluation module to perform refined index calculation and dynamic weight allocation processing on the standardized distribution data stream and the list of anomaly events to generate a comprehensive data quality score, and output a visualization report according to the comprehensive data quality score; Step 7: Construct a unified console based on the microservice architecture. The unified console and the encapsulation adapter form a cross-platform monitoring interface. Input the list of repair strategies after risk assessment and the visualization report into the cross-platform monitoring interface to generate a warning notice.
2. The intelligent data quality monitoring method according to claim 1, wherein In the said Step 1, using the unified data model to process the multi-source heterogeneous data to obtain a standardized distribution data stream specifically includes the following steps: Perform format parsing on the multi-source heterogeneous data to obtain the parsed multi-source heterogeneous data. The relational expressions existing in the corresponding process are as follows: ; Among them, represents a logical quantifier, represents the th heterogeneous data source, represents a set of heterogeneous data sources, represents the th parsed multi-source heterogeneous data, represents a set of heterogeneous data parsing functions; Perform format conversion, redundancy cleaning, and missing value processing on the parsed multi-source heterogeneous data in sequence to obtain the processed data. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the processed data, represents the metadata mapping function, represents the redundancy elimination function, represents the timestamp alignment function, represents the data stream merging operation, represents the schema alignment operation, represents the total number of data streams participating in the processing; Perform data partitioning on the processed data using the hash partitioning strategy to obtain the partitioned data. Perform distribution and backup on the partitioned data using the dynamic scaling strategy to obtain a standardized distribution data stream. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the nth data item in the processed data, represents the allocated data set, represents the data set allocated to partition , represents the consistent hashing function, represents the modulo operation, represents the partition key, represents the constraint condition, represents finding the parameter value that reaches the minimum value, represents the dynamic number of partitions, represents the routing policy function, represents the data set routed to node , represents the set of available nodes, represents node at time with the comprehensive load index, represents the average load, represents the normalized distribution data stream, represents at time the number of available computing nodes, represents the node-level normalized operation chain, represents the global data stream merging operation, represents the encryption process, represents the compression process, represents the serialization process, represents the global data stream merging operation.
3. The intelligent data quality monitoring method according to claim 2, wherein In the said Step 2, using the optimized pre-trained model to perform multi-model fusion detection on the standardized distribution data stream, and combining with the adaptive threshold module to perform threshold determination processing to obtain a processed list of anomaly events specifically includes the following steps: Perform feature extraction and text encoding on the standardized distribution data stream in sequence to obtain a set of feature vectors; Input the set of feature vectors into a supervised model for prediction to obtain an anomaly probability; Input the set of feature vectors into an unsupervised model for calculation to obtain an anomaly score; Perform multi-model fusion detection on the anomaly probability and anomaly score to obtain a preliminary anomaly label list; Perform sliding window statistics, threshold dynamic adjustment processing, and KS test processing on the preliminary anomaly label list in sequence using an adaptive threshold module to obtain a refined anomaly event list; Perform data source merging processing and additional metadata processing on the refined anomaly event list in sequence to obtain a processed anomaly event list.
4. The intelligent data quality monitoring method according to claim 3, characterized in that In step 3, input the processed anomaly event list into a context-aware and dynamic priority management module for deduplication and optimization transmission to obtain a deduplicated and optimized anomaly event list, which specifically includes the following steps: Collect context information of the processed anomaly event list in real time to obtain an anomaly event list with context information; Perform comprehensive severity score calculation, dynamic weight optimization, and priority classification processing on the anomaly event list with context information in sequence to obtain an anomaly event list with priority labels and scores; Perform similarity calculation, event duplicate merging processing, and deduplication and transmission optimization processing on the anomaly event list with priority labels and scores in sequence to obtain a deduplicated and optimized anomaly event list.
5. The intelligent data quality monitoring method according to claim 4, characterized in that Collect context information of the processed anomaly event list in real time, and the relational expressions existing in the corresponding process are as follows: ; Among them, represents the quantitative evaluation value of the context of the th abnormal event, represents the time correlation score, represents the space correlation score, , and respectively represent three weight coefficients, represents the resource correlation score; In the steps of performing comprehensive severity score calculation, dynamic weight optimization, and priority classification processing on the anomaly event list with context information in sequence, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the comprehensive severity score of the th abnormal event, , , and respectively represent four dynamic weights, represents the event weight type, represents the impact range score, represents the duration score, represents the th dynamic weight at time , represents the initial weight, represents the decay coefficient, represents the natural constant, represents the exponential decay function, represents the summation index, represents the normalized score, represents the minimum score in the current abnormal event list, represents the maximum score in the current abnormal event list; In the steps of performing similarity calculation, event duplicate merging processing, and deduplication and transmission optimization processing on the anomaly event list with priority labels and scores to obtain a deduplicated and optimized anomaly event list, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the cosine similarity between the th and the th abnormal events in the abnormal event list. represents the feature vector of the th abnormal event. represents the feature vector of the th abnormal event. represents the vector norm. represents the event merging process of the th and the th abnormal events. represents the new event after merging. represents the similarity threshold. represents the maximum operator. represents the new event after merging 's comprehensive severity score. represents the comprehensive severity score of the th abnormal event. represents the priority label of the new event after merging ; represents the highest priority operator. represents the priority label of the th abnormal event. represents the priority label of the th abnormal event. represents the comprehensive transmission cost of the abnormal event list after deduplication optimization ; represents the abnormal event list after deduplication optimization, or the th event instance of the abnormal event list ; represents the single-event transmission bandwidth requirement. represents the penalty coefficient. represents the counting operator. represents the difference set operator. represents the bandwidth resources saved after deduplication.
6. The intelligent data quality monitoring method according to claim 5, wherein In step 4, perform root cause analysis on the deduplicated and optimized anomaly event list using a root cause analysis tool to obtain a root cause analysis result, which specifically includes the following steps: Perform field parsing processing and data item association processing on the deduplicated and optimized anomaly event list in sequence to obtain an anomaly event associated data item list; Perform reverse tracing processing and path screening on the anomaly event associated data item list and the deduplicated and optimized anomaly event list using a graph database in sequence to obtain an anomaly field blood relationship path map; Construct a Bayesian network model based on historical anomaly event data, input the anomaly field blood relationship path map into the Bayesian network model for root cause ranking to obtain a root cause candidate list; Perform root cause verification processing on the root cause candidate list to obtain a verified result; Perform root cause label confirmation on the verified result to obtain a root cause analysis result.
7. The intelligent data quality monitoring method according to claim 6, characterized in that In step 5, use a repair suggestion engine to assist in generating a risk assessment-based repair strategy list for the root cause analysis result, which specifically includes the following steps: Match a preset repair rule from a rule library based on the root cause analysis result to generate a basic repair plan; Perform root cause feature extraction and historical repair record feature extraction on the root cause analysis results and the basic repair plan in sequence, and obtain the extracted root cause features and the extracted historical repair record features respectively; Construct a feature vector based on the extracted root cause features and the extracted historical repair record features to obtain a comprehensive feature vector; Input the comprehensive feature vector into a pre-trained classification model for prediction to obtain a repair strategy score, and sort the repair strategy scores to obtain a machine learning recommended strategy list; Perform strategy merging processing on the basic repair plan and the machine learning recommended strategy list, and perform weighted score calculation and screening processing in sequence to obtain a merged repair strategy list; Generate matching risk control rules according to the rule library, and calculate the risk score for the merged repair strategy list using the matching risk control rules to obtain the calculated risk score; Perform strategy correction processing on the calculated risk score to obtain a repair strategy list after risk assessment.
8. The intelligent data quality monitoring method according to claim 7, characterized in that In step 6, use a multi-dimensional evaluation module to perform refined index calculation and dynamic weight allocation processing on the standardized distribution data stream and the abnormal event list, generate a comprehensive data quality score, and output a visualization report according to the comprehensive data quality score. The specific steps are as follows: Use the multi-dimensional evaluation module to perform index calculations on the standardized distribution data stream and the abnormal event list in sequence for accuracy, integrity, consistency, timeliness, uniqueness, and compliance to obtain a real-time data quality score. The relationship in the corresponding process is as follows: ; Among them, represents the real-time data quality score, represents the accuracy score, represents the integrity score, represents the consistency score, represents the timeliness score, represents the uniqueness score, represents the compliance score; Generate business scenario labels through clustering data and manual presetting, and map the business scenario labels to business priorities in combination with a preset weight template library to obtain an initial weight vector. The relationship in the corresponding process is as follows: ; Among them, represents the initial weight vector, represents the normalization function, represents the preset weight template library, represents the clustering analysis algorithm, represents the clustering feature matrix, represents the union operator, represents the artificially preset business scenario label; Based on the historical data quality score, construct a business loss, and use the gradient descent optimization algorithm to dynamically update the real-time data quality score and the business loss in combination with the initial weight vector to obtain an optimized weight vector. The relationship in the corresponding process is as follows: ; Among them, represents the optimized weight vector, represents the 6-dimensional probability space weight vector of to obtain the optimal parameter value, represents the business loss, represents the historical rating matrix, represents the weight vector transpose of represents the regularization coefficient, represents the regularization term, represents the L2 norm; Perform comprehensive score calculation on the real-time data quality score in combination with the optimized weight vector to generate a comprehensive data quality score; Perform comprehensive score display, multi-dimensional score comparison, and construction of an abnormal event heat map on the comprehensive data quality score, the real-time data quality score, and the abnormal event list in sequence to obtain an interactive dashboard; Perform causal chain diagram construction and visualization rendering processing on the abnormal event list and the abnormal field blood relationship path map in sequence to obtain a root cause analysis report; Use the LSTM model to predict the historical data quality score and the historical business loss to obtain a trend prediction result; Perform format conversion on the interactive dashboard, the root cause analysis report, and the trend prediction result to obtain a visualization report.
9. An intelligent data quality monitoring system, characterized in that, The system applies the intelligent data quality monitoring method according to any one of claims 1 to 8 above. The system includes: A data access layer for: Construct a unified data model based on the Schema mapping rule, and use the unified data model to process multi-source heterogeneous data to obtain a standardized distribution data stream; A first core processing layer for: An adaptive threshold module is constructed based on a sliding window statistical mechanism and a threshold dynamic adjustment algorithm. The adaptive threshold module and a pre-trained model form an anomaly detection engine, and the pre-trained model includes a supervised model and an unsupervised model; Based on algorithms such as random forest, SVM, and neural network, historical data is used as a training set and input into the pre-trained model for training and optimization to obtain an optimized pre-trained model; The optimized pre-trained model is used to perform multi-model fusion detection on the standardized distributed data stream, and combined with the adaptive threshold module for threshold determination processing to obtain a processed list of anomaly events; The intelligent analysis layer is used for: A context awareness and dynamic priority management module is constructed based on multi-source heterogeneous data fusion algorithm, adaptive weight optimization algorithm, and locality sensitive hashing optimization algorithm. The processed list of anomaly events is input into the context awareness and dynamic priority management module for deduplication optimization transmission to obtain a deduplicated and optimized list of anomaly events; The deduplicated and optimized list of anomaly events is subjected to root cause analysis using a root cause analysis tool to obtain a root cause analysis result; A rule library is constructed based on general rules and preset rules. The rule library and a classification model form a repair suggestion engine, and the root cause analysis result is used to assist in generating a list of repair strategies after risk assessment using the repair suggestion engine; The second core processing layer is used for: A multi-dimensional evaluation module is constructed based on a weight allocation algorithm. The multi-dimensional evaluation module performs refined index calculation and dynamic weight allocation processing on the standardized distributed data stream and the list of anomaly events to generate a comprehensive data quality score, and outputs a visualization report according to the comprehensive data quality score; The cross-platform management layer is used for: A unified console is constructed based on a microservices architecture. The unified console and a packaged adapter form a cross-platform monitoring interface. The list of repair strategies after risk assessment and the visualization report are input into the cross-platform monitoring interface to generate a warning notice.
Citation Information
Patent Citations
Intelligent Operation and Maintenance Analysis System Based on Multi-source Heterogeneous Data Fusion, Machine Learning and Customer Service Robot
CN109343995A
Big data task monitoring method, system and device and storage medium
CN117421171A
Cited By
Unified social credit code data quality control method based on traceability technology
CN120975806A
Unified social credit code data quality control method based on traceability technology
CN120975806B
A global campus data quality management and control sharing system and method
CN122596770A