Data management method and system based on machine learning
By constructing a data causal relationship generalization diagram and using deep learning models for data repair, the existing data governance methods are solved in terms of efficiency and flexibility, and efficient and flexible data governance effects are achieved.
Patent Information
- Application Number
- CN202510069929.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
When facing massive and complex data, existing data governance methods are inefficient, poorly flexible, and are easily affected by subjective factors and changes in data sources, resulting in poor governance results.
Using machine learning-based data governance methods, we use machine learning to obtain the initial data to be managed for standardized preprocessing, build a data causal relationship generalization diagram, and use deep learning to build a data simulation model and data reconstruction model, perform point repair and path repair, thereby realizing data governance.
It improves the flexibility and efficiency of data governance, reduces the need for manual intervention, reduces the governance cost, and improves the data feature extraction and repair capabilities by learning subsequences and corrected simulated subsequences.
Smart Images

Figure CN119988847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance technology, and specifically to a data governance method and system based on machine learning. Background Art
[0002] With the rapid development of information technology, various information systems have accumulated an incalculable amount of data in daily operations. In the data-driven era, the quality of information accumulated in various information systems has a crucial impact on the results of data analysis. However, due to various reasons, such as data entry errors, system compatibility issues, human negligence, etc., these data are filled with a large amount of inaccurate or non-standard information. The existence of these non-standard data greatly increases the difficulty of data mining and analysis. It not only consumes more computing resources and leads to resource waste, but also may mislead the results of data analysis, resulting in decision-making errors, thereby bringing immeasurable losses.
[0003] A variety of data governance methods have been proposed in the prior art. Among them, the more traditional method is to manually check and modify data. Although this method can solve the problem to a certain extent, it is time-consuming and labor-intensive, and inefficient. Especially when the amount of data is huge, manual operation can hardly meet the needs of data governance. In addition, manual inspection is easily affected by subjective factors, resulting in uneven quality of data governance.
[0004] Another common data governance method is to write corresponding computer program scripts to automate data processing. This method improves the efficiency of data governance to a certain extent, but it also has problems that cannot be ignored. Due to the variety of types and rules of non-standard data, different programs or scripts often need to be written for different types of data. This not only increases the workload, but also makes the data governance process less flexible. In particular, when the data source changes, the original program or script may not be applicable and needs to be rewritten, which undoubtedly further reduces the efficiency of data governance.
[0005] In summary, the existing data governance methods have certain limitations in terms of efficiency, flexibility and applicability when facing massive and complex data. Therefore, it is particularly urgent to explore a more efficient, flexible and universal data governance method. Summary of the invention
[0006] The purpose of the present invention is to provide a data governance method and system based on machine learning. There is a certain causality between various types of data in the same system, and the principle of data causality can be used to repair the data.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] A data governance method based on machine learning, the steps of which include:
[0009] Acquire initial data to be governed, and perform standardization preprocessing on the initial data to be governed; the standardization preprocessing refers to converting the initial data to be governed into a unified standard, and performing preliminary cleaning on the initial data to be governed;
[0010] Classify the pre-processed data to be governed, determine the causal relationship between each type of data to be governed, and construct a generalized causal relationship diagram of the data;
[0011] Use deep learning to build data simulation models and data reconstruction models based on the data causal relationship generalization diagram;
[0012] Determining the types of data to be managed, wherein the types of data to be managed include point repair and path repair;
[0013] For the point repair, the preprocessed data to be managed is repaired using a data simulation model, and for the path repair, the preprocessed data to be managed is repaired using a data simulation model and a data reconstruction model, thereby performing data management.
[0014] According to the above technical solution, the step of constructing the data causal relationship generalization diagram includes:
[0015] Each type of data to be governed is regarded as a node, and the causal relationship between each node is obtained using the F-GES model to construct a preliminary data causal relationship generalization graph; the preliminary data causal relationship generalization graph is a directed acyclic graph;
[0016] The sequence corresponding to each node except the initial node is regarded as the parent sequence in the grey correlation degree, and the corresponding subsequences are the sequences corresponding to the nodes at the starting positions of all directed edges pointing to the node are regarded as the subsequences of the point. The grey correlation degree is used to calculate the correlation between each parent sequence and the corresponding subsequence, and the directed edges between the nodes corresponding to the corresponding subsequences with correlation less than a threshold α and the nodes corresponding to their parent sequences are eliminated in the preliminary data causal relationship generalization graph to obtain the data causal relationship generalization graph; the data causal relationship generalization graph is a directed acyclic graph.
[0017] Among them, F-GES is a score-based causal detection method. Each time an edge (i.e., causal relationship) is added or reduced, a score will be given to measure the degree of data distribution in the causal graph. The Bayesian information criterion (BIC) is selected for scoring, which can effectively learn the structure of the causal relationship of the time series. However, the method based on the determination of the number of effective components will inevitably lead to the phenomenon of interaction and offset between data, so there is a certain degree of error in the preliminary data causal relationship generalization diagram. Therefore, the gray correlation degree can be used to explore the correlation between node data based on the causal relationship of each node in the preliminary data causal relationship generalization diagram, and the preliminary data causal relationship generalization diagram can be further modified.
[0018] According to the above technical solution, the causal weight of each subsequence corresponding to each parent sequence in the data causal relationship generalization graph is calculated, and the specific steps include:
[0019] Calculate the correlation between each parent sequence and its corresponding child sequence in the data causal relationship generalization diagram using grey correlation;
[0020] The data causal relationship generalization diagram is used as a theoretical model in a structural equation model (SEM), and the path coefficient corresponding to each directed edge is calculated using statistical software;
[0021] Add the path coefficients and corresponding associations of the directed edges between the corresponding nodes of each subsequence and the corresponding nodes of the parent sequence to obtain the preliminary weights;
[0022] Based on the preliminary weight between the parent sequence and each corresponding subsequence, the weight between the parent sequence and each corresponding subsequence is combined as 1 to obtain the causal weight between the parent sequence and each corresponding subsequence.
[0023] In the structural equation model, causal relationships are usually expressed by path coefficients, which reflect the strength and direction of the direct relationship between variables. Therefore, the path coefficient can be used as the causal correlation between nodes, and the grey correlation can be used to explore the correlation between node data. After the causal correlation and the grey correlation are standardized to explore the correlation between node data, the corresponding causal weights between nodes are calculated, making the causal weights more accurate.
[0024] According to the above technical solution, the parent sequence and each of its subsequences multiplied by the corresponding causal weight in the data causal relationship generalization graph are divided into a training set and a validation set, and the data simulation model is constructed by deep learning based on the training set and the validation set. Deep learning can be a convolutional neural network, a recurrent neural network, a long short-term memory network, an autoencoder, and the like.
[0025] According to the above technical solution, the types of data to be managed include:
[0026] Point repair, which is to repair data at a certain point in the data causal relationship generalization graph, which can be understood as repairing incomplete data where a node on a certain edge in the data causal relationship generalization graph is missing;
[0027] Path repair, performing data repair on continuous points in the data causal relationship generalization graph, can be understood as repairing incomplete data with missing continuous nodes in a path in the data causal relationship generalization graph.
[0028] According to the above technical solution, the data reconstruction model execution step includes:
[0029] For a parent sequence missing at a certain point in the data causal relationship generalization graph, a data simulation model is used to obtain a corresponding subsequence and simulate a subsequence, and a virtual causal weight corresponding to the simulated subsequence is calculated;
[0030] Based on the virtual causal weights corresponding to the simulated subsequences and using the weight difference prediction model, the simulated subsequences are predicted with weight differences;
[0031] The predicted weight difference is used to obtain the weight correction amount by using the fuzzy neural network, and the simulation subsequence is corrected according to the weight correction amount; that is, the causal weight corresponding to the point of the simulation subsequence is corrected according to the weight correction amount, and the corrected causal weight is multiplied with the simulation subsequence to obtain the corrected simulation subsequence entering the next convolutional network;
[0032] The subsequence corresponding to the missing parent sequence and the modified simulated subsequence multiplied by the corresponding causal weight are respectively extracted using convolutional layers, and the extracted feature information is input into the global maximum pooling fusion, and then the feature distribution value is calculated after passing through two fully connected layers, and the negative value is filtered out using the ReLU activation function, and the perceptual weights corresponding to the subsequence and the modified simulated subsequence are respectively obtained;
[0033] The subsequence and the modified simulated subsequence are multiplied by the corresponding perceptual weights and input into the recurrent neural network and attention mechanism for processing, and then normalized and output using the softmax function.
[0034] Since the prediction results using the data simulation model are erroneous, the causal weights between nodes can be used to correct the prediction results of the data simulation model using the fuzzy neural network. When the virtual causal weight of the prediction result is too large, the fuzzy neural network is used to obtain the weight correction amount; the virtual causal weight is corrected according to the weight correction amount, and the prediction results of the data simulation model are corrected according to the corrected virtual causal weight, so that the corrected prediction results enter the next layer network for learning.
[0035] Since the number of simulated subsequences corresponding to different parent sequences is inconsistent, the error in the learning process of the neural network may be large. Therefore, the prediction accuracy can be improved by learning the perceptual weights corresponding to the subsequences and the modified simulated subsequences.
[0036] According to the above technical solution, the steps of constructing the weight difference prediction model include:
[0037] For a parent sequence corresponding to a point in the data causal relationship generalization graph, obtaining a causal weight of each subsequence corresponding to the parent sequence;
[0038] Based on all subsequences corresponding to the parent sequence, the data simulation model is used to obtain a simulated subsequence;
[0039] Calculate the virtual causal weights between the parent sequences corresponding to the simulated subsequences;
[0040] Calculate the difference between the virtual causal weight corresponding to the simulated subsequence and the causal weight of the corresponding subsequence;
[0041] A weighted difference model is constructed using a neural network model based on simulated subsequences and their corresponding differences.
[0042] According to the above technical solution, the data simulation model is used to repair the sequence corresponding to the initial point in the path repair, and the sequence corresponding to the data reconstruction model is used to repair the points other than the initial point in the path repair.
[0043] In another embodiment, a data governance system based on machine learning includes:
[0044] A preprocessing module, which obtains initial data to be managed and performs standardized preprocessing on the initial data to be managed;
[0045] The generalization graph construction module classifies the pre-processed data to be governed, and determines the causal relationship between each type of data to be governed to construct a generalization graph of the causal relationship of the data;
[0046] The governance model building module uses deep learning to build data simulation models and data reconstruction models based on the data causal relationship generalization diagram;
[0047] The management data type determination module determines the types of data to be managed, and the types of data to be managed include point repair and path repair; the types of data to be managed include:
[0048] Point repair, performing data repair on a certain point in the data causal relationship generalization graph;
[0049] Path repair, performing data repair on continuous points in the data causal relationship generalization graph;
[0050] The data governance module uses a data simulation model to repair the pre-processed data to be governed for the point repair, uses a data simulation model to repair the sequence corresponding to the initial point in the path repair, and uses a data reconstruction model to repair the points other than the initial point in the path repair, thereby performing data governance.
[0051] Also included is another embodiment, a storage medium for storing computer-executable instructions, characterized in that: when the computer-executable instructions are executed, they implement a data governance method based on machine learning as described in any one of the above-mentioned schemes.
[0052] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: the present invention reveals the internal connections and laws between data by constructing a data causal relationship generalization diagram considering the causality between data, and adopts different repair strategies according to the data causal relationship generalization diagram (point repair and path repair), thereby improving the flexibility of data governance. In the process of path repair, the data entering the next network is corrected by considering the weights between nodes, thereby improving the accuracy of network prediction, and improving the ability to extract data features by learning the perception weights corresponding to the subsequences and the corrected simulated subsequences, thereby improving the ability to repair data. At the same time, the automated and intelligent data governance process reduces the need for manual intervention, thereby reducing the cost of data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0054] Figure 1 It is a flowchart of the steps of a data governance method based on machine learning of the present invention;
[0055] Figure 2 is a flow chart of the calculation steps of the perception weight of the embodiment. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] Example 1
[0058] The present invention provides a technical solution: a data governance method based on machine learning, the steps of which include:
[0059] S1. Acquire initial data to be managed, and perform standardized preprocessing on the initial data to be managed.
[0060] S2, classify the pre-processed data to be managed, determine the causal relationship between each type of data to be managed, and construct a data causal relationship generalization diagram. Calculate the causal weight of each subsequence corresponding to each parent sequence in the data causal relationship generalization diagram.
[0061] The steps for constructing a data causal relationship generalization diagram include:
[0062] Each type of data to be governed is regarded as a node, and the causal relationship between each node is obtained using the F-GES model to construct a preliminary data causal relationship generalization graph; the preliminary data causal relationship generalization graph is a directed acyclic graph;
[0063] Except for the initial node, the sequence corresponding to each node is regarded as the parent sequence in the grey correlation degree, and the corresponding subsequence is the sequence corresponding to the node at the starting position of all directed edges pointing to the node, which are all subsequences of the point. The grey correlation degree is used to calculate the correlation between each parent sequence and the corresponding subsequence, and the directed edges between the nodes corresponding to the corresponding subsequences with correlation less than the threshold α and the nodes corresponding to their parent sequences are eliminated in the preliminary data causal relationship generalization graph to obtain the data causal relationship generalization graph; the data causal relationship generalization graph is a directed acyclic graph.
[0064] The causal weight of each subsequence corresponding to each parent sequence in the data causal relationship generalization graph is calculated, and the specific steps include:
[0065] Calculate the correlation between each parent sequence and its corresponding child sequence in the data causal relationship generalization diagram using grey correlation;
[0066] The data causal relationship generalization diagram is used as a theoretical model in a structural equation model (SEM), and the path coefficient corresponding to each directed edge is calculated using statistical software;
[0067] Add the path coefficients and corresponding associations of the directed edges between the corresponding nodes of each subsequence and the corresponding nodes of the parent sequence to obtain the preliminary weights;
[0068] Based on the preliminary weight between the parent sequence and each corresponding subsequence, the weight between the parent sequence and each corresponding subsequence is combined as 1 to obtain the causal weight between the parent sequence and each corresponding subsequence.
[0069] S3. Use deep learning to build data simulation models and data reconstruction models based on the data causal relationship generalization diagram.
[0070] The parent sequence and each of its subsequences multiplied by the corresponding causal weight in the data causal relationship generalization graph are divided into a training set and a validation set, and the data simulation model is constructed by deep learning based on the training set and the validation set.
[0071] The data reconstruction model execution step includes: obtaining corresponding subsequences and simulated subsequences for a parent sequence missing at a point in the data causal relationship generalization graph, and calculating virtual causal weights corresponding to the simulated subsequences;
[0072] Based on the virtual causal weights corresponding to the simulated subsequences and using the weight difference prediction model, the simulated subsequences are predicted with weight differences;
[0073] The predicted weight difference is used to obtain the weight correction amount by using a fuzzy neural network, and the simulation subsequence is corrected according to the weight correction amount; that is, the causal weight corresponding to this point in the simulation subsequence is corrected according to the weight correction amount, and the corrected causal weight is multiplied by the simulation subsequence to obtain the corrected simulation subsequence that enters the next convolutional network.
[0074] The subsequence corresponding to the missing parent sequence multiplied by the corresponding causal weight and the modified simulated subsequence are respectively extracted with three convolutional layers, and the extracted feature information is input into the global maximum pooling layer for fusion. Then, after two fully connected layers, the simulated subsequence feature distribution value Q1 and the subsequence feature distribution value Q2 are calculated, and the negative values are filtered out using the ReLU activation function. Then, the weight distribution mechanism is used to obtain the subsequence perception weight W2 and the perception weight W1 corresponding to the modified simulated subsequence, as shown in Figure 2 Shown
[0075] in,
[0076] The subsequence and the modified simulated subsequence are multiplied by the corresponding perceptual weights and input into the recurrent neural network and attention mechanism for processing, and then normalized and output using the softmax function.
[0077] Among them, the steps of constructing the weight difference prediction model include:
[0078] For a parent sequence corresponding to a point in the data causal relationship generalization graph, obtaining a causal weight of each subsequence corresponding to the parent sequence;
[0079] Based on all subsequences corresponding to the parent sequence, the data simulation model is used to obtain a simulated subsequence;
[0080] Calculate the virtual causal weights between the parent sequences corresponding to the simulated subsequences;
[0081] Calculate the difference between the virtual causal weight corresponding to the simulated subsequence and the causal weight of the corresponding subsequence;
[0082] A weighted difference model is constructed using a neural network model based on simulated subsequences and their corresponding differences.
[0083] S4. Determine the type of data to be governed, the type of data to be governed includes point repair and path repair; point repair is to perform data repair on a certain point in the data causal relationship generalization graph; path repair is to perform data repair on continuous points in the data causal relationship generalization graph.
[0084] S5. For the point repair, the pre-processed data to be managed is repaired using the data simulation model. For the path repair, the pre-processed data to be managed is repaired using the data simulation model and the data reconstruction model, thereby performing data management. Specifically, for the initial point in the path repair, the sequence corresponding to the initial point is repaired using the data simulation model. For the points other than the initial point in the path repair, the sequence corresponding to the data reconstruction model is repaired.
[0085] Example 2
[0086] A data governance system based on machine learning, comprising:
[0087] A preprocessing module, which obtains initial data to be managed and performs standardized preprocessing on the initial data to be managed;
[0088] The generalization graph construction module classifies the pre-processed data to be governed, and determines the causal relationship between each type of data to be governed to construct a generalization graph of the causal relationship of the data;
[0089] The governance model building module uses deep learning to build data simulation models and data reconstruction models based on the data causal relationship generalization diagram;
[0090] The management data type determination module determines the types of data to be managed, and the types of data to be managed include point repair and path repair; the types of data to be managed include:
[0091] Point repair, performing data repair on a certain point in the data causal relationship generalization graph;
[0092] Path repair, performing data repair on continuous points in the data causal relationship generalization graph;
[0093] The data governance module uses a data simulation model to repair the pre-processed data to be governed for the point repair, uses a data simulation model to repair the sequence corresponding to the initial point in the path repair, and uses a data reconstruction model to repair the points other than the initial point in the path repair, thereby performing data governance.
[0094] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0095] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data governance method based on machine learning, characterized in that: The steps include: Acquire initial data to be managed, and perform standardized preprocessing on the initial data to be managed; Classify the pre-processed data to be governed, determine the causal relationship between each type of data to be governed, and construct a generalized causal relationship diagram of the data; Use deep learning to build data simulation models and data reconstruction models based on the data causal relationship generalization diagram; Determining the types of data to be managed, wherein the types of data to be managed include point repair and path repair; For the point repair, the preprocessed data to be managed is repaired using a data simulation model, and for the path repair, the preprocessed data to be managed is repaired using a data simulation model and a data reconstruction model, thereby performing data management.
2. A data governance method based on machine learning according to claim 1, characterized in that: The steps of constructing the data causal relationship generalization diagram include: Each type of data to be governed is regarded as a node, and the causal relationship between each node is obtained using the F-GES model to construct a preliminary data causal relationship generalization graph; the preliminary data causal relationship generalization graph is a directed acyclic graph; Except for the initial node, the sequence corresponding to each node is regarded as the parent sequence in the grey correlation degree, and the corresponding subsequence is the sequence corresponding to the node at the starting position of all directed edges pointing to the node. The grey correlation degree is used to calculate the correlation between each parent sequence and the corresponding subsequence, and the directed edges between the nodes corresponding to the corresponding subsequences with correlation less than the threshold α and the nodes corresponding to the parent sequence are eliminated in the preliminary data causal relationship generalization diagram to obtain the data causal relationship generalization diagram; The data causal relationship generalization graph is a directed acyclic graph.
3. A data governance method based on machine learning according to claim 2, characterized in that: Calculating the causal weight of each subsequence corresponding to each parent sequence in the data causal relationship generalization graph, the specific steps include: Calculate the correlation between each parent sequence and its corresponding child sequence in the data causal relationship generalization diagram using grey correlation; The data causal relationship generalization diagram is used as a theoretical model in a structural equation model (SEM), and the path coefficient corresponding to each directed edge is calculated using statistical software; Add the path coefficients and corresponding associations of the directed edges between the corresponding nodes of each subsequence and the corresponding nodes of the parent sequence to obtain the preliminary weights; Based on the preliminary weight between the parent sequence and each corresponding subsequence, the weight between the parent sequence and each corresponding subsequence is combined as 1 to obtain the causal weight between the parent sequence and each corresponding subsequence.
4. A data governance method based on machine learning according to claim 3, characterized in that: The parent sequence and each of its subsequences multiplied by the corresponding causal weight in the data causal relationship generalization graph are divided into a training set and a validation set, and the data simulation model is constructed by deep learning based on the training set and the validation set.
5. The data governance method based on machine learning according to claim 1 is characterized in that: The types of data to be managed include: Point repair, performing data repair on a certain point in the data causal relationship generalization graph; Path repair, performing data repair on continuous points in the data causal relationship generalization graph.
6. The data governance method based on machine learning according to claim 1 is characterized in that: The data reconstruction model execution step includes: For a parent sequence missing at a certain point in the data causal relationship generalization graph, a data simulation model is used to obtain a corresponding subsequence and simulate a subsequence, and a virtual causal weight corresponding to the simulated subsequence is calculated; Based on the virtual causal weights corresponding to the simulated subsequences and using the weight difference prediction model, the simulated subsequences are predicted with weight differences; The predicted weight difference is used to obtain the weight correction amount by using the fuzzy neural network, and the simulation subsequence is corrected according to the weight correction amount; The subsequence corresponding to the missing parent sequence and the modified simulated subsequence multiplied by the corresponding causal weight are respectively extracted using convolutional layers, and the extracted feature information is input into the global maximum pooling fusion, and then the feature distribution value is calculated after passing through two fully connected layers, and the negative value is filtered out using the ReLU activation function, and the perceptual weights corresponding to the subsequence and the modified simulated subsequence are respectively obtained; The subsequence and the modified simulated subsequence are multiplied by the corresponding perceptual weights and input into the recurrent neural network and attention mechanism for processing, and then normalized and output using the softmax function.
7. The data governance method and system based on machine learning according to claim 6, characterized in that: The steps of constructing the weight difference prediction model include: For a parent sequence corresponding to a point in the data causal relationship generalization graph, obtaining a causal weight of each subsequence corresponding to the parent sequence; Based on all subsequences corresponding to the parent sequence, the data simulation model is used to obtain a simulated subsequence; Calculate the virtual causal weights between the parent sequences corresponding to the simulated subsequences; Calculate the difference between the virtual causal weight corresponding to the simulated subsequence and the causal weight of the corresponding subsequence; A weighted difference model is constructed using a neural network model based on simulated subsequences and their corresponding differences.
8. The data governance method based on machine learning according to claim 1, characterized in that: The data simulation model is used to repair the sequence corresponding to the initial point in the path repair, and the sequence corresponding to the data reconstruction model is used to repair the points other than the initial point in the path repair.
9. A data governance system based on machine learning, characterized in that: include: A preprocessing module, which obtains initial data to be managed and performs standardized preprocessing on the initial data to be managed; The generalization graph construction module classifies the pre-processed data to be governed, and determines the causal relationship between each type of data to be governed to construct a generalization graph of the causal relationship of the data; The governance model building module uses deep learning to build data simulation models and data reconstruction models based on the data causal relationship generalization diagram; A management data type determination module determines the type of data to be managed, wherein the type of data to be managed includes point repair and path repair; The types of data to be managed include: Point repair, performing data repair on a certain point in the data causal relationship generalization graph; Path repair, performing data repair on continuous points in the data causal relationship generalization graph; The data governance module uses a data simulation model to repair the pre-processed data to be governed for the point repair, uses a data simulation model to repair the sequence corresponding to the initial point in the path repair, and uses a data reconstruction model to repair the points other than the initial point in the path repair, thereby performing data governance.
10. A storage medium for storing computer executable instructions, characterized in that: When executed, the computer executable instructions implement a data governance method based on machine learning as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Low power consumption wireless personal area network tree-like routing method based on IPv6
CN107197480A
Large-scale equipment fault prediction method with causal and attention emphasized in industrial internet
CN114580472A
Power grid big data security detection early warning method and system based on causal machine learning
CN115526407A
Power utilization data restoration method and system based on Bayesian Gaussian tensor decomposition model
CN116610911A
Model correction method and device, equipment and storage medium
CN116757293A