An integrated weighted master data recognition method based on machine learning
Through the comprehensive weighted master data recognition method based on machine learning, the random forest model and decision tree are used for classification, and the comprehensive weighting method is used to improve the accuracy of classification results, the problem of inaccuracy and efficiency in master data recognition is solved, and more efficient master data management is achieved.
Patent Information
- Application Number
- CN202111201242.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-15
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-15
AI Technical Summary
The prior art has problems of accuracy and efficiency in master data identification, especially when dealing with complex and diverse business entities, it is difficult to ensure the accuracy and efficiency of the identification.
The comprehensive weighted master data recognition method based on machine learning is adopted. By sorting out the business entity domain, selecting representative features of the master data, building a random forest model, using the decision tree for classification, and using the comprehensive weighting method to assign different weights to multiple decision trees, and finally obtaining the final classification result through the voting rule.
It improves the accuracy and speed of searching for enterprise business master data, effectively improves the management efficiency of enterprise master data, reduces the impact of manual participation and unreasonable scoring on the identification results, and improves the accuracy of predicted results and the stability of the model.
Smart Images

Figure CN113920366B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology and relates to a comprehensive weighted master data recognition method based on machine learning. Background Art
[0002] Master data is the basic information of an enterprise (organization) that can meet the cross-departmental collaboration needs of an enterprise and reflect the status attributes of core business entities, with relatively stable attributes, higher accuracy requirements, and uniquely identifiable data. Enterprise master data is usually scattered in various systems, covering all aspects of business processes such as human resources and finance, and there are numerous associated business entities. How to accurately identify enterprise master data from complex and numerous business entities is a complex project.
[0003] Currently, some literature has studied methods for master data recognition. For example, the Analytic Hierarchy Process (AHP) is a multi-objective and multi-criteria decision-making method, which has adaptability, simplicity, practicability, and systematicness. However, when there are too many evaluation indicators in a single layer, the scale of the judgment matrix will be too large, and it is difficult to check whether the judgment matrix is consistent when the matrix scale is too large. When there are consistency problems in the judgment matrix that need to be adjusted, the adjustment workload is relatively large and complex, and it is difficult to ensure the consistency between the judgment matrix and human thinking. The currently commonly used method is to evaluate through methods such as expert scoring, using qualitative characteristics such as data independence, sharing, and life cycle as evaluation indicators, and then determining master data items.
[0004] Although the method is simple and fast, it has high requirements for experts. If they are not familiar with the uses and status of each data item in all application systems, there will be misjudgments or omissions, and it is difficult to ensure the accuracy of master data recognition.
[0005] In summary, master data recognition is still a major challenge that enterprises urgently need to solve, and it is very necessary to explore an effective theory and method. Summary of the Invention
[0006] To solve the above problems, the purpose of the present invention is to provide a comprehensive weighted master data recognition method based on machine learning, which can improve the accuracy and speed of enterprise business master data search and effectively improve the management efficiency of enterprise master data. Select N master data tables according to the sorted and divided N business entity domains to construct a sample set, clean and filter out empty tables and tables with extremely small amounts of data, obtain the characteristics of the master data tables based on the sample set, randomly select 70% of the tables in the sample set as the training set, and 30% as the validation set for random forest training to construct decision trees, and use the constructed decision trees to predict the test set. Use the comprehensive weighting method to assign different weights to multiple decision trees, and use the voting rule to obtain the final classification result.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A comprehensive weighted master data recognition method based on machine learning, comprising the following steps:
[0009] Step 1: Sort out the business entity domain and select the most representative features of the master data according to the characteristics of the master data. The features at least include the out-degree and in-degree of references;
[0010] Step 2: Use the recognition features extracted from the master data obtained in Step 1 as the features for random forest classification. Select a training set, perform data cleaning, and based on the random forest algorithm, select the optimal parameters to construct a decision tree;
[0011] Step 3: Use the test set and test it with the constructed multiple decision trees to obtain the corresponding classification categories;
[0012] Step 4: Use the comprehensive weighting method to assign different weights to the decision trees, and use the voting rule to obtain the final classification result.
[0013] In a preferred embodiment of the present invention, the master data recognition features include table information features and data features. The table features include but are not limited to table name, creation time, table comment, table data volume, out-degree and in-degree of references; the data features include but are not limited to: field name, field type, field comment, number of field value records, distinct records of field values, primary key information.
[0014] In a preferred embodiment of the present invention, in Step 2, the CART algorithm is selected to divide the data set for the internal nodes of the decision tree.
[0015] In a preferred embodiment of the present invention, the number of decision trees is set to 100.
[0016] In a preferred embodiment of the present invention, a method of weighting the Gini coefficient is used to construct a decision tree: The Boostrap machine sampling method is used to extract a data set D from the master data. The data set D consists of x training samples and M features, and the weight W of each class of samples k is inversely proportional to the frequency P k (k = 1, 2,..., K) of the classification appearing in the data set D, where K is the number of sample categories, then:
[0017]
[0018] During the growth process of the decision tree, the weighted Gini coefficient G W is used to find the optimal splitting feature:
[0019]
[0020] Among them, n k is the number of various samples in the node; W k is the weight value assigned to each category.
[0021] In a preferred embodiment of the present invention, the probability that a point in the dataset D belongs to the k-th category is P k , then the Gini index of this probability distribution is:
[0022]
[0023] In a preferred embodiment of the present invention, the dataset D can be divided into two parts D1 and D2 according to the feature A, and G W (D, A) is minimized to obtain the optimal partition, and a weighted decision tree is constructed:
[0024]
[0025] G W (D, A) The minimum value is the optimal feature of this node, and A (A ∈ M) is the splitting feature.
[0026] In a preferred embodiment of the present invention, weight values are assigned to each decision tree:
[0027]
[0028] TP represents the number of samples that the model predicts to be true and are actually true, where true is the main data, FP represents the number of samples that the model predicts to be true but are actually false, and FN represents the number of samples that the model predicts to be false but are actually true; then the weight assignment formula is: Among them, F i is the F1-Score value of the i-th decision tree.
[0029] The beneficial effects of the present invention are:
[0030] 1. Usually, we set the main data recognition weight according to the enterprise situation requirements and expert opinions, construct a main data recognition scoring template, and then obtain the score of the suspected main data. The random forest prediction of the main data weakens the design of the scoring template in the traditional main data recognition, thus reducing the influence of manual participation or unreasonable scoring on the recognition result.
[0031] 2. Traditionally, the random forest algorithm votes by counting the categories output in the decision tree. When the number of votes for the two categories of samples is very close, it is predicted as the category with more votes. This method may lead to unsatisfactory results. The comprehensive weighting method will avoid the above situation and improve the accuracy of the result and the stability of the model in the final prediction result classification voting. Description of the Drawings
[0032] Figure 1 This is a diagram of a comprehensive weighted master data recognition method based on machine learning in this application;
[0033] Figure 2 This is a diagram of the characteristics of the master data table. Detailed implementation manners
[0034] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.
[0035] Master data is the basic data with sharing nature, which can be reused across various business departments within an enterprise and is in a state of high value, high sharing and relatively stable, such as Figure 2 .
[0036] High sharing is a particularly important feature of master data. According to the distribution of the company's business departments, the sorted systems include human resources system, project management system, customer management system, product management system, financial system and logistics management system. Then, the employee table, department table and customer table are very likely to be master data tables. Therefore, the training set data table for master data table recognition can be sorted out.
[0037] According to the above characteristics and table information, it is determined that the master data recognition features include table information features and data features. The table features are table name, creation time, table comment, table data volume, out-degree and in-degree of reference; the data features are: field name, field type, field comment, number of field value records, distinct records of field values, primary key information. The above information is extracted from the relational database as the features for random forest classification.
[0038] Since there are empty tables or tables with extremely small data volume, such as tables with only one row of data volume, in the original database of the enterprise, such tables need to be screened out from the data tables during the data cleaning process. The training set data required for the final algorithm is obtained.
[0039] Due to its characteristics such as simple and flexible, not easy to overfit, and high accuracy, the random forest algorithm has shown good results in many applications. It consists of multiple decision trees, and the CART algorithm is selected to divide the internal nodes of the decision trees. The measurement index of the CART decision tree algorithm is the Gini coefficient. We divide the data set into three categories: high, medium and low. Then the probability that a point in the data set D belongs to the k-th category is P k , then the Gini index of this probability distribution is:
[0040]
[0041] To find the optimal number of decision trees, the number of trees was set to 10, 50, 100, 150, 300, and 500 respectively, and the dataset was classified and verified. It was found that the classification accuracy increased with the increase in the number of decision trees. When the number of trees reached 100, the classification accuracy tended to be stable, and the error remained within a certain accuracy range. When the number of trees was large enough, the influence of feature variables on the classification accuracy was controlled within a certain range. Therefore, the optimal number was determined to be 100.
[0042] In the standard random forest algorithm, features are randomly selected, so the probability of each sample feature being selected is the same. However, in fact, the importance of each feature's impact on the result is different.
[0043] The present invention constructs a decision tree by using a method of weighting the Gini coefficient: The Boostrap machine sampling method is used to extract the dataset D from the original data. The dataset D consists of x training samples and M (where M is 12) features. The weight W of each class of samples k is inversely proportional to the frequency P k (k = 1, 2,..., K) of the classification in the sample set, and K is the number of sample categories. Then:
[0044]
[0045] During the growth process of the decision tree, the weighted Gini coefficient G W is used to find the optimal splitting feature:
[0046]
[0047] where n k is the number of various samples in the node; W k is the weight value assigned to each class. Assuming that the dataset D can be split into D1 and D2 according to the feature A, the minimum value of G W (D, A) is obtained to get the optimal splitting and construct the weighted decision tree:
[0048]
[0049] G W (D, A) minimum value is the optimal feature of this node, and A (A ∈ M) is the splitting feature.
[0050] After determining the feature parameters, we train on the training set based on the weighted random forest algorithm. After obtaining the forest, when judging or predicting a new sample, each decision tree in the forest judges the sample separately, and finally obtains multiple different classification results.
[0051] Due to its own randomness, the prediction of the random forest algorithm may lead to certain fluctuations in the results of each test, and this instability sometimes also affects the overall prediction performance. Therefore, different weights should be assigned to the corresponding decision trees to improve the classification effect.
[0052] Precision and recall are two important indicators for evaluating the model effect. However, sometimes there will be contradictory situations, so it is necessary to consider them comprehensively. The most common method is to calculate the F-Score, which is usually used to evaluate the quality of model classification.
[0053] Use the method of comprehensively weighting precision and recall, that is, the F1-Score value, to assign weights to each decision tree:
[0054]
[0055] Among them, "true" means a table is the main data table, and "false" means a table is a non-main data table. Then, TP represents the number of samples that the model predicts as true (main data) and are actually true, FP represents the number of samples that the model predicts as true but are actually false, and FN represents the number of samples that the model predicts as false but are actually true.
[0056] Then the weight assignment formula is: Among them, F i is the F1-Score value of the i-th decision tree.
[0057] Input the data of the validation set into each decision tree. Then each decision tree will output a class prediction for each record in the validation set, and compare the prediction result of the decision tree with the true result. If the effect of the model is not good, the parameters in the feature extraction and random forest model mentioned above need to be adjusted.
[0058] The improved random forest algorithm reduces the influence brought by the average voting mechanism, reduces the influence of weak classifiers on the results, and improves the overall performance of the algorithm. Finally, input the test set into the algorithm model to obtain the classification result and save the result to the database.
Claims
1. An integrated weighted master data recognition method based on machine learning, characterized in that, It includes the following steps: Step 1: Comb the business entity domain and select the most representative identification features of the master data according to the characteristics of the master data. The features include at least the out-degree and in-degree of references; Step 2: Use the identification features extracted from the master data obtained in Step 1 as the features for random forest classification. Select a training set, perform data cleaning, and based on the random forest algorithm, select the optimal parameters to construct decision trees; Step 3: Use the test set and test it with the constructed multiple decision trees to obtain the corresponding classification categories; Step 4: Use the comprehensive weighting method to assign different weights to the decision trees, and use the voting rule to obtain the final classification result. Assign weights to each decision tree: TP represents the number of samples that the model predicts as true and are actually true, where true is the main data, FP represents the number of samples that the model predicts as true but are actually false, and FN represents the number of samples that the model predicts as false but are actually true; then the weight assignment formula is: Among them, F i is the F1-Score value of the i-th decision tree.
2. The integrated weighted master data identification method based on machine learning according to claim 1, wherein The master data identification features include table information features and data features. The table information features include but are not limited to table name, creation time, table comment, table data volume, out-degree and in-degree of references; the data features include but are not limited to: field name, field type, field comment, number of field value records, distinct records of field values, primary key information.
3. A comprehensive weighted master data identification method based on machine learning according to claim 1, characterized in that In Step 2, the CART algorithm is selected to divide the data set for the internal nodes of the decision tree.
4. A comprehensive weighted master data identification method based on machine learning according to claim 1, characterized in that, The number of decision trees is set to 100.
5. A comprehensive weighted master data recognition method based on machine learning according to claim 1, characterized in that Construct a decision tree by using the method of weighting the Gini coefficient: Use the Boostrap machine sampling method to extract the data set D from the main data. The data set D consists of x training samples and M features, and the weight W of each class of samples k is inversely proportional to the frequency P k (k = 1, 2,..., K) at which this classification appears in the data set D, where K is the number of sample categories, then: During the growth process of the decision tree, the weighted Gini coefficient G is used W to find the optimal splitting feature: Among them, n k is the number of various samples in the node; W k is the weight value assigned to each category.
6. The integrated weighted master data recognition method based on machine learning according to claim 5, characterized in that The probability that a point in the dataset D belongs to the k-th class is P k , then the Gini index of the probability distribution is:
7. A comprehensive weighted master data recognition method based on machine learning according to claim 5, characterized in that, The dataset D can be divided into two parts, D1 and D2, according to the feature A, and find G W (D, A) to obtain the optimal partition by minimizing the value, and construct a weighted decision tree: G W (D, A) The minimum value is the optimal feature of the node, and A (A ∈ M) is the splitting feature.
Citation Information
Patent Citations
Optimized classification method and optimized classification device based on random forest algorithm
CN105844300A
Double-accuracy weighted random forest algorithm based on particle swarm optimization
CN111428790A