Big data automatic classification and grading method, system and device based on dynamic update feedback mechanism and medium
By extracting field feature vectors in the big data platform and building a multi-dimensional scoring and grading model, combining data distribution detection and model incremental training, the problem of dynamic changes in traditional methods is solved, and the automation and real-time update of big data classification and grading is realized, which improves the efficiency and security of data management.
Patent Information
- Application Number
- CN202510358768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology is difficult to deal with the dynamic changes in data characteristics and business scenarios in terms of big data classification and grading. Traditional methods rely on manual rules or expert experience, and the automation methods lack real-time feedback mechanism, resulting in lagging model updates and difficult to deploy in resource-constrained environments.
By extracting the feature vectors of fields in different tables and scenarios, a multi-dimensional scoring and grading model is constructed, combining data distribution detection and model incremental training, dynamic updates and closed-loop feedback are realized, and full and incremental training is used for AutoGluon framework, lightweight incremental models are generated and shadow mode deployment is performed.
The automation, dynamic update and efficient management of data classification and hierarchy are realized, and the dependence on manual rules is reduced, ensuring that the model remains efficient and stable in the big data platform, adapts to new data characteristics, and meets the requirements of the actual production environment.
Smart Images

Figure CN120296501A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data classification and grading, and particularly relates to a big data automatic classification and grading method, system, device and medium based on a dynamic update feedback mechanism. Background Art
[0002] In recent years, the rapid development and wide application of big data technology have greatly promoted the growth of data storage and processing requirements. In the fields of finance, healthcare, government management, and e-commerce, the collection, storage, and analysis of a large amount of data have become the core. However, data management in a big data environment not only faces the challenge of scale expansion but also has the following remarkable characteristics: diverse types, dynamic changes, and high sensitivity. These characteristics further exacerbate the complexity of data classification and grading and make traditional data management strategies difficult to meet current requirements.
[0003] In actual operation, data classification and grading have been widely defined as the basic link of the data security system by national and industry standards. By accurately classifying and grading data, highly sensitive data (such as customer privacy and financial data) can be effectively identified and protected, data with lower sensitivity but higher importance in resource management (such as business logs and operation records) can be reasonably allocated, and a loose management strategy can be adopted for data with low sensitivity and low importance (such as public statistical data). The results of data classification and grading directly affect the setting of data access permissions, the selection of desensitization strategies, and the effective implementation of the overall data security strategy.
[0004] Currently, in terms of data classification and grading, the following several technical solutions are widely applied.
[0005] Classification and grading method based on manual rules: Classify and grade data by manually defining rules (such as keyword matching and regular expressions); however, once the rules are formulated, they are fixed and cannot cope with the dynamic changes of data characteristics and business scenarios; a large number of experts are required to manually maintain and update the rules.
[0006] Classification and grading method based on feature extraction: Use the statistical features of data (such as distribution, entropy value, and frequency) for classification and grading; however, simply relying on static feature extraction may be difficult to comprehensively express data distribution and semantic information and is easily limited by prior knowledge; it is insensitive to small changes in data distribution and statistical characteristics and cannot achieve dynamic adjustment.
[0007] Classification and grading methods based on traditional machine learning: By training traditional machine learning models (such as random forest, support vector machine, logistic regression, etc.), using data features (such as sensitivity, importance, access frequency, etc.) to classify and grade data. However, the feature selection, parameter tuning, and training process of the model rely on expert experience, resulting in high development costs; traditional models are often trained in one-time full quantity, making it difficult to support incremental updates, and the retraining cost is high when facing big data.
[0008] Methods based on automated machine learning (AutoML): Using automated machine learning tools (such as Google AutoML, AutoGluon), through automatically completing data preprocessing, feature selection, model training, and parameter optimization. However, in the face of constantly changing new data, traditional AutoML methods may lack a real-time feedback mechanism, resulting in model update lags; the models generated automatically are often more complex and difficult to deploy in resource-constrained environments. Summary of the Invention
[0009] In order to overcome the above-mentioned disadvantages in the prior art, the purpose of the present invention is to provide a big data automated classification and grading method, system, device, and medium based on a dynamic update feedback mechanism. By extracting the feature vectors of fields in different tables and different scenarios, and according to the scoring results in three dimensions of sensitivity, importance, and usage frequency, determine the grading results of the data in the fields, and construct a multi-dimensional scoring and grading model; perform incremental training of the model through data distribution detection to achieve dynamic update of the model; through grading anomaly detection, perform retraining of the model to achieve closed-loop feedback of the model. The classification and grading method of the present invention has the characteristics of full-process automation of classification, dynamic update, closed-loop feedback, and adaptation to big data platforms, significantly improving the efficiency, security, and flexibility of data classification and grading, and providing a comprehensive, accurate, and efficient data management solution for enterprises and organizations.
[0010] In order to achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0011] A big data automated classification and grading method based on a dynamic update feedback mechanism, comprising the following steps:
[0012] Step 1: Analyze the table structure information and data itself features in Hive metadata, extract the feature vectors of fields in different tables and different business scenarios, and obtain the vector analysis result of the fields;
[0013] Step 2: Score the fields in Hive metadata respectively according to the grading criteria of the corresponding dimensions from three dimensions of sensitivity, importance, and usage frequency, set fusion weights for weighted combination, and then generate the initial grading result of the data in the fields according to the comprehensive rating threshold;
[0014] Step 3: Use the data itself, the vector parsing results of the fields, and the scoring results of each dimension as training samples for full-scale training. Respectively construct the initial scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform;
[0015] Step 4: Collect and calculate the increment of the new data in the fields. When the increment of the new data volume in the fields exceeds the increment threshold of the original data volume in the fields, perform data distribution detection; Determine whether to perform model incremental update according to the data distribution detection result: If so, use the new data itself, the vector parsing results of the fields, and the scoring results of each dimension as new training samples for incremental training to obtain the incremental scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform; If not, update the original data statistics in the fields;
[0016] Step 5: Simultaneously perform hierarchical anomaly detection on the new data in the collected fields. If the detection result feedback is a low-risk warning or a high-risk warning, re-score all the data in the field and re-train the model; If the hierarchical level of the new data in the field increases after re-scoring and feedbacks a high-risk warning, monitor the access frequency of the historical access logs of the field through a weighted directed graph. If an abnormal warning appears in the access frequency, re-score all the data in the field and re-train the model again; If the access frequency is normal, no processing is performed; Then load the scored and graded model after re-training into the multi-dimensional scoring and grading model of the big data platform to replace the original scored and graded model.
[0017] The feature vectors extracted in Step 1 include field metadata and field distribution parameters; The field metadata includes column names, column types, and table names; The field distribution parameters include the minimum value, maximum value, mean value, and standard deviation of the numerical columns;
[0018] In Step 2, the sensitivity, importance, and usage frequency are all divided into five levels, namely levels 1-5. The higher the level, the greater the impact of this dimension;
[0019] The full-scale training in Step 3, the incremental training in Step 4, and the re-training in Step 5 are all carried out using the AutoGluon framework.
[0020] In Step 4, the KL divergence is used for data distribution detection. The calculation formula of the KL divergence is as follows:
[0021]
[0022] where p(x) is the probability distribution of the original data in the field on the feature space X, and q(x) is the probability distribution of the new data in the field on the same feature space X;
[0023] When the KL value ≤ 0.1, update the original data statistics in the field; when the KL value > 0.1, perform model incremental update.
[0024] Perform model distillation on the incremental scoring and grading model for confidentiality, importance, and usage frequency described in step 4 to obtain a lightweight incremental scoring and grading model; deploy the lightweight incremental scoring and grading model on the big data platform in shadow mode. If the average inference time of the lightweight incremental scoring and grading model is less than or equal to the average inference time of the initial scoring and grading model and the memory and CPU usage increase by no more than 5%, it is considered to meet the online performance requirements, then use the lightweight incremental scoring and grading model to replace the initial scoring and grading model for model update; otherwise, if it does not meet the online performance requirements, no model update is performed.
[0025] The process of hierarchical anomaly detection in step 5 is as follows:
[0026] Input the vector parsing results of the new data itself and the field into the initial scoring and grading model for confidentiality, importance, and usage frequency to score the new data, and obtain the corresponding scoring prediction results for confidentiality, importance, and usage frequency; use the fusion weights set in step 2 for weighted combination to determine the predicted grading result of the new data; according to the initial grading result in step 2, perform hierarchical anomaly detection on the predicted grading result. If the amount of data for which the predicted grading result of the new data is higher than the initial grading result of the original data exceeds the high-risk threshold of the original data volume in the field, a high-risk warning is fed back; if the amount of data for which the predicted grading result of the new data is lower than the initial grading result of the original data exceeds the low-risk threshold of the original data volume in the field, a low-risk warning is fed back.
[0027] Each node in the weighted directed graph in step 5 represents a field, and the edge represents the access correlation between fields; calculate the importance degree of each node through PageRank iteration. The calculation formula of PageRank is as follows:
[0028]
[0029] In the formula, PR(v) represents the PageRank value of node v, ln(u) represents the set of nodes pointing to node v, Out(u) represents the number of out-edges of node u, and α represents the damping coefficient;
[0030] When there is an abnormal fluctuation of the PageRank weight exceeding the threshold in the weighted directed graph, trigger an access frequency anomaly warning.
[0031] When new data tables or fields appear in the big data platform, the current scoring and grading model on the big data platform is used to perform multi-dimensional scoring on the current data. The most frequently occurring scoring result in each dimension is taken as the final scoring result for that dimension. Weighted merging is performed according to the fusion weights set in step 2 to determine the grading result of the data in this field, and subsequent grading processing is the same as that of the original fields.
[0032] The present invention also provides a big data automated classification and grading system based on a dynamic update feedback mechanism, including:
[0033] Initial scoring and grading model training module: used to parse the table structure information and data itself characteristics in Hive metadata, extract the feature vectors of fields in different tables and different business scenarios, and obtain the vector parsing results of fields; score the fields in Hive metadata from three dimensions of sensitivity, importance, and usage frequency according to the grading standards of the corresponding dimensions, set fusion weights for weighted merging, and then generate the initial grading results of the data in the fields according to the comprehensive rating threshold; use the data itself, the vector parsing results of the fields, and the scoring results of each dimension as training samples for full-scale training, respectively construct the initial scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform;
[0034] Grading model dynamic update module: used to collect and calculate the increment of new data in the fields. When the amount of new data in the fields exceeds the increment threshold of the original data volume in the fields, data distribution detection is performed; determine whether to perform model incremental update according to the data distribution detection result: if so, use the new data itself, the vector parsing results of the fields, and the scoring results of each dimension as new training samples for incremental training, obtain the incremental scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform; if not, update the original data statistics in the fields;
[0035] Grading model closed-loop feedback module: used to simultaneously perform grading anomaly detection on the new data collected in the fields; if the detection result feedback is a low-risk warning or a high-risk warning, re-score all the data in this field and perform model re-training; if the grading level of the new data in this field increases after re-scoring and feedback is a high-risk warning, monitor the access frequency of the historical access logs of this field through a weighted directed graph. If an abnormal warning occurs in the access frequency, re-score all the data in this field and perform model re-training again; if the access frequency is normal, no processing is performed; then load the re-trained scoring and grading model into the multi-dimensional scoring and grading model of the big data platform to replace the original scoring and grading model.
[0036] The present invention also provides a big data automated classification and grading device based on a dynamic update feedback mechanism, including:
[0037] A memory: used to store the computer program of the above-mentioned big data automated classification and grading method based on a dynamic update feedback mechanism, and is a computer-readable device;
[0038] A processor: used to implement the above-mentioned big data automated classification and grading method when executing the computer program.
[0039] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned big data automated classification and grading method based on a dynamic update feedback mechanism.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] 1. The present invention automatically captures the semantic differences of "the same word with different meanings" by extracting the feature vectors of fields in different scenarios; through data distribution detection and model incremental update, it reduces the dependence on manual rules and realizes automated and adaptive update.
[0042] 2. When the present invention extracts field metadata, it also combines the distribution parameters of numerical columns (such as minimum value, maximum value, mean value, standard deviation), and uses KL divergence to detect the data distribution. Once a significant change is found, it triggers model incremental update to realize real-time dynamic adaptation to new data features.
[0043] 3. The present invention uses the AutoGluon framework to combine full-scale training and incremental training, uses an incremental collector to collect new data in real time, and incrementally updates the model according to the data distribution detection results to ensure that the model always maintains high efficiency and stability in the big data platform.
[0044] 4. Compared with traditional single models, while the present invention uses AutoGluon to realize automated modeling, it uses model distillation technology to generate a lightweight incremental model with a smaller volume and lower computational overhead; and through shadow mode deployment, the old and new models run in parallel and compare and verify the performance and resource consumption of the new model to ensure that the new model has advantages in inference speed and stability and meets the requirements of the actual production environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a module structure diagram of the big data automated classification and grading method based on a dynamic update feedback mechanism of the present invention.
[0046] Figure 2Flowchart of the big data automated classification and grading method based on the dynamic update feedback mechanism of the present invention.
[0047] Figure 3 Flowchart of the grading of the new fields provided by the present invention. Detailed implementation manners
[0048] The technical solution of the present invention will be further introduced below in conjunction with the accompanying drawings.
[0049] As Figure 1 and Figure 2 shown, a big data automated classification and grading method based on the dynamic update feedback mechanism specifically includes the following steps:
[0050] Step 1: Analyze the table structure information and data itself characteristics in the Hive metadata, use the Column2Vec tool to extract the feature vectors of the fields in different tables and different business scenarios, obtain the vector parsing results of the fields, and realize the differential recognition of "same word with different meanings" or "different semantics of the same field name in different tables";
[0051] The extracted feature vectors include field metadata and field distribution parameters; the field metadata includes column name, column type, and table name; the field distribution parameters include the minimum value, maximum value, mean value, and standard deviation of the numerical column;
[0052] Step 2: The system administrator scores the fields in the Hive metadata according to industry or national regulations from three dimensions of sensitivity, importance, and usage frequency according to the grading standards of the corresponding dimensions, and obtains the initial scoring results of the corresponding dimensions; set the fusion weights to perform weighted merging on the initial scoring results of the corresponding dimensions, and then generate the initial grading results of the data in the fields according to the set comprehensive rating threshold; each dimension is divided into five levels, namely level 1-5, and the higher the level, the greater the influence of the dimension.
[0053] The system administrator can determine the grading standards of the three dimensions according to different industries or different business scenarios. For example, for sensitivity: the core of the confidentiality evaluation lies in assessing the sensitivity level brought about after the data is leaked, and its grading standards are:
[0054] Level 1 is the company profile, product promotion materials, and public recruitment information published on the enterprise official website;
[0055] Level 2 is the employee handbook within the enterprise, general business training materials, and regular communication emails between departments;
[0056] Level 3 is the financial budget report of the enterprise, supplier cooperation agreement, and unpublished new product R & D plan;
[0057] Level 4 includes the core technical formulas of enterprises, the detailed financial information and privacy data of customers, and strategic decision-making documents;
[0058] Level 5 includes the core operation data of national critical infrastructure enterprises, the technical data of military enterprises, and the exclusive business secrets of enterprises.
[0059] Importance: The importance evaluation of data focuses on the consequences caused by data loss or damage, and its grading criteria are as follows:
[0060] Level 1 includes test data in the software development process and regular backup copies of public data;
[0061] Level 2 includes monthly sales reports within departments and intermediate calculation data generated in business processes;
[0062] Level 3 includes order detail data in the enterprise order management system, customer contact records and preference information in the customer relationship management system;
[0063] Level 4 includes the core transaction system data of banks and the key database data of e-commerce platforms;
[0064] Level 5 includes the original gene sequencing data of enterprises and the unique digital archive data of cultural heritage protection units.
[0065] Usage frequency: The evaluation of data usage frequency aims to help enterprises reasonably manage data resources, and its grading criteria are as follows:
[0066] Level 1 includes the historical business archives of enterprises from many years ago and the early design documents of completed projects;
[0067] Level 2 includes financial data for annual financial audits and annual market research analysis reports;
[0068] Level 3 includes the regular statistical reports generated by enterprises every month and the analysis data of marketing activities;
[0069] Level 4 includes the daily reports of enterprises' daily operations and the equipment operation data in the real-time monitoring system;
[0070] Level 5 includes the real-time risk control system data of financial institutions and the intelligent recommendation system data of e-commerce platforms.
[0071] If the scoring results of the importance, confidentiality, and usage frequency of a certain field are 2, 2, and 4 respectively, and the fusion weights are 0.3:0.3:0.4, then the final comprehensive result of this field is 2.8. According to the comprehensive rating threshold, the initial grading result of this field is medium level, as shown in Table 1.
[0072] Table 1 Comprehensive rating threshold
[0073] Comprehensive result x x≤2 2<x≤4 4<x Grading result Low grade Medium grade High grade
[0074] Step 3: Use the data itself, the vector parsing results of the fields, and the scoring results for each dimension as training samples, and use the AutoGluon framework for full-scale training to respectively construct initial scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform;
[0075] Step 4: Use an incremental collector to collect and calculate the increment of new data in the fields. When the amount of new data in the fields exceeds the increment threshold of the original data volume in the fields, perform data distribution detection;
[0076] Use the incremental collector to obtain new data in real time, which is expressed as follows:
[0077] ΔD = {d|T last ≤ d.timestamp < T current}
[0078] where ΔD is the new data set in the field, T last is the last update timestamp, and T current is the current time.
[0079] The data distribution detection uses the KL divergence to judge the data distribution. The calculation formula of the KL divergence is as follows:
[0080]
[0081] where p(x) is the probability distribution of the original data in the field on the feature space X, and q(x) is the probability distribution of the new data in the field on the same feature space X.
[0082] If the KL value ≤ 0.1, update the original data statistics in the field; if the KL value > 0.1, perform model incremental update. Then, use the new data itself, the vector parsing results of the fields, and the scoring results for each dimension as new training samples, and use the AutoGluon framework for incremental training to obtain incremental scoring and grading models for confidentiality, importance, and usage frequency.
[0083] The AutoGluon framework is used to perform model distillation on the incremental scoring and grading model for confidentiality, importance, and usage frequency, resulting in a lightweight incremental scoring and grading model with a smaller volume, lower computational overhead, and high accuracy. The lightweight incremental scoring and grading model is deployed on the big data platform in shadow mode to score new data simultaneously for a period of time in the future. Shadow mode means that before switching the model, the performance and stability of the new model are verified. The new and old models are run simultaneously, and the output is compared to ensure that the results on new and old data are 98% consistent. At the same time, the inference speed and resource consumption rate of the new and old models are compared. If the average inference time of the new model is less than or equal to that of the old model and the growth of memory and CPU usage does not exceed 5%, it is considered to meet the online performance requirements, and the lightweight incremental scoring and grading model is used to replace the initial scoring and grading model for model update; otherwise, if the online performance requirements are not met, the model is not updated.
[0084] Step 5: When the amount of newly added data in the field exceeds the incremental threshold of the original data in the field, perform hierarchical anomaly detection on the newly added data in the collected field simultaneously. If the detection results are low-risk warnings and high-risk warnings, re-score all the data in the field and retrain the model.
[0085] The process of the hierarchical anomaly detection is as follows: The newly added data itself and the vector parsing results of the field are input into the initial scoring and grading model for confidentiality, importance, and usage frequency to score the newly added data, and the corresponding scoring prediction results for confidentiality, importance, and usage frequency are obtained; weighted combination is performed using the fusion weights set in Step 2 to determine the predicted hierarchical result of the newly added data; according to the initial hierarchical result in Step 2, hierarchical anomaly detection is performed on the predicted hierarchical result. If the amount of data for which the predicted hierarchical result of the newly added data is higher than the initial hierarchical result of the original data exceeds the high-risk threshold of the original data in the field, a high-risk warning is fed back, indicating a risk of insufficient data protection due to too low a level; if the amount of data for which the predicted hierarchical result of the newly added data is lower than the initial hierarchical result of the original data exceeds the low-risk threshold of the original data in the field, a low-risk warning is fed back, indicating a possible risk of over-protecting the data.
[0086] The process of the re-scoring and model retraining is as follows: Manually re-score all the data in the field and retrain the model according to the hierarchical criteria for the corresponding dimension, that is, manually re-score all the data in the field according to the hierarchical criteria for the corresponding dimension, and then use the entire data itself, the vector parsing results of the field, and the scoring results for each dimension after re-scoring as new training samples, and use the AutoGluon framework for retraining. After training is completed, replace the original scoring and grading model; at this time, the hierarchical result of this field will be different from the initial hierarchical result, and it needs to be monitored by access frequency monitoring.
[0087] For example, there are currently 10,000 pieces of data in a certain field. The results of the importance, confidentiality, and usage frequency of this field are 2, 2, and 4 respectively, and the fusion weights are 0.3:0.3:0.4. Then the final comprehensive result of this field is 2.8, which is at a medium level. The incremental threshold is 30%, the low-risk threshold is 70%, and the high-risk threshold is 20%. When the newly added data exceeds 3,000 pieces, it is necessary to perform both data distribution detection and hierarchical anomaly detection simultaneously; in the data distribution detection, if the KL value > 0.1 and it is manually determined that the incremental model meets the online performance requirements, then the model is replaced. If the KL value > 0.1 but it is manually determined that the incremental model does not meet the online performance requirements, then it is discarded; if the KL value ≤ 0.1, only the statistics of a certain field are updated; in the hierarchical anomaly detection, if the scoring results of the importance, confidentiality, and usage frequency of a new piece of data are 4, 5, and 4 respectively, and the comprehensive result is 4.3, then the classification result of this new piece of data is at a high level, higher than the medium level. When more than 600 new pieces of data in the field appear in such a situation, that is, exceeding 20% of the incremental data, then a high-risk warning is feedback; if the scoring results of the importance, confidentiality, and usage frequency of a new piece of data are 1, 3, and 2 respectively, and the comprehensive result is 2, then the classification result of this new piece of data is at a low level, lower than the medium level. When more than 2,100 pieces of data in the newly added data appear in such a situation, that is, exceeding 70% of the incremental data, then a low-risk warning is feedback; after manually re-scoring all the data in the field in each dimension, model re-training is performed, and after the training is completed, the original scoring and classification model is replaced.
[0088] Step 6: If the classification level of the newly added data in the field increases after re-scoring, a high-risk warning is feedback. The access frequency of the historical access logs of this field is monitored through a weighted directed graph. If an abnormal warning appears in the access frequency, then all the data in this field is re-scored and the model is re-trained again; if the access frequency is normal, no processing is performed;
[0089] When the classification level of the newly added data in a certain field increases after re-scoring and a high-risk warning is feedback, if the access frequency of this field drops significantly at this time, it may mean that the classification level of the newly added data in a certain field is too high, affecting the normal business process. Before hierarchical anomaly monitoring, an access graph is constructed, where each node represents a field and the edges represent the access correlation between fields, thus establishing a weighted directed graph. The importance of each node is calculated through PageRank iteration. The calculation formula of PageRank is as follows:
[0090]
[0091] In the formula, PR(v) represents the PageRank value of node v, ln(u) represents the set of nodes pointing to node v, Out(u) represents the number of out-edges of node u, and α represents the damping coefficient.
[0092] If the hierarchical level of a certain field is modified and the hierarchical level increases, and the access node frequency related to the certain field significantly decreases, it will lead to a decrease in the PageRank weight of the corresponding edge. Thus, when the abnormal fluctuation of the PageRank weight in the weighted directed graph exceeds the threshold, an abnormal access frequency warning will be triggered.
[0093] Manually re-score all the data in the field with an increased hierarchical level after re-scoring according to the hierarchical rules of the corresponding dimension in accordance with the principle of reducing the hierarchical level. Then, take all the data itself, the vector parsing results of the field, and the scoring results of each dimension after re-scoring as new training samples, and use the AutoGluon framework for re-training. After the training is completed, replace the original scoring and hierarchical model.
[0094] For example, when the original hierarchical result of a certain field is medium level, and after the hierarchical anomaly detection alarm, the hierarchical result is updated to high level manually. During the subsequent operation of the platform, when the abnormal fluctuation of the PageRank weight of this field exceeds 80% and an alarm is triggered, and after manual judgment, it is determined that there is a situation where the hierarchical level is too high and affects the business, then appropriately reduce the hierarchical result, and use all the data itself, the vector parsing results of the field, and the scoring results of each dimension after appropriate reduction as new training samples, and use the AutoGluon framework for re-training. After the training is completed, replace the original scoring and hierarchical model.
[0095] Step 7: When a new data table or field appears in the big data platform and there is no preset multi-dimensional scoring result and hierarchical result, use the current scoring and hierarchical model on the big data platform to perform multi-dimensional scoring on the current data, take the scoring result that appears the most in each dimension as the final scoring result of that dimension, and perform weighted merging according to the fusion weight set in Step 2 to determine the hierarchical result of the data in this field, as Figure 3 shown; the subsequent hierarchical processing is the same as that of the original field.
Claims
1. A big data automated classification and grading method based on a dynamic update feedback mechanism, characterized in that, The steps are as follows: Step 1: Parse the table structure information and the characteristics of the data itself in the Hive metadata, extract the feature vectors of the fields in different tables and different business scenarios, and obtain the vector parsing results of the fields; Step 2: Score the fields in the Hive metadata respectively from three dimensions of sensitivity, importance, and usage frequency according to the grading criteria of the corresponding dimensions, set the fusion weights for weighted merging, and then generate the initial grading results of the data in the fields according to the comprehensive rating threshold; Step 3: Use the data itself, the vector parsing results of the fields, and the scoring results of each dimension as training samples for full-scale training, respectively construct the initial scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform; Step 4: Collect and calculate the increment of the new data in the fields. When the amount of new data in the fields exceeds the increment threshold of the original data in the fields, data distribution detection is performed; determine whether to perform model increment update according to the data distribution detection result: if so, use the new data itself, the vector parsing results of the fields, and the scoring results of each dimension as new training samples for incremental training, obtain the incremental scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform; if not, update the original data statistics in the fields; Step 5: Perform grading anomaly detection on the new data in the collected fields at the same time. If the detection result feedback is a low-risk warning or a high-risk warning, re-score and re-train the entire data in the field; if the grading level of the new data in the field increases after re-scoring and feedback is a high-risk warning, monitor the access frequency of the historical access logs of the field through a weighted directed graph. If an abnormal warning appears in the access frequency, re-score and re-train the entire data in the field again; if the access frequency is normal, no processing is performed; then load the re-trained scoring and grading model into the multi-dimensional scoring and grading model of the big data platform to replace the original scoring and grading model.
2. The big data automatic classification and grading method based on a dynamic update feedback mechanism according to claim 1, wherein: The feature vectors extracted in Step 1 include field metadata and field distribution parameters; the field metadata includes column names, column types, and table names; the field distribution parameters include the minimum value, maximum value, mean value, and standard deviation of the numerical columns; In Step 2, sensitivity, importance, and usage frequency are all divided into five levels, namely levels 1-5, and the higher the level, the greater the influence of this dimension; The full-scale training in Step 3, the incremental training in Step 4, and the re-training in Step 5 are all performed using the AutoGluon framework.
3. The big data automated classification and grading method based on a dynamic update feedback mechanism according to claim 1, wherein: In Step 4, KL divergence is used for data distribution detection, and the calculation formula of KL divergence is as follows: where p(x) is the probability distribution of the original data in the field on the feature space X, and q(x) is the probability distribution of the new data in the field on the same feature space X; If the KL value ≤ 0.1, update the original data statistics in the field; if the KL value > 0.1, perform model increment update.
4. A big data automated classification and grading method based on a dynamic update feedback mechanism according to claim 1, characterized in that: Perform model distillation on the incremental scoring and grading model for confidentiality, importance, and usage frequency described in step 4 to obtain a lightweight incremental scoring and grading model; deploy the lightweight incremental scoring and grading model on the big data platform in shadow mode. If the average inference time of the lightweight incremental scoring and grading model is less than or equal to the average inference time of the initial scoring and grading model and the increase in memory and CPU usage does not exceed 5%, it is considered to meet the online performance requirements. Then, use the lightweight incremental scoring and grading model to replace the initial scoring and grading model for model update; otherwise, if it does not meet the online performance requirements, no model update is performed.
5. A big data automated classification and grading method based on a dynamic update feedback mechanism according to claim 1, characterized in that, The process of hierarchical anomaly detection in step 5 is as follows: Input the vector parsing results of the new data itself and its fields into the initial scoring and grading model for confidentiality, importance, and usage frequency to score the new data, and obtain the corresponding scoring prediction results for confidentiality, importance, and usage frequency. Use the fusion weights set in step 2 for weighted combination to determine the predicted hierarchical result of the new data; according to the initial hierarchical result in step 2, perform hierarchical anomaly detection on the predicted hierarchical result. If the amount of data with a predicted hierarchical result of the new data higher than the initial hierarchical result of the original data exceeds the high-risk threshold of the amount of original data in the field, a high-risk warning is fed back. If the amount of data with a predicted hierarchical result of the new data lower than the initial hierarchical result of the original data exceeds the low-risk threshold of the amount of original data in the field, a low-risk warning is fed back.
6. A big data automated classification and grading method based on a dynamic update feedback mechanism according to claim 1, characterized in that, In the weighted directed graph in step 5, each node represents a field, and the edge represents the access correlation between fields; calculate the importance of each node through PageRank iteration. The formula for PageRank is as follows: In the formula, PR(v) represents the PageRank value of node v, ln(u) represents the set of nodes pointing to node v, Out(u) represents the number of out-edges of node u, and α represents the damping coefficient. When there is an abnormal fluctuation of the PageRank weight exceeding the threshold in the weighted directed graph, an access frequency anomaly warning is triggered.
7. A big data automated classification and grading method based on a dynamic update feedback mechanism according to claim 1, characterized in that, When a new data table or field appears on the big data platform, use the current scoring and grading model on the big data platform to perform multi-dimensional scoring on the current data. Take the most frequently occurring scoring result in each dimension as the final scoring result for that dimension, and perform weighted combination according to the fusion weights set in step 2 to determine the hierarchical result of the data in this field. Subsequent hierarchical processing is the same as that of the original field.
8. A big data automated classification and grading system based on a dynamic update feedback mechanism, characterized in that, Including: Initial scoring and grading model training module: used to parse the table structure information and data itself characteristics in Hive metadata, extract the feature vectors of fields in different tables and different business scenarios, and obtain the vector parsing results of fields. Score the fields in the Hive metadata according to the grading criteria of the corresponding dimension from three dimensions: sensitivity, importance, and usage frequency, set fusion weights for weighted combination, and then generate the initial grading result of the data in the field according to the comprehensive rating threshold; use the data itself, the vector parsing result of the field, and the scoring results of each dimension as training samples for full-scale training, respectively construct the initial scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform; Grading model dynamic update module: used to collect and calculate the increment of the new data in the field. When the increment of the new data in the field exceeds the increment threshold of the original data in the field, perform data distribution detection; determine whether to perform model incremental update according to the data distribution detection result: if so, use the new data itself, the vector parsing result of the field, and the scoring results of each dimension as new training samples for incremental training to obtain the incremental scoring and grading models for confidentiality, importance, and usage frequency, and load them into the multi-dimensional scoring and grading model of the big data platform; if not, update the original data statistics in the field; Grading model closed-loop feedback module: used to perform grading anomaly detection on the new data in the collected field at the same time; if the detection result feedback is a low-risk warning or a high-risk warning, re-score all the data in the field and re-train the model; if the grading level of the new data in the field increases after re-scoring and the feedback is a high-risk warning, monitor the access frequency of the historical access logs of the field through a weighted directed graph. If an abnormal warning appears in the access frequency, re-score all the data in the field and re-train the model again; if the access frequency is normal, do not process; then load the scored and graded model after re-training into the multi-dimensional scoring and grading model of the big data platform to replace the original scored and graded model.
9. A big data automated classification and grading device based on a dynamic update feedback mechanism, characterized in that, Including: Memory: used to store the computer program of a big data automatic classification and grading method based on a dynamic update feedback mechanism according to any one of claims 1-7, and is a computer-readable device; Processor: used to implement a big data automatic classification and grading method based on a dynamic update feedback mechanism according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement a big data automatic classification and grading method based on a dynamic update feedback mechanism according to any one of claims 1-7.