A network data classification and grading processing method, device, equipment and medium

CN118981686BActive Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411042057.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-09-25
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

[0004]但是,现有的方案高度依赖于数据库字段信息的准确性,对于非结构化数据的处理能力相对有限,同时,由于数据类型的多样性和不断变化,现有的方案无法适应新的数据环境

Benefits of technology

[0040]本申请提供的网络数据的分类分级处理方法、装置、设备及介质,涉及数据安全与网络信息管理技术领域。该方法通过对获取到的网络数据进行资产识别,得到识别结果;识别结果包括数据资产、业务资产以及网络资产;根据数据资产的数据特征,使用随机森林模型对数据资产进行分类,得到数据资产的资产类别;针对每个数据资产,通过遍历预设分级规则库中的每一个规则,得到每个数据资产的分级结果;将网络资产、数据资产的资产类别和分级结果、以及业务资产进行关联,得到资产数据;对资产数据分别进行分级计算和分类计算,得到资产数据的业务类型和业务级别。根据本申请,能够将网络数据的分类分级与数据库字段信息解耦,准确且安全的实现网络数据的分类分级。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118981686B_ABST
    Figure CN118981686B_ABST
Patent Text Reader

Abstract

The application provides a network data classification and grading processing method, device, equipment and medium, and belongs to the technical field of data security and network information management. The method comprises: performing asset identification on obtained network data to obtain an identification result; the identification result comprises data assets, business assets and network assets; according to data features of the data assets, a random forest model is used to classify the data assets to obtain asset categories of the data assets; for each data asset, each rule in a preset grading rule library is traversed to obtain a grading result of each data asset; the asset categories and the grading results of the network assets and the data assets and the business assets are associated to obtain asset data; and the asset data is subjected to grading calculation and classification calculation respectively to obtain business types and business levels of the asset data. According to the application, the classification and grading of network data can be decoupled from database field information, and the classification and grading of network data can be accurately and safely realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data security and network information management technology, and in particular to a method, apparatus, device and medium for classifying and grading network data. Background Technology

[0002] In the context of the rapid development of digital information technology, data, as a core element of the digital economy, is becoming increasingly important. However, with the increase in data value, data security risks are also rising simultaneously. Security incidents such as data breaches and cross-border data transfers are frequent, posing a serious challenge to personal privacy and even national security. At the same time, the rapid development of the internet has also spurred an explosive growth in online data, which contains a wealth of critical information. How to accurately identify and extract important and core data from this massive amount of online data, and then reasonably classify and grade the data, has become a major issue in the field of cybersecurity.

[0003] Currently, commonly used data classification and grading methods primarily achieve data classification and grading through detailed identification of database fields. Specifically, each field in the database is first scanned and identified. Based on characteristics such as field name, data type, and data range, the system intelligently determines the category and level of information carried by each field. For example, some fields may contain sensitive personal information, such as ID numbers and phone numbers; these fields will be marked as high-level data. Other common descriptive fields, such as product names and colors, may be classified as low-level data. In this way, data in the database can be automatically classified and graded according to its importance and sensitivity.

[0004] However, existing solutions rely heavily on the accuracy of database field information and have relatively limited ability to process unstructured data. At the same time, due to the diversity and constant changes in data types, existing solutions cannot adapt to new data environments. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for classifying and grading network data, which can accurately and securely classify and grade network data.

[0006] Firstly, this application provides a method for identifying assets from acquired network data to obtain identification results; the identification results include data assets, business assets, and network assets.

[0007] Based on the data characteristics of the data assets, the random forest model is used to classify the data assets to obtain the asset categories of the data assets;

[0008] For each of the data assets, the classification result of each data asset is obtained by traversing each rule in the preset classification rule base;

[0009] The asset data is obtained by associating the network assets, the asset categories and classification results of the data assets, and the business assets.

[0010] The asset data is subjected to hierarchical and categorical calculations to obtain the business type and business level of the asset data.

[0011] In one possible design of the first aspect, the asset identification of the acquired network data includes:

[0012] Obtain the features and tags of the network data;

[0013] Based on the features and labels, feature weights are calculated using the nonlinear model XGBoost to obtain feature importance.

[0014] The feature weights are dynamically adjusted based on the importance of the features to obtain a nonlinear optimized weight calculation formula.

[0015] In one possible design of the first aspect, classifying the data assets using a random forest model based on the data characteristics of the data assets to obtain the asset categories of the data assets includes:

[0016] Obtain the data characteristics and preset classification standard weights of the data assets;

[0017] The data features and the preset classification standard weights are input into the random forest model to calculate the asset category of the data asset.

[0018] In one possible design of the first aspect, obtaining the classification result for each data asset by traversing each rule in a preset classification rule base for each data asset includes:

[0019] For each piece of network data, traverse each hierarchical rule in the preset hierarchical rule base;

[0020] For each of the aforementioned hierarchical rules, a preset algorithm is used for feature extraction;

[0021] Calculate the matching score between the data asset and each of the classification rules;

[0022] The matching score of each of the classification rules is weighted and summed with the weight of the classification rule to obtain the classification result of the data asset.

[0023] In one possible design of the first aspect, the preset algorithm includes one or more of the bag-of-words model algorithm, the term frequency-inverse text frequency index (TF-IDF) algorithm, and the word embedding algorithm.

[0024] In one possible design of the first aspect, calculating the matching score between the data asset and each of the hierarchical rules includes:

[0025] Calculate the similarity between the data asset and the feature representation of each of the hierarchical rules;

[0026] A matching score is calculated based on the similarity.

[0027] In one possible design of the first aspect, calculating the matching score between the data asset and each of the hierarchical rules includes:

[0028] The matching score between the data asset and each of the classification rules is calculated using the Transformer-based algorithm, the Light GBM algorithm, the Meta-Learning algorithm, and the neural network algorithm, respectively.

[0029] Secondly, this application provides a network data classification and grading processing apparatus, the apparatus comprising:

[0030] The identification module is used to identify assets from the acquired network data and obtain identification results; the identification results include data assets, business assets, and network assets.

[0031] The classification module is used to classify the data assets using a random forest model based on the data characteristics of the data assets, thereby obtaining the asset categories of the data assets;

[0032] The grading module is used to obtain the grading result of each data asset by traversing each rule in the preset grading rule base for each data asset.

[0033] The association module is used to associate the network assets, the asset categories and classification results of the data assets, and the business assets to obtain asset data;

[0034] The calculation module is used to perform hierarchical and classification calculations on the asset data to obtain the business type and business level of the asset data.

[0035] Thirdly, this application provides a computer device, including: a transceiver, a processor, and a memory communicatively connected to the processor;

[0036] The memory stores computer-executed instructions;

[0037] The processor executes computer execution instructions stored in the memory to implement the network data classification and grading processing method as described in any one of the first aspects.

[0038] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the network data classification and grading processing method as described in any one of the first aspects.

[0039] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the network data classification and grading processing method as described in any one of the first aspects.

[0040] This application provides a method, apparatus, device, and medium for classifying and grading network data, relating to the fields of data security and network information management technology. The method involves identifying assets in acquired network data to obtain identification results; these results include data assets, business assets, and network assets. Based on the data characteristics of the data assets, a random forest model is used to classify the data assets, resulting in asset categories. For each data asset, a grading result is obtained by traversing each rule in a preset grading rule base. The network assets, the asset categories and grading results of the data assets, and the business assets are associated to obtain asset data. Grading and classification calculations are performed on the asset data to obtain the business type and business level of the asset data. According to this application, the classification and grading of network data can be decoupled from database field information, achieving accurate and secure classification and grading of network data. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0042] Figure 1 The system architecture diagram for classifying and grading network data provided in this application;

[0043] Figure 2 A flowchart illustrating the classification and grading method for network data provided in this application;

[0044] Figure 3 A flowchart illustrating asset identification in the network data classification and grading method provided in this application;

[0045] Figure 4 A flowchart illustrating the data asset classification process in the network data classification and grading method provided in this application;

[0046] Figure 5 A flowchart illustrating the data asset classification process in the network data classification and classification method provided in this application;

[0047] Figure 6 A flowchart illustrating the asset data classification and grading process in the network data classification and grading method provided in this application;

[0048] Figure 7 A schematic structural diagram of the network data classification and grading processing device provided in this application;

[0049] Figure 8 A schematic structural diagram of the computer device provided in this application.

[0050] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0051] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0052] In the context of the rapid development of digital information technology, data, as a core element of the digital economy, is becoming increasingly important. However, with the increase in data value, data security risks are also rising simultaneously. Security incidents such as data breaches and cross-border data transfers are frequent, posing a serious challenge to personal privacy and even national security. At the same time, the rapid development of the internet has also spurred an explosive growth in online data, which contains a wealth of critical information. How to accurately identify and extract important and core data from this massive amount of online data, and then reasonably classify and grade the data, has become a major issue in the field of cybersecurity.

[0053] Currently, commonly used data classification and grading methods primarily achieve data classification and grading through detailed identification of database fields. Specifically, each field in the database is first scanned and identified. Based on characteristics such as field name, data type, and data range, the system intelligently determines the category and level of information carried by each field. For example, some fields may contain sensitive personal information, such as ID numbers and phone numbers; these fields will be marked as high-level data. Other common descriptive fields, such as product names and colors, may be classified as low-level data. In this way, data in the database can be automatically classified and graded according to its importance and sensitivity.

[0054] However, existing solutions rely heavily on the accuracy of database field information and have relatively limited ability to process unstructured data. At the same time, due to the diversity and constant changes in data types, existing solutions cannot adapt to new data environments.

[0055] Based on this, the inventors of this application propose a method for classifying and grading network data, such as... Figure 1 As shown in this application, the acquired network data is first identified as an asset, categorized into network assets, data assets, and business assets. For the data assets, a data classification model, a data grading model, and data processing are sequentially used to obtain the classification and grading results. Then, the classification and grading results of the network assets and data assets are correlated with the business assets to obtain asset data. The asset data is then calculated using a business classification and grading model to obtain the final classification and grading results.

[0056] This application enables the classification and grading of network data by identifying network data assets and combining this with relevant data security classification and grading standards. This provides a reference for data security classification and grading, allowing data processors to adopt different protection measures for different levels of data based on business and organizational needs. Furthermore, it can promptly identify important data not yet covered by existing classification and grading rules, ensuring that all important data, core data, and sensitive personal information are appropriately protected, thereby significantly reducing the risk of data leakage or misuse.

[0057] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0058] Figure 2 A flowchart illustrating the classification and grading method for network data provided in this application. Figure 2As shown, the network data classification and grading processing method of this embodiment may include steps 1100 to 1500:

[0059] Step 1100: Perform asset identification on the acquired network data to obtain identification results; the identification results include data assets, business assets, and network assets.

[0060] Data assets refer to important data, core data, and sensitive personal information explicitly exposed in online data. For example, sensitive personal data such as phone numbers and ID card numbers fall into this category. In addition, data assets can also include important data and core data. Important data may involve a company's financial information, such as annual financial reports and cost data; this data is crucial for a company's decision-making and strategic planning. Core data may include a company's R&D materials, proprietary technology information, or customer data; this data is the core of a company's competitiveness, and its leakage or misuse could cause significant losses to the company. Due to their high sensitivity and importance, this data requires special attention during classification and grading.

[0061] Business assets refer to the medium through which data flows in network data. For example, business assets may include domain names, application programming interfaces (APIs), etc. These assets not only carry the core information of business operations, but are also an indispensable part of data interaction.

[0062] Network assets refer to the infrastructure that constitutes network data. These assets include, but are not limited to, IP addresses, ports, operating systems, databases, and middleware. They provide fundamental support for the storage, processing, and transmission of network data and are key elements for ensuring network security and stable operation.

[0063] It is understandable that the in-depth development of network data classification and grading work focuses precisely on the diverse data assets in cyberspace. Through monitoring and analysis of network data, assets and data not yet included in the classification and grading system are comprehensively identified. This process primarily focuses on data assets, business assets, and network assets.

[0064] Specifically, in combination Figure 3 As shown, in this step, when identifying assets from network data, we can first obtain the characteristics and tags of the network data.

[0065] Specifically, data asset characteristics (X_data) are defined as: data sensitivity, data volume, and storage location security. Network asset characteristics (X_network) are defined as: importance, risk exposure, and usage frequency. Business asset characteristics (X_business) are defined as: relevance, potential impact, and scalability. The label (y) is defined as: asset rating.

[0066] The features and labels of the acquired network data are used to form three datasets: X_data (features of input data assets), X_network (features of input network assets), and X_business (features of input business assets), and corresponding label data (y) is constructed.

[0067] After obtaining the features and labels of the network data, the feature weights are calculated using the nonlinear model XGBoost based on the features and labels to obtain the feature importance.

[0068] In one possible implementation, the steps are as follows:

[0069] import xgboost as xgb

[0070] from sklearn.model_selection importtrain_test_split

[0071] from sklearn.metrics import accuracy_score

[0072] #Dataset Preparation

[0073] X_data = [data sensitivity, data volume, storage location security, ...] # Data asset characteristic matrix

[0074] X_network = [Importance, Risk Exposure, Usage Frequency, ...] #Network Asset Feature Matrix

[0075] X_business = [Relevance, Potential Influence, Scalability, ...] # Business Asset Feature Matrix

[0076] y = [Asset Rating] # Tag

[0077] #Total eigenmatrix

[0078] X_total=X_data+X_network+X_business

[0079] # Divide the dataset into training and test sets

[0080] X_train,X_test,y_train,y_test=train_test_split(X_total,y,test_size=0.2,random_state=42)

[0081] #Initialize the XGBoost model

[0082] model = xgb.XGBClassifier()

[0083] #Training the model

[0084] model.fit(X_train, y_train)

[0085] #Forecasting and Assessment

[0086] y_pred=model.predict(X_test)

[0087] print(f"Accuracy:{accuracy_score(y_test,y_pred)}")

[0088] #Importance of Feature Extraction

[0089] importances=model.feature_importances_

[0090] features = ['Data Sensitivity', 'Data Volume', 'Storage Location Security', 'Importance', 'Risk Exposure', 'Usage Frequency', 'Relevance', 'Potential Impact', 'Scalability']

[0091] feature_importance=dict(zip(features,importances))

[0092] print("Featureimportances:",feature_importance)

[0093] After obtaining the feature importance, the feature weights are dynamically adjusted based on the feature importance to obtain the nonlinear optimization weight calculation formula.

[0094] The importance of the features is as follows:

[0095] {

[0096] 'Data sensitivity': 0.2

[0097] 'Data volume': 0.1

[0098] 'Storage location security': 0.1

[0099] Importance: 0.25

[0100] Risk exposure: 0.15

[0101] 'Frequency of use': 0.1

[0102] 'Correlation': 0.05

[0103] 'Potential Influence': 0.025

[0104] 'Scalability': 0.025

[0105] }

[0106] The optimized weight allocation is based on feature importance scores:

[0107] Data asset weight (W_S):

[0108] W S = (0.2 × data sensitivity) + (0.1 × data volume) + (0.1 × storage location security)

[0109] Network asset weight (W_N):

[0110] W N = (0.25 × Importance) + (0.15 × Risk Exposure) + (0.1 × Frequency of Use)

[0111] Business Asset Weight (W_B):

[0112] W B = (0.05 × Relevance) + (0.025 × Potential Influence) + (0.025 × Scalability)

[0113] Combining the weights of all assets, the total weight is expressed using a comprehensive formula:

[0114] W total =W S +W N +W B

[0115] Step 1200: Based on the data characteristics of the data assets, use a random forest model to classify the data assets and obtain the asset categories of the data assets.

[0116] Data assets are classified according to national laws, regulations, standards, industry standards, and local standards, combined with best practices in the industry. Data characteristics are fully considered to clarify data classification standards and divisions. Data assets are then calculated using a data classification model.

[0117] Specifically, in combination Figure 4 In the process of calculating data assets, the data characteristics and preset classification standard weights of the data assets can be obtained first.

[0118] Data characteristics may include:

[0119] Data Importance: Confidential data, sensitive data, public data. The score ranges from 0 to 10, with higher scores indicating greater data importance.

[0120] Data durability: Temporary data, long-term data retention, with a score of 0 to 10. The higher the score, the longer the data is retained.

[0121] Data usage frequency: High usage frequency, low usage frequency, with a rating of 0 to 10. The higher the value, the higher the data usage frequency.

[0122] The preset classification standard weights can be: national standard weight 0.4, industry standard weight 0.3, local standard weight 0.2, and best practice weight 0.1.

[0123] After obtaining the data characteristics and preset classification standard weights of the data assets, the data characteristics and preset classification standard weights are input into the random forest model to calculate the asset category of the data assets.

[0124] In one possible implementation, the construction process of a random forest model may include:

[0125] import pandas as pd

[0126] from sklearn.model_selection import train_test_split

[0127] from sklearn.ensemble import RandomForestClassifier

[0128] from sklearn.metrics import classification_report,confusion_matrix

[0129] #Sample Data Preparation

[0130] data = {

[0131] 'data_sensitivity':[8,4,7,10,2,6,3,9,5,1],

[0132] 'data_persistence':[9,2,4,10,3,7,2,8,6,1],

[0133] 'data_frequency':[7,3,5,9,2,6,3,8,4,1],

[0134] 'access_restrictions':[9,4,6,10,3,7,2,9,5,1],

[0135] 'label':['confidential','public','sensitive','confidential','public',

[0136] 'sensitive','public','confidential','sensitive','public']

[0137] }

[0138] df = pd.DataFrame(data)

[0139] #Feature matrix (X) and label (y)

[0140] X = df.drop('label', axis = 1)

[0141] y = df['label']

[0142] #Partitioning the dataset

[0143] X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=0.2,random_state=42)

[0144] # Initialize and train the random forest model

[0145] rf_model=RandomForestClassifier(n_estimators=100, max_depth=5, random_state=42)

[0146] rf_model.fit(X_train,y_train)

[0147] #Prediction and Assessment

[0148] y_pred=rf_model.predict(X_test)

[0149] print(confusion_matrix(y_test,y_pred))

[0150] print(classification_report(y_test,y_pred))

[0151] #Feature Importance Analysis

[0152] importances=rf_model.feature_importances_

[0153] features=['data_sensitivity','data_persistence','data_frequency',

[0154] 'access_restrictions']

[0155] feature_importance=dict(zip(features,importances))

[0156] print("Feature Importances:",feature_importance)

[0157] Classification is performed using a trained random forest model; the specific formula depends on the training results and feature importance. Alternatively, the following formula can be used to represent data classification:

[0158] Data input features X = [X1, X2, X3, X4]:

[0159] Class=f(X)=RandomForestModel.predict(X)

[0160] The calculation formula for data asset classification dynamically calculates asset categories based on the model output and the importance values ​​of features.

[0161] Step 1300: For each data asset, obtain the classification result of each data asset by traversing each rule in the preset classification rule base.

[0162] In this step, data asset classification is performed. This is a system that clearly defines and classifies data based on national laws, regulations, standards, industry standards, and local standards, combined with the importance and sensitivity of the data. Through the data classification model, data can be accurately divided into different levels. The calculation formula is as follows:

[0163] [\text{Grading Result} = w_1\cdot\text{Grading Rule Base Matching Calculation} + w_2\cdot\text{Transformer-based} + w_3\cdot\text{LightGBM} + w_4\cdot\text{Meta-Learning} + w_5\cdot\text{Neural Network Grading}]

[0164] Where (w_1,w_2,w_3,w_4,w_5) represent the weights of hierarchical rule base matching calculation, Transformer-based model, LightGBM model, Meta-Learning model and neural network hierarchical in the overall hierarchical classification, respectively.

[0165] Specifically, in combination Figure 5 As shown, in the matching calculation of the hierarchical rule base, for each data asset, each hierarchical rule in the preset hierarchical rule base is traversed; for each hierarchical rule, a preset algorithm is used to extract features; the matching score between the data asset and each hierarchical rule is calculated; the matching score of each hierarchical rule and the weight of the hierarchical rule are weighted and summed to obtain the hierarchical result of the data asset.

[0166] In one example, the calculation formula can be: [\text{ranking result}{\text{rule}}=\sum{i=1}^{N_{\text{rules}}}(\text{similarity}{i}\times\text{rule weight}{i})]

[0167] Where (N_{\text{rules}}) represents the number of rules in the rule base, (\text{similarity}{i}) represents the similarity between the (i)th rule and the input data, and (\text{rule weight}{i}) represents the weight of the (i)th rule.

[0168] This algorithm can extract features from the input data and each rule, and then calculate the similarity between them; the higher the similarity, the higher the matching degree between the input data and the rule; each rule has a weight, which represents its importance in the classification; the weighted sum will take into account the matching scores of all rules and calculate the final classification result according to their weights.

[0169] The preset algorithms include one or more of the following: bag-of-words model algorithm, term frequency-inverse document frequency (TF-IDF) algorithm, and word embedding algorithm.

[0170] When calculating the matching score between the data asset and each hierarchical rule, the similarity between the feature representations of the data asset and each hierarchical rule can be calculated; the matching score is then calculated based on the similarity. In one possible implementation, the matching score between the data asset and each hierarchical rule can be calculated using Transformer-based algorithms, Light GBM algorithms, Meta-Learning algorithms, and neural network algorithms, respectively.

[0171] Specifically, the computation process of a Transformer-based algorithm can include: Input data encoding: First, the Transformer model encodes the input data. The Transformer model converts each word or token in the input data into its corresponding high-dimensional vector representation. These vectors represent the semantic and syntactic information of the input data. Representation summarization or aggregation: Next, the vector representations of these words or tokens are summarized or aggregated to obtain the representation of the entire input data. This step can be performed using methods such as average pooling or max pooling to obtain overall contextual information. Grade prediction: Finally, the overall representation of the input data is input into a grader for grade prediction. The grader can be a simple fully connected layer or other more complex neural network structures. This grader learns the features of the data based on the representation of the input data and outputs the final grade level.

[0172] In one example, the calculation formula can be: [\text{grading result}{\text{Transformer}}=f{\text{grader}}(f_{\text{Transformer}}(\text{input data}))]

[0173] Here, (f_{\text{Transformer}}) represents the encoding function of the Transformer model, which maps the input data into a high-dimensional vector representation; (f_{\text{Classifier}}) represents the classifier function, which is responsible for performing classification prediction based on the input data encoded by the Transformer model.

[0174] The Transformer model captures global dependencies and better understands contextual information when encoding input data through its self-attention mechanism; the summarization or aggregation step allows the integration of information from different parts of the input data into a holistic representation, which helps improve classification performance; the classifier associates the representation of the input data with the corresponding labels through training and learning, thereby achieving classification prediction of the input data.

[0175] Specifically, the computation process of the Light GBM algorithm can include: Model building and training: First, a LightGBM model is built based on training data, and multiple decision trees are trained step by step to enhance prediction capabilities. During training, LightGBM optimizes the tree structure and leaf node splitting of the model by minimizing loss functions (such as mean squared error, log loss, etc.). Input data prediction: For new input data, it is fed into the trained LightGBM model for prediction. Each tree traverses the nodes of the tree according to the feature values ​​of the input data, eventually reaching a leaf node. Each leaf node has a predicted value, which is usually the average of all training samples at that leaf node or the probability value of the class distribution. Weighted summation to obtain the ranking result: Finally, the prediction results of the leaf nodes of multiple trees are weighted and summed to obtain the final ranking result. The predicted value of each leaf node is multiplied by the weight of that tree and then accumulated.

[0176] In one example, the calculation formula can be: [\text{gradation result}{\text{LightGBM}}=\sum{i=1}^{N_{\text{trees}}}(\text{leaf node prediction value}{i}\times\text{leaf node weight}{i})]

[0177] Where (N_{\text{trees}}) represents the number of trees in the LightGBM model, (\text{leaf node prediction value}_{i}) represents the prediction result of the leaf node of the (i)th tree, and (\text{leaf node weight}_{i}) represents the weight of the (i)th tree, which is usually determined by the learning rate during the model training process and the contribution of the tree (such as the number of leaf nodes, depth, etc.).

[0178] The LightGBM model effectively captures complex relationships in the input data by constructing multiple decision trees and utilizing gradient boosting techniques. The prediction result of each tree is based on the predicted value of its leaf node, which represents the average class probability (for classification problems) or average predicted value (for regression problems) of the corresponding leaf node. Weighted summation generates the final classification level by taking into account the contribution of each tree, thereby improving the overall classification accuracy and generalization ability.

[0179] Specifically, the computation process of the Meta-Learning algorithm can include: Meta-learning task learning: First, the meta-learning algorithm learns the features of the task from the meta-learning task set. Feature extraction: Next, based on the learned task features, features are extracted from the input data. Hierarchical prediction: Finally, the extracted features are input into the meta-learning model, and hierarchical prediction is performed through the meta-learning algorithm.

[0180] In one example, the calculation formula can be: [\text{grading result}{\text{Meta-Learning}}=f{\text{Meta}}(\text{input data})]. Here, (f_{\text{Meta}}) represents the meta-learning algorithm model, which receives the input data and outputs the corresponding grading result.

[0181] Meta-learning models achieve adaptability to new tasks by learning features from multiple tasks. This approach allows the model to learn from a small number of samples and make fast and accurate predictions when faced with new tasks. The feature extraction stage is a crucial step, determining how the model extracts useful information from the input data for hierarchical prediction. The meta-learning algorithm model makes hierarchical predictions based on the learned task features and input data; this can be a traditional machine learning algorithm, a neural network model, or other methods.

[0182] Specifically, the computational process of a neural network algorithm can include: Model building and training: First, a neural network model (CNN, RNN) is built. The model can capture complex features in the input data through non-linear mapping via multiple layers of neurons. During the training phase, training data is input into the neural network, and the backpropagation algorithm and optimizer are used to adjust the model parameters to minimize the loss function. Input data prediction: For new input data, it is input into the trained neural network model for prediction. The input data is passed forward through the neural network to the output layer. The output layer uses the softmax function to convert the output of the neural network into class probabilities. Hierarchical prediction: Finally, based on the class probabilities of the output layer, the class with the highest probability is selected as the final hierarchical prediction result.

[0183] In one example, the calculation formula can be: [\text{grading result}{\text{neural network}}=f{\text{NN}}(\text{input data})]. Here, (f_{\text{NN}}) represents the neural network model, which receives the input data and outputs the corresponding grading result.

[0184] Neural network models learn complex features and patterns in input data through multiple layers of neurons and nonlinear activation functions. In the prediction phase, the neural network model passes the input data to each layer of neurons and generates corresponding outputs. These outputs are then processed by a softmax function to represent class probabilities. The final classification is determined based on the class with the highest probability. This approach performs well in a variety of applications, especially in vision and natural language processing tasks.

[0185] Step 1400: Associate the asset categories and classifications of network assets and data assets with business assets to obtain asset data.

[0186] By linking network, hierarchical classification data, and business assets, it is possible to perform calculations through models, sort out business assets within the specified scope, clarify the categories and levels of business assets, help data processors clearly understand the status of business assets, and ensure their security and manageability.

[0187] Step 1500: Perform hierarchical and classification calculations on the asset data to obtain the business type and business level of the asset data.

[0188] Specifically, in combination Figure 6 As shown, the formula for classifying and grading business assets is as follows:

[0189] The classification calculation of business assets includes: assuming the business system contains m types of data, namely Data_1, Data_2, ..., Data_m. For each data type, its weight is set as W_1, W_2, ..., W_m, representing the importance of this data type to the business system type; for each data type, a scoring mechanism is set to give a score value based on the characteristics and sensitivity of this data type, namely Score_1, Score_2, ..., Score_m; calculate the weighted sum of the scores for each data type, i.e.: Total_Score = Σ(W_i * Score_i), i = 1 tom; set the judgment criteria for the business system type based on the numerical range of Total_Score, which can be a threshold or a range interval to determine the type of business.

[0190] The final business system type is determined based on the Total_Score result, and possible types include, but are not limited to: financial business systems, healthcare systems, social media platforms, e-commerce platforms, etc.

[0191] The tiered calculation for business assets includes: assuming there are n data levels (Level_1, Level_2, ..., Level_n), with corresponding weights W_1, W_2, ..., W_n; and considering that the amount of data in the business system is N; for each data level, a weight is assigned to reflect the importance of the data level to the business level, namely Weight_1, Weight_2, ..., Weight_n; for each data level, a score is assigned to represent the contribution of that level, namely Score_1, Score_2, ..., Score_n; the highest score for each data level is calculated, i.e.: Max_Score = max(Score_1, Score_n). The contribution of data level to business level is calculated as follows: Contribution = Weight_i * Max_Score, where i is the index corresponding to the highest level; the contribution of data quantity to business level is calculated as: Total_Contribution_N = N; considering the contributions of data level and quantity, the final business level score is calculated as: Final_Score = Contribution + Total_Contribution_N; the business level judgment criteria are set according to the value range of Final_Score, which can set a threshold or range interval to determine the business level.

[0192] The industry and protection level of a business are determined as follows: General, Important, and Core. The final decision on the industry and level of a business, calculated by the system, rests with the organization's operator. Manual modifications to these levels will not update the industry and level labels of the business assets.

[0193] The network data classification and grading method in this embodiment identifies assets in the acquired network data to obtain identification results. These results include data assets, business assets, and network assets. Based on the data characteristics of the data assets, a random forest model is used to classify the data assets, resulting in asset categories. For each data asset, a grading result is obtained by traversing each rule in a preset grading rule base. The network assets, the asset categories and grading results of the data assets, and the business assets are associated to obtain asset data. Grading and classification calculations are performed on the asset data to obtain the business type and business level of the asset data. According to this application, the classification and grading of network data can be decoupled from database field information, accurately and securely implementing the classification and grading of network data.

[0194] Figure 7 A schematic structural diagram of the network data classification and grading processing device provided in this application. Figure 7As shown, the network data classification and grading processing device 200 provided in this embodiment may include: an identification module 210, a classification module 220, a grading module 230, an association module 240, and a calculation module 250.

[0195] The identification module 210 is used to identify assets from the acquired network data and obtain identification results; the identification results include data assets, business assets and network assets.

[0196] The classification module 220 is used to classify data assets according to their data characteristics using a random forest model to obtain the asset categories of the data assets.

[0197] The grading module 230 is used to obtain the grading result of each data asset by traversing each rule in the preset grading rule base for each data asset.

[0198] The association module 240 is used to associate the asset categories and classification results of network assets and data assets with business assets to obtain asset data;

[0199] The calculation module 250 is used to perform hierarchical and classification calculations on asset data to obtain the business type and business level of the asset data.

[0200] In one embodiment, the identification module 210 can be used to acquire features and labels of network data; calculate feature weights using the nonlinear model XGBoost based on the features and labels to obtain feature importance; and dynamically adjust the feature weights based on feature importance to obtain a nonlinear optimized weight calculation formula.

[0201] In one embodiment, the classification module 220 can be used to obtain the data features and preset classification standard weights of the data assets; input the data features and preset classification standard weights into the random forest model to calculate the asset category of the data assets.

[0202] In one embodiment, the grading module 230 can be used to: traverse each grading rule in the preset grading rule library for each piece of network data; extract features using a preset algorithm for each grading rule; calculate the matching score between the data asset and each grading rule; and perform a weighted summation of the matching score of each grading rule and the weight of the grading rule to obtain the grading result of the data asset.

[0203] The preset algorithms include one or more of the bag-of-words model algorithm, TF-IDF algorithm, and word embedding algorithm.

[0204] In one embodiment, when calculating the matching score between the data asset and each grading rule, the grading module 230 can specifically calculate the similarity between the feature representation of the data asset and each grading rule; and calculate the matching score based on the similarity.

[0205] Specifically, when calculating the matching score between the data asset and each grading rule, the grading module 230 can calculate the matching score between the data asset and each grading rule based on the Transformer-based algorithm, the Light GBM algorithm, the Meta-Learning algorithm, and the neural network algorithm, respectively.

[0206] The network data classification and grading processing device provided in this embodiment can be used to execute the network data classification and grading processing method in any of the aforementioned method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0207] Figure 8 A schematic structural diagram of the computer device provided in this application. (e.g.) Figure 8 As shown, the computer device may specifically include a transceiver 301, a processor 302, and a memory 303. The transceiver 301 receives a first input question, and the memory 303 stores computer execution instructions. The processor 302 executes the computer execution instructions stored in the memory 303 to implement the network data classification and grading processing method described in the above embodiment.

[0208] This embodiment provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, these instructions are used to implement the network data classification and grading processing method described in the above embodiment.

[0209] This embodiment also provides a computer program product, including a computer program that, when executed by a processor, implements the network data classification and grading processing method provided in any of the above embodiments.

[0210] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0211] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for classifying and grading network data, characterized in that, The method includes: Asset identification is performed on the acquired network data to obtain identification results; the identification results include data assets, business assets, and network assets; wherein, data assets refer to important data, core data, and sensitive personal information exposed in the network data; business assets refer to the medium for data flow in the network data; and network assets refer to the infrastructure that constitutes the network data. Based on the data characteristics of the data assets, the random forest model is used to classify the data assets to obtain the asset categories of the data assets; For each of the data assets, the classification result of each data asset is obtained by traversing each rule in the preset classification rule base; The asset data is obtained by associating the network assets, the asset categories and classification results of the data assets, and the business assets. The asset data is subjected to hierarchical and categorical calculations to obtain the business type and business level of the asset data.

2. The method according to claim 1, characterized in that, The asset identification process based on the acquired network data includes: Obtain the features and tags of the network data; Based on the features and labels, feature weights are calculated using the nonlinear model XGBoost to obtain feature importance. The feature weights are dynamically adjusted based on the importance of the features to obtain a nonlinear optimized weight calculation formula.

3. The method according to claim 1, characterized in that, The step of classifying the data assets using a random forest model based on their data characteristics to obtain asset categories includes: Obtain the data characteristics and preset classification standard weights of the data assets; The data features and the preset classification standard weights are input into the random forest model to calculate the asset category of the data asset.

4. The method according to claim 1, characterized in that, The step of obtaining the classification result for each data asset by traversing each rule in the preset classification rule base includes: For each piece of network data, traverse each hierarchical rule in the preset hierarchical rule base; For each of the aforementioned hierarchical rules, a preset algorithm is used for feature extraction; Calculate the matching score between the data asset and each of the classification rules; The matching score of each of the classification rules is weighted and summed with the weight of the classification rule to obtain the classification result of the data asset.

5. The method according to claim 4, characterized in that, The preset algorithm includes one or more of the following: bag-of-words model algorithm, term frequency-inverse text frequency index (TF-IDF) algorithm, and word embedding algorithm.

6. The method according to claim 4, characterized in that, The calculation of the matching score between the data asset and each of the hierarchical rules includes: Calculate the similarity between the data asset and the feature representation of each of the hierarchical rules; A matching score is calculated based on the similarity.

7. The method according to claim 4, characterized in that, The calculation of the matching score between the data asset and each of the hierarchical rules includes: The matching score between the data asset and each of the classification rules is calculated using the Transformer-based algorithm, the Light GBM algorithm, the Meta-Learning algorithm, and the neural network algorithm, respectively.

8. A network data classification and grading processing device, characterized in that, The device includes: The identification module is used to identify assets in the acquired network data and obtain identification results. The identification results include data assets, business assets, and network assets. Among them, data assets refer to important data, core data, and sensitive personal information exposed in the network data; business assets refer to the medium for data flow in the network data; and network assets refer to the infrastructure that constitutes the network data. The classification module is used to classify the data assets using a random forest model based on the data characteristics of the data assets, thereby obtaining the asset categories of the data assets; The grading module is used to obtain the grading result of each data asset by traversing each rule in the preset grading rule base for each data asset. The association module is used to associate the network assets, the asset categories and classification results of the data assets, and the business assets to obtain asset data; The calculation module is used to perform hierarchical and classification calculations on the asset data to obtain the business type and business level of the asset data.

9. A computer device, characterized in that, include: A transceiver, a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the network data classification and grading processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method for classifying and grading network data as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Database masking system and method based on big data

    CN106599713A

  • Context-sensitive switching in a computer network environment

    US20050262005A1