A Structured Data Intelligent Classification and Grading System for Deep Neural Network Models
By combining deep neural networks and multi-layer perceptron regression model, semi-automated labeling and feature extraction of attribute column data is achieved, solving the problems of strong artificial dependence and insufficient data utilization in the existing technology, and improving the efficiency and accuracy of data classification and grading.
Patent Information
- Application Number
- CN202310215953.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-03-08
AI Technical Summary
The existing data classification and grading methods rely on manual discrimination and labeling processing time, the data utilization is insufficient, and the attribute column data feature extraction method is single, making it difficult to efficiently complete data classification and grading.
Combining deep neural network and multi-layer perceptron regression model, semi-automated labeling processing and feature extraction are realized through windowed transformation of attribute column data, self-coding feature extraction and regularization methods, and instantiating attribute column data with sliding windows, a data classification neural network and a data-level multi-layer perceptron regression model are constructed.
It reduces the time of manual discrimination and labeling, fully explores the data information of attribute columns, and improves the efficiency and accuracy of data classification and grading.
Smart Images

Figure CN116257759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of data classification and grading prediction. Background Art
[0002] The state has successively issued the "Cybersecurity Law" and the "Data Security Law", requiring the classification and grading management of data. Relevant industries should pay attention to national regulations and strengthen the standardized control of data. Data is an asset, and data security is related to legal persons, individuals, as well as public and national interests. With the advent of the big data era, the management of data is becoming increasingly important. Data classification and grading is the basis of data security control. If data is not stored by category, it will lead to the mixing and confusion of relevant domain data, and the security management of data cannot be completed, and in severe cases, data loss and leakage problems will occur.
[0003] Deep learning is a key research direction in the field of machine learning, and it has made breakthrough progress in many aspects [1-8]. The deep neural network establishes a model to simulate the neural connection structure of the human brain. When processing various input signals, it describes the data features through multiple neurons, and then gives an interpretation of the data. The machine learning method of the deep neural network shows a high accuracy rate in class discrimination, providing new ideas and concepts for data grading and classification. There are many data classification methods. Among them, the rule-based classification method [9,10] etc. requires a macroscopic control of the data stream, and screens the data according to rules during the data flow process. Its rules are mainly based on the relevant regulatory standards. For such methods, multiple systems or methods often need to be designed according to the complexity of the data to complete the screening and discrimination of the data. For the machine learning method, it can realize the category judgment of relevant data through the feature learning of the data. Both supervised and unsupervised machine learning methods
[11] can complete data classification to a certain extent. Compared with the traditional rule-based data classification and grading method, the machine learning method requires less human participation and relies on relatively fewer rules, while the rule-based data classification and grading method has poor scalability. For new data sets, rules often need to be re-established and the discrimination system needs to be re-built.
[0004] Although there are various existing data classification methods, most data classification methods are still rule-based. They often focus on specific rules and methods, and the concerned domain is relatively limited. Changing the data requires re-establishing the rules. At the same time, rule changes are often abstracted as changes in different systems or functions of functions at the software level, which also means that the rule-based method has poor expansion ability. Although the exploration of using machine learning methods to achieve data classification and grading at the present stage [12-15]Continuing to delve deeper, there are still the following deficiencies in the process: (1) The process of data tagging still heavily relies on manual methods, making it time-consuming to process large volumes of data; (2) In the data classification and grading method, the attribute column data is not fully utilized, often only focusing on the attribute names; (3) It is relatively difficult to extract the recognition features of column data based on machine learning methods.
[0005] The prior art closest to the present application
[0006] In the prior art close to it, the technical solution of "a method for metadata classification and grading based on machine learning algorithms"
[13] A method for metadata classification and grading of a proposed machine learning algorithm creates a frequent item thesaurus using a sensitive original metadata set in the financial field, and uses this thesaurus to convert the features of the class text fields in the corresponding data set into numerical features. Then a binary classification model is constructed to perform sensitive discrimination on the metadata, and finally a multi-classification model is used to complete the subdivision of the data. This method can solve the dependence on manual labor for data classification and grading and improve the classification efficiency. The specific solution process is as Figure 1 .
[0007] The deficiencies of the above technical solution are as follows: (1) When using the method of word sets to perform feature vectorization on class text data, the considered text features are only manifested in the frequency of the data, and other data features of the class text data cannot be concerned; (2) Building the feature vector of class text data with frequent word sets is relatively limited, and it is necessary to collect all the data in this domain as much as possible to achieve accurate vectorization of the data to be characterized; (3) The combinations of words concerned in the construction of the frequent item thesaurus are of three types, and the concerned types can be appropriately expanded.
[0008] In the prior art close to it, the technical solution of "a data classification and grading system and method based on data security and privacy protection"
[16] Using a method that combines rules and industry standards to complete the classification and grading of data. This solution includes multiple subsystems, namely a data receiving subsystem, a data recognition subsystem, a data screening subsystem, a data classification subsystem, and a data grading subsystem; the data receiving subsystem completes data reception; the data recognition subsystem identifies industry data; the data screening subsystem screens industry data; the data classification subsystem uses industry standards to complete data classification; the data grading subsystem uses industry standards to complete data grading. This solution can achieve data classification and grading, and can refine the data classification and grading according to industry standards. The specific solution process is as Figure 2 .
[0009] The deficiencies of the above-mentioned "data classification and grading system and method based on data security and privacy protection" technical solution are as follows: (1). The classification and grading method of this solution is implemented according to industry standards. Since the implementation of the rule standards is system-based, some systems may change with the change of the rule standards, and the scalability is relatively poor to a certain extent. (2). The dependence on manual work is relatively large, and manual input of standard parameters is required. (3). The scenario data that the system focuses on and processes is relatively fixed and is often closely combined with industry standards. Summary of the Invention
[0010] Technical problems to be solved by the present application
[0011] Based on the disadvantages of the existing technologies, the technical problems to be solved by the present invention are summarized as follows:
[0012] (1) There are problems such as too long time for manual discrimination and labeling processing in the existing data classification and grading methods. In the discrimination process, it not only depends on manual discrimination but also on the data information accessed by users in the system for a long time. The process of constructing the classification and grading framework is long and very time-consuming.
[0013] (2) The data utilization is insufficient. The existing methods mainly focus on using data attributes to construct classification and grading models, ignoring the metadata of the attributes, resulting in waste of data information.
[0014] (3) The method for extracting the feature of the attribute column data is relatively single, and more models mainly use the data attribute names for modeling and extracting the feature of the attribute names.
[0015] The present invention aims to solve the pain points in the application of the above-mentioned related methods to the data classification and grading process, and combine the deep neural network method with the data classification and grading architecture to make more full use of data information and complete the classification and grading of data.
[0016] The technical solution of the present invention is summarized as:
[0017] A structured data intelligent classification and grading system for a deep neural network model, characterized by comprising modules: I. Structured data processing module; II. Data labeling processing module; III. Windowed conversion module for attribute column data; IV. Construction and training module for an autoencoder feature extractor; V. Feature transformation module for windowed attribute data; VI. Construction and training module for a data classification neural network model; VII. Construction and training module for a data grading multi-layer perceptron regression model; VIII. Data classification and grading prediction module.
[0018] Overview of the main inventive points of the present invention
[0019] Combined with the above problems, the main innovative technical points of the present invention are summarized as follows:
[0020] (I) Provide a keyword thesaurus and regularization methods to distinguish and judge the categories of columns with existing attribute names, and perform semi-manual labeling processing on the data to be input into the model, so as to solve the problem of long time for full manual judgment and labeling processing.
[0021] (ii) A method for instantiating attribute column data based on a sliding window is proposed to convert a single attribute column into a multi-instance representation, and a feature extraction scheme combining statistical features of the attribute column data and an autoencoder feature extractor is constructed to fully utilize the attribute column data information and mine the potential information of the attribute column data.
[0022] (3) Design and construct a deep neural network for data classification and a multi-layer perceptron regression model for data grading to complete data classification and grading.
[0023] Beneficial effects of this application:
[0024] 1) For the existing data classification method that relies on manual discrimination and labeling, the present invention uses the established regularization method and keyword thesaurus to distinguish and judge the attribute names, and labels the data to be input into the model training to solve the problem of relying on manual labeling. It is convenient and fast, reducing the time of manual discrimination and labeling processing, and reducing the burden on technical personnel.
[0025] 2) In order to solve the problem of insufficient data utilization in existing machine learning methods, the present invention designs a sliding window sampling method to instantiate the attribute column data, and trains an autoencoder feature extractor based on the instance, uses this pre-trained autoencoder to vectorize the instance, and combines the instance artificial statistical feature extraction method to fully mine the internal data features of the instance.
[0026] 3) Considering the actual application scenarios, the combination of artificial features and deep neural network automatic features provides ideas for the feature extraction of structured data attribute columns and can effectively complete data classification and grading. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A flowchart of a technical solution for a metadata classification method based on a machine learning algorithm in the prior art 1
[0028] Figure 2 Flowchart of data classification system and method scheme based on data security and privacy protection in prior art 2
[0029] Figure 3 This is a schematic diagram of the overall technical solution system module of the present invention
[0030] Figure 4 This is a schematic diagram of the overall technical solution of the present invention.
[0031] Figure 5Specific model schematic diagrams of modules related to structuring data tagging, instantiation, and feature extraction
[0032] Figure 6 Specific model schematic diagrams of modules for constructing, training, and predicting data classification and grading models Specific implementation manners
[0033] Combined with the attached Figure 3 、 Figure 4 , describe the implementation process of the overall technical solution
[0034] A structured data intelligent classification and grading system for a deep neural network model, as shown in Figure 3 , the solution includes modules: I. Structured data processing module; II. Data tagging processing module; III. Windowing conversion module for attribute column data; IV. Autoencoder feature extractor construction and training module; V. Windowed attribute data feature transformation module; VI. Data classification neural network model construction and training module; VII. Data grading multi-layer perceptron regression model construction and training module; VIII. Data classification and grading prediction module.
[0035] The implementation manners of each module are introduced below.
[0036] I. Structured data processing module: includes steps S0, S1, where
[0037] Step S0: Formation of structured data. Use the constructed domain raw data set, and adopt ETL, stream data processing, and manual entry methods to integrate the scattered stored data into structured data, and complete data extraction and loading.
[0038] Step S1: During the processing of the structured data set, check the integrity and standardization of the structured data, and process the abnormal row data.
[0039] II. Data tagging processing module: includes steps S2, S3, S4, S5, S6, S7, S8, where
[0040] Step S2: Extract the attribute names of the attribute column data in the data set;
[0041] Step S3: Use the data attribute recognition rules to match the extracted attribute names, and tag the attribute columns if the match is successful;
[0042] Step S4: If not tagged, complete the tagging of the attribute columns using the attribute keyword library constructed in step S5. If tagged, enter S8;
[0043] Step S5: Retrieve the extracted attribute names using the keyword library, and tag the attribute columns if the retrieval is successful;
[0044] Step S6: If not labeled, use the manual method in Step S7 to label the attribute column; if labeled, go to Step S8 to form a labeled attribute data set.
[0045] In the above data labeling processing module, through Steps S3, S4, S5, S6, and S7, a labeled attribute data set in Step S8 is formed. Among them, the keyword library S5 of attribute names is constructed. The keyword library is expanded by using natural language processing technology as the main method and brainstorming as the auxiliary method to complete the synonyms and similar names of a single attribute name, and to identify the attribute names and attribute security levels belonging to each attribute family, thus forming a keyword library. The labeling rule S3 is constructed based on the keyword library, converting each attribute family in the keyword library into a regular expression, and setting the attribute name and attribute security level for each attribute family regular expression.
[0046] III. Attribute column data windowing conversion module: Steps S9, S10, S11, where
[0047] Step S9: Set the window, including setting the window size and window step size;
[0048] Step S10: Use the window size with the window step size as the sliding distance to extract the attribute column data;
[0049] Step S11: Combine with the column label to obtain a set of labeled instances.
[0050] IV. Autoencoder feature extractor construction and training module: Includes Steps S14, S15, S16, S18, where
[0051] Step S14: Instance normalization performs normalization processing on the instances in Step S10 by attribute column;
[0052] Step S15: Construct an autoencoder feature extractor;
[0053] Step S16: Use the normalized instances in Step S14 to train the autoencoder feature extractor constructed in Step S15;
[0054] Step S18: Obtain a pre-trained autoencoder feature extractor model.
[0055] Among them, the construction of the autoencoder feature extractor model in Step S15 is specifically as follows:
[0056] 1) The encoder is described as:
[0057] Let h0 = I′, there is:
[0058]
[0059] The encoder output is expressed as
[0060] 2) Let The decoder is expressed as:
[0061]
[0062] 3) Let The loss function is expressed as:
[0063]
[0064] (The encoder part I′ represents the normalized instance; in the network parameters, W i e represents the encoding weight, b i e represents the encoder bias, σ i e represents the encoder activation function, L e represents the number of layers of the encoder neural network; in the network parameters of the decoder part, W i d represents the decoding weight, b i d represents the decoding bias, σ i d represents the encoder activation function, L d represents the number of layers of the decoder neural network; ζ represents the penalty coefficient.)
[0065] V. Windowed Attribute Data Feature Transformation Module: It includes steps S12, S13, and steps S17, S18, S19, where
[0066] Step S12: Instance statistical feature extraction designs basic statistical features and their transformations using expert knowledge to characterize the instance.
[0067] Step S13: For the instance-transformed statistical features, perform normalization processing on the instance statistical features of the attribute column data.
[0068] Step S17: Before instance vectorization, perform normalization processing on the instance and provide it to S18;
[0069] Step S18: Use the pre-trained autoencoder feature extractor model and enter S19;
[0070] Step S19: Encode the normalized instance using the pre-trained autoencoder feature extractor for the labeled and normalized instance.
[0071] In the above windowed attribute data feature transformation module, steps S12, S13, S17, and S19 transform the same instance into statistical features and autoencoder vectorized feature representations. Among them, the statistical features in step S12 are: arithmetic mean, median, mode, quartiles (3 numbers), interquartile range, range, standard deviation (degrees of freedom is n), skewness, kurtosis, coefficient of variation, standard deviation (degrees of freedom is n - 1), variation ratio, and midrange.
[0072] VI. Data Classification Neural Network Construction and Training Module: It includes steps S20, S21, and S24. Among them,
[0073] Step S20: Construct a data classification deep neural network model for training in step S21;
[0074] Step S21: Input the normalized statistical features of the instances in step S13 into the statistical feature learning layer and input the labeled instance vectors in step S19 into the encoded instance learning layer. Fuse the two types of vectors learned by the statistical feature learning layer and the encoded instance learning layer, and input them into the fused feature learning layer. Adjust the model parameters with the predicted instance attribute category and the actual instance attribute category output by the fused feature learning layer under the cross-entropy loss function to complete the training of the data classification neural network model. During the training process, introduce a temperature coefficient to act on the model and adjust the adaptability of the data classification neural network model to different classifications.
[0075] Step S24: Obtain the trained data classification deep neural network model for entering step S26.
[0076] Among them, the data classification deep neural network model in S20 is specifically:
[0077] 1) The learning function of the statistical feature learning layer is The model learns the statistical characteristics of the instance features and transforms the m-dimensional instance statistical features into k-dimensional vectors.
[0078] 2) The learning function of the encoded instance learning layer is The model adjusts the influence of the encoded vector on the model here.
[0079] 3) The learning function of the fused feature learning layer is The model learns the classification function κ from the fused feature vectors of the instance feature vectors and the statistical feature vectors, realizes the learning and classification of the two types of transformed vectors, and completes the category judgment of the data.
[0080] 4) The model optimization problem is:
[0081]
[0082]
[0083] (y is the predicted class set of the classifier, θ1, θ2, and θ3 are the hyperparameters of the corresponding models respectively, y′ is the original class set, T is the temperature coefficient, M is the number of labels, and N is the number of sample instances.)
[0084] VII. Data Hierarchical Multilayer Perceptron Regression Model Construction Module: It includes steps S22, S23, and S25, where
[0085] Step S22: Construct a data hierarchical multilayer perceptron regression model and proceed to step S23 for training;
[0086] Step S23: Concatenate the instance normalized statistical features in step S13 and the labeled instance vectors in step S19 to represent the instance conversion features. Input these instance conversion features into the model, and use the model to predict the security level of the attribute instance and the actual security level of the attribute instance. Complete the training of the data hierarchical multilayer perceptron regression model constructed in step S22 under the mean squared error loss function, and proceed to step S25.
[0087] Step S25: Obtain the trained data hierarchical multilayer perceptron regression model for use in step S26.
[0088] Among them, each layer of the neural network in the S22 data hierarchical multilayer perceptron regression model is defined as:
[0089]
[0090] ( is the activation function, L is the number of neural network layers, n (l) is the number of neurons in the l-th layer, O i (l) represents the output of the i-th neural network in the l-th layer, w j,i (l) and w 0,i (l) are the weight parameter and bias parameter corresponding to the neuron respectively.)
[0091] VIII. Data Hierarchical Classification Prediction Module: It includes step S26, where
[0092] Step S26: Take the most frequent number of output attribute classes of the S24 data classification neural network model as the final discriminant attribute class; in the attribute security level judgment, take the most frequent number of output attribute security levels of the S25 data hierarchical multilayer perceptron regression model as the final discriminant attribute security level, output the result, and the program ends.
[0093] For the data of the attribute column to be predicted, generate attribute column instances according to steps S9 and S10, obtain the normalized instances using step S14, and input the normalized instances into the S18 autoencoder feature extractor to obtain the vectorized instances. For the attribute column instances generated in steps S9 and S10, generate instance statistical features using step S12 and normalize the instance statistical features in step S13. Input the normalized instance statistical features in step S13 and the vectorized instances generated by the S18 autoencoder feature extractor into the S24 data classification neural network model respectively to complete the judgment of the attribute category of the instance, and input them into the S25 data hierarchical multi-layer perceptron regression model to complete the judgment of the attribute security level of the instance. In the attribute category judgment, take the most frequent number of times the S24 data classification neural network model outputs the attribute category as the final discriminant attribute category; in the attribute security level judgment, take the most frequent number of times the S25 data hierarchical multi-layer perceptron regression model outputs the attribute security level as the final discriminant attribute security level.
[0094] The technical details of the important modules in the system are further described in detail below.
[0095] Such as Figure 5 shown, the detailed implementation processes of the structured data processing module and the data tagging processing module:
[0096] 1) Adopt the methods of ETL, streaming data processing, and manual entry to integrate the scattered stored data to form the S0 structured data set SD = {(attr0, x′0),...,(attr d , x′ d )}. In step S1, perform integrity and normalization detection on the attribute column data x′ = {x′0,..., x′ d} in the structured data set SD, and delete the abnormal data rows and missing value data rows therein to obtain x = {x0,..., x d}. Then the processed structured data set is D = {(attr0, x0),...,(attr d , x d )}.
[0097] 2) Take the attribute set A = {attr0,..., attr d} in the structured data set D, and sequentially match the attributes in A using the S3 tagging rules constructed by regular expressions. If the match is successful, add the attribute and the corresponding tag tag in the tagging rules to the attribute tagging set AT part , otherwise input the attribute into the S5 keyword library for search. If the search is successful, add the attribute and the corresponding tag tag in the keyword library to AT part , otherwise complete the identification using manual methods and integrate AT partObtain \(AT = {(attr0, tag0),..., (attr d , tag d ) \}\). In this stage, combine \(x=\{x0,..., x d \}\) and \(AT\) to obtain the S8-tagged attribute dataset \(DTag=\{(attr0, x0, tag0),..., (attr d , x d , tag d ) \}\). (Where \(AT part = {(attr i , tag i ),..., (attr j , tag j ) | 0 ≤ i, j ≤ m, i ≤ j \}\), the tag \(tag\) contains two parts, one part is the family name of the attribute family corresponding to the structured attribute name, and the other part is the security level of the attribute family corresponding to the structured attribute name.)
[0098] 3) In step S9, set the window size to \(w\) and the window step to \(s\). In step S10, take the attribute column data \(x i \) in \(DTag\), for the data in \(x i \), intercept the data of window size \(w\) as an instance; move the step \(s\), repeat the selection of instances; combine the tags in \(DTag\) to tag the instances, and obtain the S11-tagged instance set \(ITag=\{(I0, tag0),..., (I d , tag d ) \}\).
[0099] As Figure 5 shown, the detailed implementation processes of the attribute column data windowing conversion module, the autoencoder feature extractor construction and training module, and the windowed attribute data feature transformation module:
[0100] 1) The normalization method for the instance set \(I = \{I0,..., I d \}\) in the tagged instance set \(ITag\) in steps S14 and S17 is as follows: Take the attribute column data \(x = \{x0,..., x d \}\), calculate the maximum absolute value for each column of data to obtain \(CL = [m0,..., m d \). Divide all the data in the attribute column instance set \(I0\) by the maximum absolute value \(m0\) to obtain the normalized instance set \(I0'\). Repeat the above steps until all the attribute column instances are normalized, and obtain all the attribute normalized instance set \(I'=\{I'0,...., I' d \}\).
[0101] 2) In the construction of the S15 auto-encoder feature extractor model, the activation functions of the encoder and decoder are tanh, the loss function is MSE, and the optimizer is SGD. In step S16, the model is trained on the training set, the optimal model is found on the validation set, and the optimal model is verified on the test set. A penalty coefficient ζ is introduced during the training process to adjust the model.
[0102] 3) For any instance (I j,i , tag j,i ) in the labeled instance set ITag, I j,i is normalized according to the method described in step S17, and in step S19, the trained S18 auto-encoder feature extractor is used to encode I j,i into E j,i ; in step S12, using expert knowledge, the basic statistical features of the instance in instance I j,i are extracted, added, and converted into S j,i . The vectorized feature E j,i of the instance and the basic feature S j,i of the instance are set with the label tag j,i . The above steps are repeated until all labeled instances are converted into instance statistical features and instance vectorized features, denoted as ESTag = {(E0, S0, tag0),..., (E d , S d , tag d )}. Here, the expert knowledge is the basic statistical features and their conversions for constructing the description of the data distribution. The skewness, kurtosis, coefficient of variation, and coefficient of dispersion are used as the instance conversion statistical features, and the results obtained by dividing the remaining statistical values in pairs are also used as the instance conversion statistical features.
[0103] 4) The normalization method in step S13 is as follows: The attribute column statistical features S i in the set ESTag are vertically concatenated into VS, and the maximum absolute value of each column of data in VS is calculated to obtain SL = [m0,..., m v . Each column of data in S i is divided by the corresponding maximum absolute value in SL to obtain the normalized set S′ i of statistical features. The above steps are repeated until all attribute column statistical features are normalized, and ES′Tag = {(E0, S′0, tag0),...,(E d , S′ d , tag d )} is obtained.
[0104] As Figure 6 shown, the detailed implementation process of the data classification and grading model construction, training, and prediction module:
[0105] 1) The learning function of the statistical feature learning layer in the data classification neural network model constructed in step S20 The learning function φ of the encoding instance learning layer and the learning function κ of the fusion feature learning layer are composed of stacked multi-layer fully connected neural networks. Among them, the activation function of the hidden layer is relu, the activation function of the classifier layer is softmax, the loss function uses cross-entropy, and the optimizer is SGD. Combining the attribute column feature set ES′Tag = {(E0, S′0, tag0),..., (E d , S′ d , tag d )} obtained in steps S13 and S19 and the data classification deep neural network model in S20, the model training process in step S21 is described as follows: The data set is divided into a training set, a validation set, and a test set. For the instance feature (E i,j , S′ i,j , tag i,j ) representing this attribute column, input E i,j into the function φ, and input S′ i,j into the function Concatenate the result φ(E i,j ) and and input them into the function κ. In the classifier layer, divide by the temperature coefficient T before the data passes through the loss function softmax to adjust the model. Train the model on the training set, find the optimal model on the validation set, and complete the verification of the optimal model on the test set.
[0106] 2) The activation function in the data grading multi-layer perceptron regression model constructed in step S22 is relu, and the optimizer is Adam. Combining the attribute column feature set ES′Tag obtained in steps S13 and S19 and the data grading multi-layer perceptron regression model in S22, the model training process in step S23 is described as follows: The data set is divided into a training set and a test set. For the instance feature (E i,j , S′ i,j , tag i,j ) representing this attribute column, concatenate E i,j and S′ i,j and input them into the model, and use the corresponding attribute security level in tag i,j to update the model parameters. Train the model on the training set and find the optimal model on the test set.
[0107] 3) The model prediction process in step S26 is as follows:
[0108] (1) For the prediction column data, obtain multiple instance sets I pre representing this column through steps S9 and S10.
[0109] (2) Use steps S12, S13, steps S14, S18, and S19 to convert all instances I pre into the normalized set S′ of instance statistical features pre and the vectorized set E of instances pre .
[0110] (3) Input (E pre , S′ pre ) into the S24 data classification deep neural network model to obtain the prediction result C pre . Statistically analyze the classification situation predicted in C pre , and select the attribute corresponding to the most predicted classification as the final discriminant attribute.
[0111] (4) Input (E pre , S′ pre ) into the S25 data grading multi-layer perceptron regression model to obtain the prediction result Level pre . Round all the data in the predicted grading result to the nearest integer to obtain Level′ pre . Statistically analyze the grading situation predicted in Level′ pre , and select the most predicted grading as the final discriminant safety level.
[0112] The key innovative technical points of this application include:
[0113] 1) Utilize prior knowledge to complete the construction of domain data attribute name keywords and regularization methods, perform semi-automated labeling on data attribute columns, rely less on the manual labeling process, and play a guiding role in the classification and grading of data and the labeling of training model data.
[0114] 2) Establish a conversion model from attribute column data to instances, from instances to vectors and basic features, providing a method for the conversion of attribute column data. The conversion model includes windowing of attribute column data, vectorization of windowed instances, and extraction of statistical features of windowed instances. Among them, the extraction of statistical features of windowed instances utilizes scene expert knowledge to make the feature extraction more in line with the scene and the features better represent the data characteristics.
[0115] 3) Combine artificial features and automated feature extraction by deep neural networks to fuse scene knowledge and potential information in the data.
[0116] 4) Construct two parts of models for data classification and data grading, and effectively utilize the important information features of the data to complete data classification and grading.
[0117] Literature materials helpful for understanding the technology of the present invention
[0118] [1] Covington P, Adams J, Sargin E. Deep neural networks for YouTube recommendations. ACM conference on recommender systems. New York, NY, USA: Association for Computing Machinery. 2016: 191 - 198.
[0119] [2] Larochelle H, Bengio Y, Louradour J, et al. Exploring strategies for training deep neural networks. Journal of Machine Learning Research, 2009, 10(1): 1 - 40.
[0120] [3] Montavon G, Samek W, Müller K - R. Methods for interpreting and understanding deep neural networks. Digital Signal Processing, 2018, 73: 1 - 15.
[0121] [4] Montufar G F, Pascanu R, Cho K, et al. On the number of linear regions of deep neural networks. International Conference on Neural Information Processing Systems. Cambridge, MA, United States: MIT Press. 2014: 2924 - 2932
[0122] [5] Moosavi Dezfooli S M, Fawzi A, Frossard P. Deepfool: a simple and accurate method to fool deep neural networks. the IEEE conference on computer vision and pattern recognition. Las Vegas, NV, USA: IEEE Press. 2016: 2574 - 2582
[0123] [6]SZE V,CHEN Y-H,YANG T-J,et al.Efficient processing of deep neuralnetworks:A tutorial and survey.Proceedings of the IEEE 2017,105(12):2295-2329.
[0124] [7]SZEGEDY C,TOSHEV A,ERHAN D.Deep neural networks for objectdetection.International Conference on Neural Information ProcessingSystems.Red Hook,NY,United States Curran Associates Inc.2013:2553-2561.
[0125] [8]YOSINSKI J,CLUNE J,BENGIO Y,et al.How transferable are features indeep neural networks?Proceedings of the 27th International Conference onNeural Information Processing Systems.Cambridge,MA,United States:MITPress.2014:3320-3328
[0126] [9]Gao Lei,Zhao Zhangjie,Lin Yeli,et al.Research on Data Classification and Grading Method Based on the Data Security Law.Information Security Research,2021,7(10):933-940
[0127]
[10] Song Shaohong,Chen Zhang.A Data Classification and Grading Method Based on Financial Industry Data Security:China,Application Date CN202111539492.2021.12.15
[0128]
[11] He Wenzhu,Peng Changgen,Wang Maoni,et al.Sensitive Attribute Recognition and Grading Algorithm for Structured Datasets.Computer Applications in Research,2020,37(10):3077-3082.
[0129]
[12] Lu Hongtai. Urban Data Classification and Grading Method Based on Deep Learning Clustering Algorithm. Industrial Technology Innovation, 2021, 8(4): 73-78.
[0130]
[13] Wu Mingguang, Guo Huiru, Liu Qiong, et al. A Metadata Classification and Grading Method Based on Machine Learning Algorithm: China, CN202210300625. Application Date: March 25, 2022
[0131]
[14] ZHANG Q, ZHANG C, NI J, et al. Data Sensitivity Measurement and Classification Model of Power IOT Based on Information Entropy and BP Neural Network. International Conference on Advanced Algorithms and Control Engineering. IOP Publishing Ltd. 2021.
[0132]
[15] Yu Yihan, Fu Yu, Wu Xiaoping. Privacy Data Measurement and Grading Model Based on Shannon Information Entropy and BP Neural Network. Journal of Communications, 2018, 39(12): 10-17.
[0133]
[16] Jin Huasong, He Ying, Lai Xiaoyou, et al. Data Classification and Grading System and Method Based on Data Security and Privacy Protection: China, CN202110923721. Application Date: August 12, 2021
[0134] Abbreviations and Definitions of Key Terms
[0135] ETL Process: The entire process in which data enters the data warehouse after being extracted, transformed, and loaded.
[0136] Deep Neural Network: A neural network with two or more hidden layers is usually called a deep neural network.
[0137] Streaming Data: Refers to data continuously generated by thousands of data sources, usually sent in the form of data records at the same time, and the scale is relatively small.
[0138] Regular Expression: A logical formula that operates on strings (including ordinary characters (e.g., letters between a and z) and special characters (called "meta-characters")).
[0139] Natural Language Processing: Using machine learning to analyze the structure and meaning of text. With natural language processing applications, organizations can analyze text and extract information about people, places, and events to better understand the sentiment of social media content and customer conversations.
[0140] Training set: Refers to the set of samples used for training, mainly used to train the parameters in a neural network.
[0141] Validation set: A set of samples used to validate the performance of a model.
[0142] Test set: A set of samples used to test the performance of a model.
[0143] Feature extraction: It starts from an initial set of measured data and then constructs informative and non-redundant derived values, called feature values.
[0144] Database: An organized collection of structured information or data (usually stored electronically in a computer system), typically controlled by a database management system.
[0145] Classifier: A general term for methods in data mining that classify samples, including algorithms such as decision trees, logistic regression, naive Bayes, and neural networks.
[0146] Dense: A commonly used fully connected layer. The operation it implements is output = activation(dot(input, kernel)+bias). Here, activation is an element-wise activation function, kernel is the weight matrix of this layer, and bias is the bias vector.
[0147] Activation function: A function that runs on the neurons of an artificial neural network and is responsible for mapping the input of a neuron to the output end.
[0148] Loss function: A function that maps the values of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of the random event.
[0149] Optimizer: During the backpropagation process in deep learning, it guides each parameter of the loss function (objective function) to update in the correct direction by an appropriate amount, so that the updated parameters make the value of the loss function (objective function) continuously approach the global minimum.
[0150] Temperature coefficient: Acts on the softmax activation function to adjust the degree of attention to difficult samples: the smaller the temperature coefficient, the more it focuses on separating this sample from the most similar other samples.
[0151] Cross entropy: An important concept in Shannon information theory, mainly used to measure the differential information between two probability distributions.
[0152] MSE: The mean square error is a measure that reflects the degree of difference between an estimator and the parameter being estimated. Let t be an estimator of the population parameter θ determined from a sample, and (θ - t) 2 The mathematical expectation of is called the mean square error of the estimator t. It is equal to σ 2 + b 2 , where σ 2 and b are the variance and bias of t respectively.
Claims
1. A structured data intelligent classification and grading system for a deep neural network model, characterized in that, Included modules: Ⅰ. Structured data processing module; Ⅱ. Data tagging processing module; Ⅲ. Windowed conversion module for attribute column data; Ⅳ. Autoencoder feature extractor construction and training module; Ⅴ. Windowed attribute data feature transformation module; Ⅵ. Data classification neural network model construction and training module; Ⅶ. Data grading multi-layer perceptron regression model construction and training module; Ⅷ. Data classification and grading prediction module; The windowed conversion module for attribute column data: Steps S9, S10, S11, where Step S9: Set the window, including setting the window size and window step; Step S10: Use the window size with the window step as the sliding distance to extract the attribute column data; Step S11: Combine with the column label to obtain the tagged instance set; The windowed attribute data feature transformation module: Includes steps S12, S13, and steps S17, S18, S19, where Step S12: Extract instance statistical features, design basic statistical features and their transformations using expert knowledge to characterize the instance; Step S13: For the transformed statistical features of the instance, normalize the instance statistical features of the attribute column data; Step S17: Before instance vectorization, normalize the instance and provide it to S18; Step S18: Use the pre-trained autoencoder feature extractor model and enter S19; Step S19: Encode the normalized instance using the pre-trained autoencoder for the tagged and normalized instance; In the above windowed attribute data feature transformation module, steps S12, S13, S17, S19 transform the same instance into statistical features and autoencoder vectorized feature representations.
2. The intelligent classification and grading system according to claim 1, wherein The structured data processing module: Includes steps S0, S1, where Step S0: Formation of structured data. Use the constructed domain raw data set, adopt ETL, stream data processing, and manual entry methods to integrate the dispersed stored data into structured data, and complete data extraction and loading; Step S1: During the structured data set processing, check the integrity and normativity of the structured data, and process the abnormal row data.
3. The intelligent classification and grading system according to claim 1, characterized in that, The data tagging processing module: Includes steps S2, S3, S4, S5, S6, S7, S8, where Step S2: Extract the attribute names of the attribute column data in the data set; Step S3: Use the data attribute recognition rules to match the extracted attribute names. If the match is successful, tag the attribute column; Step S4: If not tagged, complete the tagging of the attribute column using the attribute keyword library constructed in step S5. If tagged, enter S8; Step S5: Retrieve the extracted attribute names using the keyword library. If the retrieval is successful, tag the attribute column; Step S6: If not tagged, use the manual method in step S7 to tag the attribute column. If tagged, enter step S8 to form the tagged attribute data set; In the above data tagging processing module, the tagged attribute data set in step S8 is formed through steps S3, S4, S5, S6, S7.
4. The intelligent classification and grading system according to claim 1, characterized in that, The autoencoder feature extractor construction and training module: Includes steps S14, S15, S16, S18, where Step S14: Instance normalization performs normalization processing on the instances in Step S10 by attribute columns; Step S15: Construct an autoencoder feature extractor; Step S16: Use the normalized instances in Step S14 to train the autoencoder feature extractor constructed in Step S15; Step S18: Obtain a pre-trained autoencoder feature extractor model; The construction of the autoencoder feature extractor model in Step S15 is specifically as follows: The encoder is described as: Let , there is: <1> The encoder output is represented as ; Let , the decoder is expressed as: <2> Let , the loss function is expressed as: <3> Encoder section Represents a normalization instance; among the network parameters Represents the encoding weights, Represents the encoder bias, Represents the encoder activation function, Represents the number of layers of the encoder neural network; among the network parameters of the decoder section Represents the decoding weights, Represents the decoding bias, Represents the encoder activation function, Represents the number of layers of the decoder neural network; Represents the penalty coefficient.
5. The intelligent classification and grading system according to claim 1, wherein: The data classification neural network model construction and training module: includes Steps S20, S21, and S24, where Step S20: Construct a data classification deep neural network model for training in Step S21; Step S21: Input the instance normalization statistical features in Step S13 into the statistical feature learning layer and input the labeled instance vectors in Step S19 into the encoded instance learning layer. Fuse the two types of vectors learned by the statistical feature learning layer and the encoded instance learning layer, and input them into the fused feature learning layer. Adjust the model parameters with the predicted instance attribute category output by the fused feature learning layer and the actual instance attribute category under the cross-entropy loss function to complete the training of the data classification neural network model; Step S24: Obtain a trained data classification deep neural network model; The data classification deep neural network model in Step S20 is specifically as follows: 1) The learning function of the statistical feature learning layer is , the model learns the statistical characteristics of the instance, and converts the m-dimensional instance statistical features into k-dimensional vectors; 2) The learning function of the encoding example learning layer is , and the model adjusts the influence of the encoding vector here; 3) The learning function of the fusion feature learning layer is , and the model learns the classification function from the fusion feature vector of the instance feature vector and the statistical feature vector , realizes the learning and classification of the two types of transformation vectors, and completes the category judgment of the data; 4) The model optimization problem is: <4> <5> is the classifier prediction class set, are the hyperparameters of the corresponding models respectively, is the original class set, is the temperature coefficient, is the number of labels, is the number of sample instances.
6. The intelligent classification and grading system according to claim 1, characterized in that, The data grading multi-layer perceptron regression model construction module: includes Steps S22, S23, and S25, where Step S22: Construct a data grading multi-layer perceptron regression model for training in Step S23; Step S23: Concatenate the instance normalization statistical features in Step S13 and the labeled instance vectors in Step S19 to represent the instance conversion features. Input these instance conversion features into the model, use the model to predict the attribute instance security level and the actual attribute instance security level, and complete the training of the data grading multi-layer perceptron regression model constructed in Step S22 under the mean squared error loss function and enter Step S25; Step S25: Obtain a trained data grading multi-layer perceptron regression model Each layer of the neural network in the data grading multi-layer perceptron regression model in S22 is defined as: <6> is the activation function, is the number of neural network layers, is the number of neurons in the th layer, represents the output of the th neural network in the , are the weight parameter and bias parameter corresponding to the neuron respectively.
7. The intelligent classification and grading system according to claim 1, characterized in that, The data grading classification prediction module includes Step S26, where Step S26: Use the most frequent number of times the attribute category is output by the data classification neural network model in S24 as the final discriminant attribute category; in the attribute security level judgment, use the most frequent number of times the attribute security level is output by the data grading multi-layer perceptron regression model in S25 as the final discriminant attribute security level, output the result, and the program ends; Input the normalized instance statistical features in Step S13 and the vectorized instances generated by the autoencoder feature extractor in S18 into the data classification neural network model in S24 to complete the judgment of the instance attribute category, and input them into the data grading multi-layer perceptron regression model in S25 to complete the judgment of the instance attribute security level; In the attribute category judgment, use the most frequent number of times the attribute category is output by the data classification neural network model in S24 as the final discriminant attribute category; In the judgment of the attribute security level, the maximum number of times that the S25 data classification multi-layer perceptron regression model outputs the attribute security level is used as the final judgment of the attribute security level.
Citation Information
Patent Citations
Data classification method based on data security and privacy protection
CN113627535B
Metadata grading and classifying method based on machine learning algorithm
CN114676253A
A method of Clothing Attribute Prediction with Auto-Encoding Transformations
AU2020102476A4
Dialogue information extraction method and system, and computer-readable storage medium
CN113822058A