An enterprise data asset hierarchical management method and system
By collecting and processing enterprise data assets and using natural language processing and K-means technology to classify data, the problem of unclear data classification in enterprise data asset grading management is solved, precise data grading and protection is achieved, and management efficiency and security are improved.
Patent Information
- Application Number
- CN202510148876.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-11
AI Technical Summary
There are problems with unclear data classification in enterprise data asset hierarchical management, which leads to security and management challenges, manual classification is inefficient and error-prone, and different departments use different classification methods and lack a unified framework, resulting in inconsistent data protection measures.
A corporate data asset grading management method is adopted to collect and initially process the enterprise's comprehensive asset data, use improved natural language processing technology and K-means clustering analysis to obtain similar grading data sets, and hierarchical matching is carried out through the grading scheme design model to achieve accurate data classification and grading.
It realizes more accurate data classification and grading, replaces a large number of manual operations, improves data processing efficiency, avoids resource waste and increases in security management costs, and ensures reasonable data grading and protection.
Smart Images

Figure CN119621973B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of enterprise data asset hierarchical management, and more specifically, to a method and system for enterprise data asset hierarchical management. Background Art
[0002] Enterprise data assets are important data of the company. When hierarchical management is required, the problem of unclear data classification will arise, which will lead to a series of security and management challenges.
[0003] Manual classification systems are inefficient, their classification functions are inaccurate and prone to errors, or the definition of data classification standards is vague or too broad, which may lead to the incorrect classification of sensitive data as non-sensitive, or vice versa, resulting in increased risk of data leakage or waste of resources on protecting unnecessary data. The sensitivity of data may change over time or with changes in the business environment, resulting in data that is no longer sensitive being over-protected, while newly generated sensitive data may not receive the protection it deserves.
[0004] In addition, in the process of data classification, different departments or teams may use different classification methods. There is no unified enterprise-level framework, which leads to confusion when data flows across departments and makes it difficult to implement consistent data protection measures.
[0005] Therefore, based on the above problems, it is necessary to design hierarchical management of enterprise data assets. Summary of the invention
[0006] In view of the problems existing in the prior art, the purpose of the present invention is to provide a method and system for hierarchical management of enterprise data assets to achieve hierarchical management of enterprise data assets.
[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solution: The enterprise data asset hierarchical management method comprises:
[0008] Step S1: Collect the comprehensive asset data and data classification standards of the enterprise within the cycle time, and perform preliminary processing on the comprehensive asset data to obtain preliminary comprehensive data;
[0009] Step S2: Use improved natural language processing technology to perform key feature analysis on each text in the preliminary comprehensive data, collect word conversion vectors corresponding to feature words in all texts, and obtain feature comprehensive data;
[0010] Step S3: Use K-means technology to perform cluster analysis on each text in the feature comprehensive data to obtain a similar graded data set;
[0011] Step S4: Use the grading scheme design model to perform grade matching on each group in the similar graded data set to obtain the data classification level corresponding to each group;
[0012] Step S5: Manage the data using the protection measures in the data classification standard according to the data classification level corresponding to each group.
[0013] Preferably, the comprehensive asset data includes basic asset data and metadata and usage data of the basic asset data;
[0014] Basic asset data refers to the data itself that needs to be managed hierarchically;
[0015] The metadata of basic asset data includes attribute data and open source data. Attribute data includes file format, size, creation date, modification date, owner and access rights. Open source data includes data lineage.
[0016] Usage data includes access logs and access frequency;
[0017] Data classification standards refer to the standards and rules that an enterprise clearly defines for different data classification levels, as well as the protection measures corresponding to the data classification levels.
[0018] Preferably, the method for preliminarily processing the comprehensive asset data comprises:
[0019] For numerical data in comprehensive asset data, use linear interpolation to fill missing values; use box plots or Z-scores to identify outliers, delete them, and treat them as missing values; and perform standardization and normalization;
[0020] And label mapping is performed on the processed numerical data, that is, a label is added to each value. The label consists of words and is used to express the value.
[0021] For text data in comprehensive asset data, special characters and punctuation marks are removed through text cleaning, and the format is unified and spaces and line breaks are removed; spelling check tools are used to correct spelling errors in the text, and syntax analysis tools are used to correct grammatical errors in the text, and missing words and sentences are filled in.
[0022] Preferably, the use of improved natural language processing technology to perform key feature analysis on each text in the preliminary comprehensive data includes:
[0023] For each text in the preliminary synthesis data;
[0024] Step B1: Use Jieba or HanLP word segmentation tools to segment the text into independent words or phrases and filter out stop words;
[0025] Step B2: Use vector-feature technology to perform feature processing on the text after word segmentation to obtain a text conversion vector;
[0026] Step B3: Build and train a CRF model, input the text conversion vector into the trained CRF model, obtain the keywords in each text, and form a keyword set;
[0027] Step B4: Use TF-Te technology to rank the keywords in the keyword set in terms of importance, and select the first F_P feature words to form a feature set.
[0028] Preferably, the method of using the vector-feature technology to perform feature processing on the segmented text to obtain the text conversion vector includes:
[0029] Step C1: Use the word vector model to represent each word as a vector, recorded as a word vector;
[0030] Step C2: For any word vector in any text, preset it as a basis vector, and select D_W adjacent word vectors before and after the basis vector;
[0031] ,in, Indicates the first The number of adjacent word vectors selected for each word corresponding to the word vector, and represents the adjustment factor, Indicates the number of non-word characters in the text, Represents the total number of word vectors corresponding to words in the text, Represents the number of categories of word vectors in the text;
[0032] Step C3: Use semantic technology to define the possible meanings of each word, and repeat step C1 to assign a word vector to each possible meaning, which is recorded as the meaning vector;
[0033] Step C4: Calculate the average vector of the basis vector and D_W adjacent word vectors. For each basis vector, use the cosine similarity to calculate the similarity between the average vector and each sense vector. Select the sense vector with the highest similarity as the semantic representation of the word, recorded as the word sense vector.
[0034] Step C5: For each word, concatenate the word vector and the word meaning vector in parallel to form a new vector, which is recorded as the word conversion vector;
[0035] Step C6: For each text, collect all word conversion vectors to form a text conversion vector.
[0036] Preferably, the step of constructing and training a CRF model, inputting the text conversion vector into the trained CRF model, and obtaining keywords in each text includes:
[0037] Construct a sample set: The sample set includes T_U groups of samples at historical time. Each group of samples includes the text transformation vector of each text and the corresponding text label set.
[0038] Use the BIOES scheme to manually annotate the keywords in each text, associate a label with each word, assign the label to the word conversion vector corresponding to the keyword, and collect the labels corresponding to all the word conversion vectors of the text, which are recorded as the text label set;
[0039] Among them, for the BIOES scheme, B-KEY represents the starting word of the keyword, I-KEY represents the middle word of the keyword, E-KEY represents the ending word of the keyword, S-KEY represents the keyword composed of a single word, and O represents a non-keyword;
[0040] Construct a CRF model: set u_h input layers and output layers in parallel. The input of each input layer is the text conversion vector of each text. The output of each output layer is the text label set corresponding to each text. The optimization goal is to maximize the conditional probability.
[0041] Training the CRF model: Import the sample set into the constructed CRF model in batches until the maximum number of iterations is reached, and then obtain the trained CRF model;
[0042] Use the CRF model: Extract the company's current preliminary comprehensive data, input the text conversion vector of each text in the preliminary comprehensive data into the trained CRF model in parallel, and obtain the keywords in each text.
[0043] Preferably, the use of TF-Te technology to rank the keywords in the keyword set by importance includes:
[0044] TF-Te technology is based on the combination of TF-IDF technology and TextRank technology;
[0045] Extract keywords from each text in the company's current preliminary comprehensive data;
[0046] Calculate the TF-IDF value of each keyword, set the TF-IDF threshold, and extract keywords with TF-IDF values higher than the TF-IDF threshold as candidate words;
[0047] The selected candidate words are used as nodes, and the nodes are connected by undirected edges. The product of the TF-IDF values between two candidate words is used as the weight of the corresponding undirected edge. The TextRank graph of the text is constructed based on the nodes and undirected edges.
[0048] Use the PageRank algorithm to iteratively calculate, obtain the TextRank value of each node, sort the TextRank values of the nodes in descending order, and select the candidate words corresponding to the first F_P TextRank value nodes as the feature words of the text.
[0049] Preferably, the method of performing cluster analysis on each text in the feature comprehensive data using K-means technology comprises:
[0050] For each text in the feature synthesis data;
[0051] Step E1: preset K_v value and randomly select K_v cluster centers;
[0052] Step E2: Calculate the distance between each text and each cluster center ,in, Indicates the first word conversion vectors, Indicates Cluster centers, Indicates The text and The distance between cluster centers is calculated, and the text is assigned to the cluster center with the closest distance, and the text in the feature comprehensive data is divided into K groups;
[0053] Step E3: Calculate the average of the text transformation vectors of all the texts in each group, and use the averaged text transformation vector to update the cluster center;
[0054] Step E4: Repeat steps E2 and E3 until all cluster centers no longer change, obtain the text divided into K_v groups, and then form a similarity classification data set.
[0055] Preferably, the grading scheme design model is based on a neural network learning technology, and includes an input layer, a hidden layer, and an output layer. The input of the input layer is set as a similar graded data set, and the output of the output layer is the data classification level corresponding to each group in the similar graded data set.
[0056] The mean square error function is used as the loss function, and the optimization goal is to minimize the loss function;
[0057] A training set is formed by using a similar hierarchical data set in historical time and a data classification level corresponding to each group, and a hierarchical scheme design model is trained using the training set to obtain a trained hierarchical scheme design model;
[0058] Input the current similar classification data set into the trained classification scheme design model to obtain the corresponding data classification level.
[0059] An enterprise data asset hierarchical management system, comprising:
[0060] Data collection module: used to collect comprehensive asset data and data classification standards of enterprises within the cycle time;
[0061] Data analysis and processing module: used to perform preliminary processing on comprehensive asset data, obtain preliminary comprehensive data, use improved natural language processing technology to perform key feature analysis on preliminary comprehensive data, obtain feature comprehensive data, use K-means technology to perform cluster analysis on feature comprehensive data, and obtain similar graded data sets;
[0062] Model design and output module: used to use the grading scheme design model to perform grading matching on similar grading data sets, obtain the data classification level corresponding to each group, and manage according to the data classification level corresponding to each group.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] Through data collection, preprocessing and key feature analysis, enterprises can more clearly understand the composition, characteristics and value of their own data assets, laying the foundation for subsequent data analysis and application.
[0065] By utilizing improved natural language processing technology and machine learning algorithms, we can deeply understand the data content, extract key features, and achieve more accurate data classification and grading, replacing a large number of manual operations and significantly improving data processing efficiency. Through automatic word segmentation and keyword extraction, we can achieve concise and accurate analysis of text, reduce manual intervention, and reduce the possibility of human error.
[0066] Based on the data classification results, corresponding protection strategies can be applied to data of different levels, avoiding the "one-size-fits-all" security management model and achieving reasonable allocation and effective utilization of resources. Hierarchical management enables enterprises to allocate security resources according to the importance of data, avoiding waste of resources and reducing security management costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 The present invention proposes a structural schematic diagram of an enterprise data asset hierarchical management system;
[0068] Figure 2 It is a schematic diagram of the method applied to the enterprise data asset hierarchical management system in the present invention. DETAILED DESCRIPTION
[0069] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.
[0070] Embodiment 1
[0071] Reference Figure 1 and Figure 2 Embodiment 1 further illustrates an enterprise data asset hierarchical management system proposed by the present invention.
[0072] Enterprise data assets are important data of the company. When hierarchical management is required, the problem of unclear data classification will arise, which will lead to a series of security and management challenges.
[0073] Manual classification systems are inefficient, inaccurate, and prone to errors, or data classification standards are vaguely or overly broadly defined, which may lead to the misclassification of sensitive data as non-sensitive, or vice versa, resulting in increased risk of data leakage or waste of resources on protecting unnecessary data. For example, labels such as "confidential" or "internal use" may not have clear standards to define which data should be classified into these categories. The sensitivity of data may change over time or with changes in the business environment, resulting in data that is no longer sensitive being overprotected, while newly generated sensitive data may not receive the protection it deserves.
[0074] In addition, in the process of data classification, different departments or teams may use different classification methods. There is no unified enterprise-level framework, which leads to confusion when data flows across departments and makes it difficult to implement consistent data protection measures.
[0075] By addressing these issues, organizations can more effectively manage and protect their data assets, reduce the risk of data breaches, and improve data availability and compliance.
[0076] Step S1: Collect the comprehensive asset data and data classification standards of the enterprise within the cycle time, and perform preliminary processing on the comprehensive asset data to obtain preliminary comprehensive data;
[0077] Comprehensive asset data includes basic asset data and metadata and usage data of basic asset data;
[0078] Basic asset data refers to the data itself that needs to be managed in a hierarchical manner. It is the most basic data and requires the actual data content to be obtained for content analysis and understanding. For example, the content of text files, records in databases, images or video files, etc.
[0079] The metadata of basic asset data includes attribute data and open source data. Attribute data includes file format, size, creation date, modification date, owner and access rights. Open source data includes data lineage (the source and destination of the data).
[0080] Usage data includes access logs and access frequency;
[0081] Data classification standards refer to the standards and rules that an enterprise clearly defines for different data classification levels, such as "confidential", "internal use", "public", etc., or 1, 2, 3, 4, 5, 6, etc., as well as the protection measures corresponding to the data classification levels, such as access control, encryption, data desensitization, etc.
[0082] Methods for initial processing of comprehensive asset data include:
[0083] For numerical data in comprehensive asset data, use linear interpolation to fill missing values; use box plots or Z-scores to identify outliers, delete them, and treat them as missing values; and perform standardization and normalization;
[0084] And label mapping is performed for the processed numerical data, that is, adding a label to each numerical value. The label is composed of words and is used to express the numerical value. For example, for a numerical value 37, which represents the access frequency, the access frequency is used as the label of 37, so that when analyzing the data, text analysis is uniformly used to reduce processing obstacles and improve analysis efficiency.
[0085] For text data in comprehensive asset data, special characters and punctuation marks are removed through text cleaning, and the format is unified and spaces and line breaks are removed; spelling check tools are used to correct spelling errors in the text, and syntax analysis tools are used to correct grammatical errors in the text, and missing words and sentences are filled in.
[0086] Among them, removing special characters and punctuation marks is to use regular expressions or string operations to remove special characters, punctuation marks, HTML tags, etc. in the text, and only retain meaningful text content. Case conversion is to convert all text to lowercase or uppercase to maintain consistency and avoid vocabulary recognition problems caused by case differences. Removing spaces and line breaks is to remove extra spaces, tabs, and line breaks in the text to ensure that the text format is standardized.
[0087] Use spelling checkers to correct spelling errors in text. Common methods include dictionary-based methods (comparing words with dictionaries to identify spelling errors), rule-based methods (identifying spelling errors based on language rules, such as repeated letters, vowel errors, etc.), and statistical-based methods (using language models and probabilistic statistical methods to identify and correct spelling errors).
[0088] Use grammatical analysis tools to correct grammatical errors in text. Common methods include rule-based grammatical analysis (using grammatical rules and part-of-speech tagging to identify grammatical errors), statistical-based grammatical analysis (using probability models and machine learning methods to identify grammatical errors), and using pre-trained language models (using BERT, GPT and other pre-trained models to detect and correct grammatical errors).
[0089] To fill in missing words and sentences, common methods include context inference (inferring missing words and sentences based on context information. For example, a language model can be used to predict missing words or phrases), using placeholders (using specific placeholders (such as <unk>or <missing>) indicates missing words and sentences, allowing the model to learn how to deal with missing values), and if the missing words and sentences have little effect on understanding the text, you can consider deleting or ignoring them.
[0090] The goal of data preprocessing is to convert raw data into a format suitable for model training and to improve the quality of the data.
[0091] Step S2: Use improved natural language processing technology to perform key feature analysis on each text in the preliminary comprehensive data, collect word conversion vectors corresponding to feature words in all texts, and obtain feature comprehensive data;
[0092] Improved natural language processing techniques are used to analyze key features of each text in the preliminary comprehensive data, including:
[0093] For each text in the preliminary synthesis data;
[0094] Step B1: Use Jieba or HanLP word segmentation tools to segment the text into independent words or phrases and filter out stop words;
[0095] Step B2: Use vector-feature technology to perform feature processing on the text after word segmentation to obtain a text conversion vector;
[0096] Step B3: Build and train a CRF model, input the text conversion vector into the trained CRF model, obtain the keywords in each text, and form a keyword set;
[0097] Step B4: Use TF-Te technology to rank the keywords in the keyword set in terms of importance, and select the first F_P feature words to form a feature set.
[0098] Use vector-feature technology to perform feature processing on the segmented text to obtain the text conversion vector, including:
[0099] Step C1: Use a word vector model (such as Word2Vec, GloVe or FastText) to represent each word as a vector, recorded as a word vector;
[0100] In actual text, you may encounter unregistered words (OOV, Out-of-Vocabulary). You can assign random vectors to unregistered words, or use the FastText method to generate word vectors based on characters or subwords. You can also use the average of all word vectors in the corpus as the vector of the unregistered word.
[0101] Since a single word vector cannot fully represent the semantic information of a word in a sentence, it can be modeled using a context window.
[0102] Step C2: For any word vector in any text, preset it as a basis vector, and select D_W adjacent word vectors before and after the basis vector, that is, The previous neighbor word vectors and The vectors of the next adjacent words;
[0103] ,in, Indicates the first The number of adjacent word vectors selected for each word corresponding to the word vector, and represents the adjustment factor, Indicates the number of non-word characters in the text, Represents the total number of word vectors corresponding to words in the text, Indicates the number of categories of word vectors in the text, that is, word vectors with the same expression are classified into one category, and based on this, the types of word vectors with different expressions in the text are identified;
[0104] Many words have multiple meanings (e.g. "apple" can mean a fruit or a company). The goal of word sense disambiguation is to select the correct meaning based on the context.
[0105] Step C3: Use semantic technology (such as WordNet, HowNet) to define the possible meanings of each word, and repeat step C1 to assign a word vector to each possible meaning, which is recorded as the meaning vector;
[0106] Step C4: Calculate the average vector of the basis vector and D_W adjacent word vectors. ,in, Indicates word vectors, and for each basis vector, the similarity between the average vector and each sense vector calculated using cosine similarity, ,in, represents the basis vector sense vectors, select the sense vector with the highest similarity as the semantic representation of the word, and record it as the word sense vector;
[0107] Step C5: For each word, the word vector and the word sense vector are concatenated in parallel to form a new vector, which is recorded as the word conversion vector. The disambiguated sense vector is used to accurately represent the meaning of the word.
[0108] Step C6: For each text, collect all word conversion vectors to form a text conversion vector.
[0109] Build and train the CRF model, input the text conversion vector into the trained CRF model, and obtain the keywords in each text, including:
[0110] Construct a sample set: The sample set includes T_U groups of samples at historical time. Each group of samples includes the text transformation vector of each text and the corresponding text label set.
[0111] Use the BIOES scheme to manually annotate the keywords in each text, associate a label with each word, assign the label to the word conversion vector corresponding to the keyword, and collect the labels corresponding to all the word conversion vectors of the text, which are recorded as the text label set;
[0112] Among them, for the BIOES scheme, B-KEY represents the starting word of the keyword, I-KEY represents the middle word of the keyword, E-KEY represents the ending word of the keyword, S-KEY represents the keyword composed of a single word, and O represents a non-keyword;
[0113] Construct a CRF model: set u_h input layers and output layers in parallel. The input of each input layer is the text conversion vector of each text. The output of each output layer is the text label set corresponding to each text. The optimization goal is to maximize the conditional probability.
[0114] Training the CRF model: Import the sample set into the constructed CRF model in batches until the maximum number of iterations is reached, and then obtain the trained CRF model;
[0115] Use the CRF model: Extract the company's current preliminary comprehensive data, input the text conversion vector of each text in the preliminary comprehensive data into the trained CRF model in parallel, and obtain the keywords in each text.
[0116] Use TF-Te technology to rank the keywords in the keyword set in terms of importance, including:
[0117] TF-Te technology is based on the combination of TF-IDF technology and TextRank technology;
[0118] Extract keywords from each text in the company's current preliminary comprehensive data;
[0119] Calculate the TF-IDF value of each keyword, set the TF-IDF threshold, and extract keywords with TF-IDF values higher than the TF-IDF threshold as candidate words;
[0120] The TF-IDF value is the product of the TF value and the IDF value. TF (term frequency) is obtained by calculating the frequency of each word in the document, and IDF (inverse document frequency) is obtained by calculating the inverse document frequency of each word in the entire corpus.
[0121] The selected candidate words are used as nodes, and the nodes are connected by undirected edges. The product of the TF-IDF values between two candidate words is used as the weight of the corresponding undirected edge to reflect the importance of the words. The TextRank graph of the text is constructed based on the nodes and undirected edges.
[0122] Use the PageRank algorithm to iteratively calculate, obtain the TextRank value of each node, sort the TextRank values of the nodes in descending order, and select the candidate words corresponding to the first F_P TextRank value nodes as the feature words of the text.
[0123] Both F_P and TF-IDF thresholds can be set based on experimental data analysis or historical data analysis.
[0124] When using the PageRank algorithm for iterative calculation, the TextRank value of all nodes in the graph is initialized to 1. According to the iterative formula, the TextRank value of each node is iteratively calculated until convergence or the maximum number of iterations is reached, and the final TextRank value of each node is output.
[0125] The convergence condition is that the change of the TextRank value of all nodes is less than the preset threshold, or the number of iterations reaches the maximum limit.
[0126] Combining TF-IDF and TextRank is an effective text analysis method that can fully utilize their respective advantages to improve the accuracy and efficiency of keyword extraction and obtain text analysis results that better meet the needs.
[0127] Step S3: Use K-means technology to perform cluster analysis on each text in the feature comprehensive data to obtain a similar graded data set;
[0128] The method of using K-means technology to perform cluster analysis on each text in the feature comprehensive data includes:
[0129] For each text in the feature synthesis data;
[0130] Step E1: preset K_v value and randomly select K_v cluster centers;
[0131] Step E2: Calculate the distance between each text and each cluster center ,in, Indicates the first word conversion vectors, Indicates Cluster centers, Indicates The text and The distance between cluster centers is calculated, and the text is assigned to the cluster center with the closest distance, and the text in the feature comprehensive data is divided into K groups;
[0132] Step E3: Calculate the average of the text transformation vectors of all the texts in each group, and use the averaged text transformation vector to update the cluster center;
[0133] Step E4: Repeat steps E2 and E3 until all cluster centers no longer change, obtain the text divided into K_v groups, and then form a similarity classification data set.
[0134] The value of K_v is the same as the number of levels of data classification.
[0135] Step S4: Use the grading scheme design model to perform grade matching on each group in the similar graded data set to obtain the data classification level corresponding to each group;
[0136] The grading scheme design model is based on neural network learning technology, including input layer, hidden layer and output layer. The input of the input layer is set as a similar grading data set, and the output of the output layer is the data classification level corresponding to each group in the similar grading data set.
[0137] The mean square error function is used as the loss function, and the optimization goal is to minimize the loss function;
[0138] A training set is formed by using a similar hierarchical data set in historical time and a data classification level corresponding to each group, and a hierarchical scheme design model is trained using the training set to obtain a trained hierarchical scheme design model;
[0139] Input the current similar classification data set into the trained classification scheme design model to obtain the corresponding data classification level.
[0140] Embodiment 2
[0141] Reference Figure 1 and Figure 2 , Example 2 further illustrates an enterprise data asset hierarchical management system proposed by the present invention.
[0142] A method for hierarchical management of enterprise data assets comprises the following steps:
[0143] Step S1: Collect the comprehensive asset data and data classification standards of the enterprise within the cycle time, and perform preliminary processing on the comprehensive asset data to obtain preliminary comprehensive data;
[0144] Step S2: Use improved natural language processing technology to perform key feature analysis on each text in the preliminary comprehensive data, collect word conversion vectors corresponding to feature words in all texts, and obtain feature comprehensive data;
[0145] Step S3: Use K-means technology to perform cluster analysis on each text in the feature comprehensive data to obtain a similar graded data set;
[0146] Step S4: Use the grading scheme design model to perform grade matching on each group in the similar graded data set to obtain the data classification level corresponding to each group;
[0147] Step S5: Manage the data using the protection measures in the data classification standard according to the data classification level corresponding to each group.
[0148] An enterprise data asset hierarchical management system, used to implement an enterprise data asset hierarchical management method, includes:
[0149] Data collection module: used to collect comprehensive asset data and data classification standards of enterprises within the cycle time;
[0150] Data analysis and processing module: used to perform preliminary processing on comprehensive asset data, obtain preliminary comprehensive data, use improved natural language processing technology to perform key feature analysis on preliminary comprehensive data, obtain feature comprehensive data, use K-means technology to perform cluster analysis on feature comprehensive data, and obtain similar graded data sets;
[0151] Model design and output module: used to use the grading scheme design model to perform grading matching on similar grading data sets, obtain the data classification level corresponding to each group, and manage according to the data classification level corresponding to each group;
[0152] The modules are connected to each other via wired and / or wireless means.
[0153] In addition, according to the implementation of the present application, the process described in the figure of an enterprise data asset hierarchical management system can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium, which stores machine-readable instructions, and the machine-readable instructions can be executed by a processor to execute instructions corresponding to the method steps provided by the present application. Of course, the architecture shown in the figure of an enterprise data asset hierarchical management system is only exemplary. When implementing different devices, adaptive selection or adjustment can be made according to actual needs.
[0154] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.
[0155] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.< / missing> < / unk>
Claims
1. A method for hierarchical management of enterprise data assets, characterized in that: The enterprise data asset hierarchical management method comprises: Step S1: Collect the comprehensive asset data and data classification standards of the enterprise within the cycle time, and perform preliminary processing on the comprehensive asset data to obtain preliminary comprehensive data; The comprehensive asset data includes basic asset data and metadata and usage data of the basic asset data; Basic asset data refers to the data itself that needs to be managed hierarchically; The metadata of basic asset data includes attribute data and open source data. Attribute data includes file format, size, creation date, modification date, owner and access rights. Open source data includes data lineage. Usage data includes access logs and access frequency; Data classification standards refer to the standards and rules that enterprises clearly define for different data classification levels, as well as the corresponding protection measures for data classification levels; The method for preliminarily processing the comprehensive asset data comprises: For numerical data in comprehensive asset data, use linear interpolation to fill missing values; use box plots or Z-scores to identify outliers, delete them, and treat them as missing values; and perform standardization and normalization; And label mapping is performed on the processed numerical data, that is, a label is added to each value. The label consists of words and is used to express the value. For text data in the comprehensive asset data, special characters and punctuation marks are removed through text cleaning, and the format is unified and spaces and line breaks are removed; spelling errors in the text are corrected using spelling check tools, grammatical errors in the text are corrected using grammar analysis tools, and missing words and sentences are filled in; Step S2: Use improved natural language processing technology to perform key feature analysis on each text in the preliminary comprehensive data, collect word conversion vectors corresponding to feature words in all texts, and obtain feature comprehensive data; The improved natural language processing technology is used to analyze key features of each text in the preliminary comprehensive data, including: For each text in the preliminary synthesis data; Step B1: Use Jieba or HanLP word segmentation tools to segment the text into independent words or phrases and filter out stop words; Step B2: Use vector-feature technology to perform feature processing on the text after word segmentation to obtain a text conversion vector; Step B3: Build and train a CRF model, input the text conversion vector into the trained CRF model, obtain the keywords in each text, and form a keyword set; Step B4: Use TF-Te technology to rank the keywords in the keyword set in terms of importance, and select the first F_P feature words to form a feature set; Step S3: Use K-means technology to perform cluster analysis on each text in the feature comprehensive data to obtain a similar graded data set; Step S4: Use the grading scheme design model to perform grade matching on each group in the similar graded data set to obtain the data classification level corresponding to each group; Step S5: Manage the data using the protection measures in the data classification standard according to the data classification level corresponding to each group.
2. The enterprise data asset hierarchical management method according to claim 1, characterized in that: The method of using the vector-feature technology to perform feature processing on the segmented text to obtain the text conversion vector includes: Step C1: Use the word vector model to represent each word as a vector, recorded as a word vector; Step C2: For any word vector in any text, preset it as a basis vector, and select D_W adjacent word vectors before and after the basis vector; Among them, D_W i represents the number of adjacent word vectors selected for the word vector corresponding to the i-th word in the text, α1 and α2 represent adjustment coefficients, Fw represents the number of non-word characters in the text, CY_z represents the total number of word vectors corresponding to the words in the text, and CY_j represents the number of categories of word vectors in the text; Step C3: Use semantic technology to define the possible meanings of each word, and repeat step C1 to assign a word vector to each possible meaning, which is recorded as the meaning vector; Step C4: Calculate the average vector of the basis vector and D_W adjacent word vectors. For each basis vector, use the cosine similarity to calculate the similarity between the average vector and each sense vector. Select the sense vector with the highest similarity as the semantic representation of the word, recorded as the word sense vector. Step C5: For each word, concatenate the word vector and the word meaning vector in parallel to form a new vector, which is recorded as the word conversion vector; Step C6: For each text, collect all word conversion vectors to form a text conversion vector.
3. The enterprise data asset hierarchical management method according to claim 2, characterized in that: The CRF model is constructed and trained, and the text conversion vector is input into the trained CRF model to obtain the keywords in each text, including: Construct a sample set: The sample set includes T_U groups of samples at historical time. Each group of samples includes the text transformation vector of each text and the corresponding text label set. Use the BIOES scheme to manually annotate the keywords in each text, associate a label with each word, assign the label to the word conversion vector corresponding to the keyword, and collect the labels corresponding to all the word conversion vectors of the text, which are recorded as the text label set; Among them, for the BIOES scheme, B-KEY represents the starting word of the keyword, I-KEY represents the middle word of the keyword, E-KEY represents the ending word of the keyword, S-KEY represents the keyword composed of a single word, and O represents a non-keyword; Construct a CRF model: set u_h input layers and output layers in parallel. The input of each input layer is the text conversion vector of each text. The output of each output layer is the text label set corresponding to each text. The optimization goal is to maximize the conditional probability. Training the CRF model: Import the sample set into the constructed CRF model in batches until the maximum number of iterations is reached, and then obtain the trained CRF model; Use the CRF model: Extract the company's current preliminary comprehensive data, input the text conversion vector of each text in the preliminary comprehensive data into the trained CRF model in parallel, and obtain the keywords in each text.
4. The enterprise data asset hierarchical management method according to claim 3, characterized in that: The use of TF-Te technology to rank the keywords in the keyword set by importance includes: TF-Te technology is based on the combination of TF-IDF technology and TextRank technology; Extract keywords from each text in the company's current preliminary comprehensive data; Calculate the TF-IDF value of each keyword, set the TF-IDF threshold, and extract keywords with TF-IDF values higher than the TF-IDF threshold as candidate words; The selected candidate words are used as nodes, and the nodes are connected by undirected edges. The product of the TF-IDF values between two candidate words is used as the weight of the corresponding undirected edge. The TextRank graph of the text is constructed based on the nodes and undirected edges. Use the PageRank algorithm to iteratively calculate, obtain the TextRank value of each node, sort the TextRank values of the nodes in descending order, and select the candidate words corresponding to the first F_P TextRank value nodes as the feature words of the text.
5. The enterprise data asset hierarchical management method according to claim 4, characterized in that: The method of using K-means technology to perform cluster analysis on each text in the feature comprehensive data includes: For each text in the feature synthesis data; Step E1: preset K_v value and randomly select K_v cluster centers; Step E2: Calculate the distance between each text and each cluster center in, represents the hth word conversion vector in the text, represents the g-th cluster center, JL(v,g) represents the distance between the v-th text and the g-th cluster center, and the text is assigned to the cluster center with the nearest distance, and the text in the feature comprehensive data is divided into K groups; Step E3: Calculate the average of the text transformation vectors of all the texts in each group, and use the averaged text transformation vector to update the cluster center; Step E4: Repeat steps E2 and E3 until all cluster centers no longer change, obtain the text divided into K_v groups, and then form a similarity classification data set.
6. The enterprise data asset hierarchical management method according to claim 5, characterized in that: The grading scheme design model is based on neural network learning technology, including an input layer, a hidden layer and an output layer. The input of the input layer is set as a similar graded data set, and the output of the output layer is the data classification level corresponding to each group in the similar graded data set. The mean square error function is used as the loss function, and the optimization goal is to minimize the loss function; A training set is formed by using a similar hierarchical data set in historical time and a data classification level corresponding to each group, and a hierarchical scheme design model is trained using the training set to obtain a trained hierarchical scheme design model; Input the current similar classification data set into the trained classification scheme design model to obtain the corresponding data classification level.
7. An enterprise data asset hierarchical management system, used to implement the enterprise data asset hierarchical management method according to any one of claims 1 to 6, characterized in that: The enterprise data asset hierarchical management system includes: Data collection module: used to collect comprehensive asset data and data classification standards of enterprises within the cycle time; Data analysis and processing module: used to perform preliminary processing on comprehensive asset data, obtain preliminary comprehensive data, use improved natural language processing technology to perform key feature analysis on preliminary comprehensive data, obtain feature comprehensive data, use K-means technology to perform cluster analysis on feature comprehensive data, and obtain similar graded data sets; Model design and output module: used to use the grading scheme design model to perform grading matching on similar grading data sets, obtain the data classification level corresponding to each group, and manage according to the data classification level corresponding to each group.
Citation Information
Patent Citations
Method and device for automatically acquiring multi-level classification training data of enterprise
CN112287075A
Data security management and control system based on data classification and grading method
CN119167422A