A data classification and grading method, system, device and medium based on big data

By automatically processing data characteristics and sensitivity evaluation based on big data, generating hierarchical index tables and assigning access paths, the problem of inefficient data classification and grading in the prior art is solved, and efficient and accurate data grading management is achieved.

CN119441564BActive Publication Date: 2025-06-17BEIJING WISEDA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510044543.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-11
Publication Date
2025-06-17
Estimated Expiration
2045-01-11

AI Technical Summary

Technical Problem

Existing data classification and grading methods rely on manual analysis, which makes it inefficient in large-scale enterprise data and is difficult to quickly process and gradle large amounts of data.

Method used

Using a big data-based method, by obtaining the data to be hierarched by multiple data nodes and the benchmark index table of the central management node, extracting the content characteristics and business attributes of the data, generating data feature identifiers, and using the preset sensitive thesaurus and data evaluation matrix for automated evaluation, generating a hierarchical index table and assigning access paths.

Benefits of technology

The rapid processing of large-scale data sets is achieved, reducing the process of human judgment, improving the accuracy and efficiency of grading, and avoiding the inefficiency problem in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441564B_ABST
    Figure CN119441564B_ABST
Patent Text Reader

Abstract

A data classification and grading method, system, device and medium based on big data, which relates to the technical field of data processing. The method includes: obtaining the data to be graded of multiple data nodes and the benchmark index table of the central management node; extracting the content features and business attributes of each data to be graded, and generating corresponding data feature identifiers according to each content feature and each business attribute; scanning each data feature identifier with a preset sensitive word library to generate corresponding data sensitivity scores, and evaluating each data sensitivity score through a preset data evaluation matrix to obtain corresponding data level information; associating the benchmark index table of the central management node based on the data feature identifiers and data level information of each data node to obtain a graded index table, and allocating corresponding access paths to each data node in the graded index table. Implementing the technical solution provided by this application can improve the efficiency of data classification and grading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data classification and grading method, system, device and medium based on big data. Background Art

[0002] In the digital age, data has become an important asset. As enterprises increasingly rely on the collection, processing and storage of a large amount of data, data security and compliance issues have also increased. All kinds of data, whether personal information, financial records or trade secrets, need to be properly managed and protected to prevent data leakage, abuse or other forms of security threats.

[0003] Currently, the existing data classification and grading methods mainly rely on collecting the operation data of enterprises through documents, databases, etc., and manually analyzing these data by the staff related to data governance to achieve the purpose of data classification and grading. However, in practical applications, due to the large scale of the operation data of enterprises, only by manually classifying and grading data often requires a large amount of processing time, thus reducing the efficiency of data classification and grading. Summary of the Invention

[0004] This application provides a data classification and grading method, system, device and medium based on big data, which can improve the efficiency of data classification and grading.

[0005] In the first aspect, this application provides a data classification and grading method based on big data, including:

[0006] Obtain the data to be graded of multiple data nodes and the benchmark index table of the central management node;

[0007] Extract the content features and business attributes of each piece of data to be graded, and generate corresponding data feature identifiers according to each content feature and each business attribute;

[0008] Scan and process each data feature identifier with a preset sensitive word library to generate corresponding data sensitivity scores, and evaluate each data sensitivity score through a preset data evaluation matrix to obtain corresponding data level information;

[0009] Associate the benchmark index table of the central management node based on the data feature identifiers and data level information of each data node to obtain a graded index table, and allocate corresponding access paths to each data node in the graded index table.

[0010] In the second aspect of this application, a data classification and grading system based on big data is provided. The system includes:

[0011] A data acquisition module, configured to acquire the data to be classified of multiple data nodes and the benchmark index table of the central management node;

[0012] A feature identification determination module, configured to extract the content features and service attributes of each piece of the data to be classified, and generate corresponding data feature identifiers according to each of the content features and each of the service attributes;

[0013] A level information determination module, configured to perform a scanning process on each of the data feature identifiers by using a preset sensitive word library, generate corresponding data sensitivity scores, and evaluate each of the data sensitivity scores through a preset data evaluation matrix to obtain corresponding data level information;

[0014] A data classification module, configured to associate the benchmark index table of the central management node based on the data feature identifiers and data level information of each of the data nodes, obtain a classified index table, and allocate corresponding access paths to each of the data nodes in the classified index table.

[0015] In a third aspect of the present application, an electronic device is provided, including a memory, a processor, and a program stored on the memory and executable on the processor, and the program can implement a data classification and grading method based on big data when being loaded and executed by the processor.

[0016] In a fourth aspect of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements a data classification and grading method based on big data.

[0017] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0018] By adopting the above technical solutions, first, the data to be classified of multiple data nodes and the benchmark index table of the central management node are acquired, establishing a unified data management foundation; then, by extracting the content features and service attributes of the data to be classified to generate data feature identifiers, the standardized expression of data features is realized; afterwards, automatic scanning is performed by using a preset sensitive word library and combined with a data evaluation matrix for evaluation to obtain objective data level information, ensuring the accuracy and consistency of the classification results. Finally, the method forms a classified index table by associating the data feature identifiers and data level information with the benchmark index table, and allocates corresponding access paths to each data node, constructing a complete data classification and management system. Compared with the existing method of manually classifying and grading data, this classification method based on automated processing can quickly process large-scale data sets, reduce the process of manual judgment, improve the classification accuracy, and avoid the problem of low efficiency in the traditional data classification and grading process. Description of the Drawings

[0019] Figure 1 is a schematic flowchart of a data classification and grading method based on big data provided by an embodiment of the present application;

[0020] Figure 2 is a schematic structural diagram of a data classification and grading system based on big data provided by an embodiment of the present application;

[0021] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0022] Description of the reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Embodiments

[0023] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0024] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0025] In the description of the embodiments of the present application, the meaning of the term "a plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0026] The embodiments of the present application provide a data classification and grading method based on big data. In one embodiment, please refer to Figure 1 , Figure 1FIG. 0 is a schematic flowchart of a data classification and grading method based on big data provided by an embodiment of the present application. This method can be implemented relying on a computer program, which can be integrated into an application or run as an independent tool - type application. This method can also be implemented relying on a single - chip microcomputer and can run on a big - data - based data classification and grading system based on the von Neumann architecture. Specifically, this method may include the following steps:

[0027] Step 101: Obtain the data to be graded of multiple data nodes and the reference index table of the central management node.

[0028] Among them, a data node refers to a data storage and processing unit distributed in various business systems of an enterprise, which can be different types of data sources, such as business database servers, file storage servers, application system servers, etc. Each data node independently manages and stores data resources in a specific business area, and they together constitute the enterprise's distributed data storage architecture.

[0029] The data to be graded refers to the original data stored in each data node that needs to be classified and graded, including but not limited to: business data generated during the operation of the enterprise (such as customer information, transaction records), management data (such as personnel files, financial statements), technical data (such as R & D documents, design drawings), etc. These data may exist in different formats, including both structured data and unstructured data.

[0030] The central management node refers to the core control unit of the enterprise data classification and grading system, which is responsible for uniformly managing and coordinating the entire classification and grading process. It maintains enterprise - level data classification and grading policies, security rules, and access control policies, and is responsible for communicating and exchanging data with each data node. The central management node undertakes core functions such as formulating data grading standards, storing grading results, and managing access permissions.

[0031] The reference index table is the basic data structure stored in the central management node, which defines the basic framework and rule system for data classification and grading. The reference index table contains information such as the hierarchical structure of data classification, the definition of basic attributes of various types of data, grading standards, sensitivity assessment rules, etc. It serves as a reference standard for data classification and grading, guiding the system to extract features and evaluate the levels of the data to be graded.

[0032] Specifically, it is first necessary to obtain the data to be classified from multiple data nodes and the reference index table of the central management node. Since enterprise data is usually distributed in different business systems and storage locations, a centralized data collection mechanism needs to be established to achieve unified data classification and grading management. The system connects the data nodes of each business system by configuring data collection interfaces. These data nodes can be file servers, database servers, or other data storage devices. The system adopts a distributed collection method to obtain the data to be classified from multiple data nodes simultaneously. This data may include structured data (such as database records) and unstructured data (such as documents, pictures, etc.). While obtaining the data to be classified, the system also needs to obtain the reference index table from the central management node. The reference index table is an important reference basis for data classification and grading, which contains the basic rules of data classification, index structure, and key attribute definitions. The reference index table is usually pre-developed by the data governance team and stored in the central management node. The system reads the reference index table from the central management node through a secure data transmission channel to ensure the integrity and consistency of the index information. This centralized data acquisition method can ensure that all data to be classified can be uniformly managed and evaluated. By establishing a connection between the data nodes and the central management node, the system can obtain the latest data content and index rules in real time, providing a reliable data basis for subsequent data feature extraction and grading evaluation.

[0033] Step 102: Extract the content features and business attributes of each piece of data to be classified, and generate corresponding data feature identifiers according to each content feature and each business attribute.

[0034] Among them, the content feature refers to a preset number of high-frequency keywords selected from the keyword set obtained after processing the data to be classified by text tokenization. The system will count the occurrence frequency of keywords and regard the keywords with an occurrence frequency greater than the frequency threshold as high-frequency keywords. These high-frequency keywords directly reflect the core content and theme information of the data to be classified, and can accurately represent the specific content characteristics of the data.

[0035] The business attribute refers to the business feature determined by matching the basic information such as the creation time, business scope, and using department of the data to be classified with a preset business type rule template. The preset business type rule template contains various combinations of different creation times, different business scopes, and different using departments and their corresponding business attribute definitions. These business attributes reflect the position and importance of the data in the enterprise business process.

[0036] A data feature identifier refers to a standardized feature vector generated by vectorizing content features and business attributes through feature vectorization. This feature vector combines the content features reflected by high-frequency keywords and the business attributes determined by business type rules, forming a unified numerical feature representation. This way of representing features through vectorization enables the system to quantitatively analyze and compare different types of data using a unified standard.

[0037] Specifically, after obtaining the data to be classified, it is necessary to extract features from the data to identify and evaluate its importance and sensitivity. The system first performs content feature extraction on the data to be classified through text tokenization, splitting the text content into a set of keywords. For these keywords, the system counts the occurrence frequency of each keyword and selects the keywords whose occurrence frequency exceeds a preset frequency threshold as high-frequency keywords. These high-frequency keywords can effectively reflect the core content and theme features of the data, and the system selects a preset number of high-frequency keywords as the content features of the data to be classified. At the same time, the system also needs to extract the business attribute information of the data to be classified. The system obtains basic information such as the creation time, business scope, and using department of the data to be classified, and matches this information with a preset business type rule template. The business type rule template pre-defines the business attributes corresponding to different combinations of creation time, business scope, and using department. Through this matching, the system can accurately identify the business attributes of the data to be classified, thereby understanding the importance and usage scenarios of the data in the business process.

[0038] After obtaining the content features and business attributes, the system needs to convert this information into a standardized data feature identifier. The system uses a feature vectorization processing method to convert text features and business attributes into numerical feature vectors. The feature vectorization process takes into account factors such as the importance of high-frequency keywords and the weights of business attributes, and integrates these features into a unified data feature identifier through a mathematical model. This feature vectorization processing enables different types of data to be compared and evaluated under the same standard.

[0039] Based on the above embodiments, as an optional embodiment, in step 102: extracting the content features and business attributes of each data to be classified, and generating corresponding data feature identifiers according to each content feature and each business attribute, this step may further include the following steps:

[0040] Step 201: For each data to be classified, perform text tokenization on the data to be classified to obtain a set of keywords.

[0041] Specifically, the system first determines the format type of the data to be classified. For structured data, the system directly extracts the content of its text fields; for unstructured data such as documents and emails, the system needs to first perform format conversion and extract the text content therein. After obtaining the text content, the system uses a preset word segmentation dictionary to perform word segmentation on the text, splitting the continuous text content into a sequence of words. During the word segmentation process, the system filters out stop words (such as function words like "de", "le", etc.) and special characters, and only retains the words with actual semantic meaning. At the same time, the system also performs part-of-speech tagging on the word segmentation results to identify different types of words such as nouns and verbs, and this information helps to judge the importance of the words in the subsequent process. After the text word segmentation process, the system obtains a keyword set composed of words with independent semantic meanings. Each word in this keyword set carries its position information and part-of-speech information in the original text, providing basic data for the subsequent extraction of high-frequency keywords. Compared with the original text content, the keyword set is more structured, facilitating the system to perform statistical analysis and feature extraction.

[0042] Step 202: Count the occurrence frequencies of each keyword in the keyword set, and select a preset number of high-frequency keywords as content features. The high-frequency keywords are the keywords whose occurrence frequencies are greater than the frequency threshold.

[0043] Specifically, after obtaining the keyword set, the system needs to identify the keyword that best represents the core content of the data to be classified as the content feature. Since the keyword set may contain a large number of words, and the importance of these words in expressing the text content is not the same, it is necessary to find the most representative high-frequency keywords through statistical analysis. The system first performs frequency statistics on each keyword in the keyword set, calculating the number of times each keyword appears in the text. To eliminate the influence of text length on frequency statistics, the system divides the original occurrence times by the total number of words in the text to obtain the standardized occurrence frequency. During the statistical process, the system will consider the part-of-speech information of the words, assigning higher weights to content words such as nouns and proper nouns, and lower weights to general verbs, adjectives, etc., so as to highlight the importance of the keywords expressing the core content. After the statistics are completed, the system sorts all the keywords in descending order of occurrence frequency. The system pre-sets a frequency threshold, and only the keywords whose occurrence frequency exceeds this threshold will be regarded as high-frequency keywords. The setting of this frequency threshold needs to consider the length and domain characteristics of the text, and is usually determined through data analysis and practical experience. From the high-frequency keywords, the system further selects a preset number of keywords as content features. The determination of the preset number needs to balance the representativeness of the features and the computational efficiency, ensuring that the selected keywords can fully express the core content of the text while avoiding the reduction of subsequent processing efficiency due to too high feature dimensions. The content features selected in this way have good representativeness and discrimination, and can accurately reflect the core content of the data to be classified. These features provide a content basis for subsequent data classification and help the system accurately evaluate the sensitivity and importance of the data.

[0044] Step 203: Obtain the creation time, business scope, and using department of the data to be classified, and determine the business attribute based on a preset business type rule template. The preset business type rule template includes business attributes corresponding to various combinations of different creation times, different business scopes, and different using departments; perform feature vectorization processing on the content feature and the business attribute to generate a data feature identifier.

[0045] Specifically, the system first extracts basic information such as creation time, business scope, and using department from the metadata information of the data to be classified. The creation time reflects the timeliness of the data, and the system converts it into a standard time format; the business scope indicates the specific business area to which the data belongs, such as finance, human resources, R & D, etc.; the using department indicates the main users and managers of the data. The system matches this information with a preset business type rule template, in which various combinations of different time periods, different business scopes, and different using departments are pre-defined with corresponding business attribute definitions. For example, the budget data of the finance department created within the last month may be defined as "highly sensitive real-time financial information", while the daily notice of the general office department two years ago may be defined as "ordinary historical information". After determining the business attributes, the system needs to uniformly convert the content features and business attributes into numerical feature vectors. For the high-frequency keywords in the content features, the system uses a word vector model to convert each keyword into a numerical vector of a fixed dimension, and calculates the weighted average in combination with the word frequency information to obtain a vector representing the content features. For the business attributes, the system converts them into vector representations in a preset encoded format. These two parts of vectors are normalized and then concatenated according to the preset dimensions and weights to form the final data feature identifier. In this way, the system can more accurately evaluate the sensitivity of the data and provide a basis for formulating corresponding protection strategies.

[0046] Step 103: Scan and process each data feature identifier using a preset sensitive word library to generate corresponding data sensitivity scores, and evaluate each data sensitivity score through a preset data evaluation matrix to obtain corresponding data level information.

[0047] Among them, the preset sensitive word library refers to a systematically organized collection of sensitive word data, which contains sensitive word entries stored classified by different business fields and information types. Each sensitive word entry consists of three key elements: sensitive word text (specific sensitive words or phrases), sensitive word weight (a numerical value reflecting the sensitivity of the word), and sensitive word category (such as category identifiers for trade secrets, personal information, financial data, etc.). This sensitive word library supports dynamic update and maintenance, and new sensitive words can be added or the attributes of existing sensitive words can be adjusted according to enterprise requirements.

[0048] The data sensitivity score refers to a numerical indicator obtained by scanning and calculating sensitive words in the text content. It consists of two levels of calculations: first is the sub-sensitivity score, which is the product of the sensitive word weight of a single sensitive word and its occurrence times in the text; then is the comprehensive sensitivity score, which is the final score obtained by weighted summing of each sub-sensitivity score according to the category coefficient of its belonging category. This score reflects the overall sensitivity of the data.

[0049] The preset data evaluation matrix refers to a mapping table that defines the correspondence between sensitivity scores and data levels. It contains the definitions of multiple score intervals (such as [0 - 30], [31 - 70], [71 - 100], etc.) and the data level information corresponding to each interval. This matrix serves as an evaluation criterion for converting numerical sensitivity scores into specific data levels. The design of the evaluation matrix takes into account the characteristics of the actual data distribution, and the interval division and level definition can be adjusted as needed.

[0050] Data level information refers to the data classification result obtained by matching through the data evaluation matrix based on the data sensitivity score. It usually includes the level name (such as "highly sensitive", "relatively sensitive", "generally sensitive", etc.), level description, specific protection requirements (such as access control requirements, encrypted storage requirements, etc.). These level information provide specific guiding basis for data security management and protection measures.

[0051] Specifically, the system first uses the preset sensitive word library to scan the content feature part in the data feature identifier. The preset sensitive word library contains keywords and phrases with different sensitivity levels, and these sensitive words are divided into different levels according to their sensitivity levels, and each level corresponds to a different basic score. During the scanning process, the system will identify the sensitive words that appear in the content feature and calculate the initial sensitivity score according to the level and frequency of the sensitive words. At the same time, the system also considers the context in which the sensitive words appear and adjusts the initial score through semantic analysis to avoid misjudgment caused by word ambiguity. After obtaining the initial sensitivity score, the system uses the preset data evaluation matrix to perform multi-dimensional evaluation on the data. The data evaluation matrix is a multi-dimensional evaluation model that comprehensively considers the business attribute information in the data feature identifier, including factors such as the timeliness of the data, the importance of the business, and the scope of use. The evaluation matrix sets different weight coefficients for each dimension, and obtains the final data sensitivity score through weighted calculation. For example, data related to core business will obtain a higher weight, while general daily data will obtain a lower weight. Based on the final data sensitivity score, the system determines the data level information with reference to the preset grading standard. This grading standard defines the data levels corresponding to different score intervals, usually including multiple levels, such as highly sensitive, relatively sensitive, generally sensitive, etc. The system will map the data sensitivity score to the corresponding level interval to obtain the final data level information.

[0052] Based on the above embodiments, as an alternative embodiment, in step 103: using the preset sensitive word library to scan each data feature identifier to generate the corresponding data sensitivity score, this step may further include the following steps:

[0053] Step 301: Determine the corresponding text scanning range according to each data feature identifier; for each data feature identifier, retrieve multiple sensitive word entries corresponding to the text scanning range from a preset sensitive word library, where the sensitive word entries include sensitive word text, sensitive word weight, and sensitive word category.

[0054] Specifically, the system first analyzes the content features and business attribute information in the data feature identifier and determines the text scanning range based on this information. The high-frequency keywords in the content features can reflect the subject area to which the data belongs, while the business attributes indicate the business type and usage scenario of the data. For example, when the data feature identifier shows that the data belongs to the financial field and is used by the finance department, the system will determine the scanning range as the sensitive information category related to finance; when the data involves personnel information, the scanning range will include the sensitive information category related to personal privacy. After determining the scanning range, the system retrieves the corresponding sensitive word entries from the preset sensitive word library. The preset sensitive word library adopts a hierarchical classification structure and stores sensitive words classified according to different business fields and information types. Each sensitive word entry contains three key elements: sensitive word text, sensitive word weight, and sensitive word category. The sensitive word text is the specific word or phrase; the sensitive word weight reflects the sensitivity of the word, usually represented by a numerical value; the sensitive word category indicates the information category to which the word belongs, such as trade secrets, personal information, financial data, etc. In practical applications, this processing method shows good flexibility and scalability. The system can dynamically adjust the structure and content of the sensitive word library according to the specific needs of the enterprise, add new sensitive word categories or update the weights of existing sensitive words. At the same time, the sensitive word library supports multiple matching modes, including exact matching, fuzzy matching, and regular expression matching, which can meet the sensitive information recognition requirements in different scenarios. The text scanning range and the retrieved sensitive word entries determined in this way provide a data basis for the subsequent sensitivity scoring.

[0055] Step 302: Match the text within the text scanning range with the sensitive word text in the sensitive word entries to obtain sensitive word hit records; multiply the sensitive word weight and the sensitive word hit times of each target sensitive word entry in the sensitive word hit records to obtain a sub-sensitivity score.

[0056] Specifically, the system first performs sensitive word matching on the content within the text scanning range. The matching process uses a multi-pattern string matching algorithm, such as the Aho-Corasick automaton algorithm, which can perform parallel matching of multiple sensitive words simultaneously, improving the processing efficiency. During the matching process, the system considers variants and combination forms of words, such as homophones, synonyms and other variants, as well as different permutations and combinations of sensitive words. When a sensitive word is found in the text, the system records the specific location of the sensitive word, the information of the sensitive word entry it belongs to, and the number of occurrences, forming a sensitive word hit record. To accurately evaluate the sensitivity of the data, the system needs to comprehensively consider the importance and occurrence frequency of sensitive words. For each target sensitive word entry in the sensitive word hit record, the system multiplies its sensitive word weight by the number of hits to obtain the sub-sensitivity score of the sensitive word. The sensitive word weight reflects the sensitivity of the word, and usually more sensitive words are assigned higher weight values; while the number of hits reflects the usage frequency of the sensitive word in the text, and a sensitive word that appears multiple times usually indicates a higher degree of association between the data and the sensitive information. In practical applications, this processing method exhibits good performance and accuracy. The system uses an efficient matching algorithm and can quickly process a large amount of text data. At the same time, by recording detailed hit information, the system can provide a visual display of the distribution of sensitive information for managers to help understand the sensitive characteristics of the data. The system also regularly analyzes the scoring results for optimizing the configuration of sensitive word weights to make the scoring results more in line with the actual situation. The sub-sensitivity scores calculated in this way can more accurately reflect the distribution of various types of sensitive information in the text. Especially when dealing with complex data containing multiple types of sensitive information, this scoring method based on weight and frequency can effectively distinguish the importance of different sensitive words and provide a basis for subsequent comprehensive evaluation.

[0057] Step 303: Based on the category to which the sensitive word of each target sensitive word entry belongs, determine the corresponding category coefficient, and perform a weighted sum of each sub-sensitivity score with the category coefficients as weights to obtain the sensitivity score corresponding to the data feature identifier.

[0058] Specifically, the system first determines the corresponding category coefficient according to the category to which the target sensitive term belongs. Sensitive information of different categories has different degrees of influence on data sensitivity. For example, information involving corporate secrets may have a higher sensitivity level than general business information. The setting of the category coefficient is usually based on the enterprise's data security policy and compliance requirements, reflecting the relative importance of different categories of sensitive information. The system maintains a category coefficient configuration table that defines the corresponding weight values for each sensitive information category, and more important categories are assigned higher category coefficients. After determining the category coefficient, the system calculates the weighted sum of each sub-sensitivity score. For each target sensitive term, the system multiplies its sub-sensitivity score by the corresponding category coefficient, and then adds up all the weighted scores to obtain the final sensitivity score. This method of weighted summation ensures that the scoring result can reflect the comprehensive influence of different categories of sensitive information, and at the same time maintains the dominant role of the higher sensitivity category in the final score. In practical applications, this processing method shows good adaptability and maintainability. The system can adjust the category coefficient in a timely manner according to new security requirements or business changes to ensure that the scoring standard always meets the current management needs. At the same time, the system will record the detailed scoring process, including the contribution value of each category. This information can help managers understand the sensitive characteristics of the data and provide a basis for formulating targeted protection measures. The sensitivity score calculated in this way has high accuracy. The scoring result not only reflects the overall level of sensitive information in the data, but also reflects the relative importance of different categories of sensitive information.

[0059] Based on the above embodiments, as an optional embodiment, in step 103: evaluating each data sensitivity score through a preset data evaluation matrix to obtain the corresponding data level information. This step may further include the following steps:

[0060] Step 304: Obtain a preset data evaluation matrix, where the preset data evaluation matrix includes multiple scoring intervals and the data level information corresponding to each scoring interval.

[0061] Specifically, the system first determines the corresponding category coefficient according to the category to which the target sensitive term belongs. Sensitive information of different categories has different degrees of influence on data sensitivity. For example, information involving national secrets may have a relatively high category coefficient, while general business information has a relatively low category coefficient. The setting of the category coefficient is usually based on the enterprise's security policies and compliance requirements and can be dynamically adjusted according to actual needs. The system multiplies each sub-sensitivity score by the corresponding category coefficient and then performs weighted summation to obtain the overall sensitivity score corresponding to the data feature identifier. This weighted calculation method ensures that the scoring result can accurately reflect the comprehensive influence of sensitive information of different categories. After obtaining the sensitivity score, the system determines the data level through a preset data evaluation matrix. The data evaluation matrix is a mapping relationship table that defines the data level information corresponding to different scoring intervals. The division of the scoring intervals needs to consider the distribution characteristics of data sensitivity and usually adopts a non-uniform division method to better distinguish data with different sensitivity levels. Each scoring interval corresponds to specific data level information, such as "highly sensitive", "relatively sensitive", "generally sensitive", etc., and also includes the processing suggestions and protection requirements for data at this level. This evaluation method has multiple advantages. First, by introducing the category coefficient, the system can reflect the differences in the importance of sensitive information of different categories, making the evaluation result more in line with actual security needs. Second, the preset data evaluation matrix provides a standardized basis for level division, ensuring the consistency and comparability of the evaluation results. In addition, the configurability of the evaluation matrix enables the system to flexibly adjust the rating criteria according to business requirements.

[0062] Step 305: Match each data sensitivity score with each scoring interval in the data evaluation matrix to obtain the data level information corresponding to the scoring interval where each data sensitivity score is located.

[0063] Specifically, the system first compares the data sensitivity score with the scoring intervals defined in the data evaluation matrix. These scoring intervals are usually defined in the form of numerical ranges. For example, [0 - 30] represents the low - sensitivity interval, [31 - 70] represents the medium - sensitivity interval, and [71 - 100] represents the high - sensitivity interval. The system uses an interval detection algorithm to determine the specific interval to which the score belongs. This process needs to consider the handling principles of interval boundaries. For example, a left - closed - right - open or left - open - right - closed interval definition method is adopted to ensure that the score can be uniquely mapped to a certain interval. After determining the scoring interval, the system directly obtains the corresponding data level information. The data level information usually includes content such as level name, level description, and protection requirements. For example, when the score falls into the high - sensitivity interval, the system will extract the "highly sensitive" level information corresponding to this interval, which may include specific protection requirements such as "strict access control is required" and "must be encrypted for storage". These level information provide clear guidelines for the subsequent processing and protection of data. This matching method has multiple advantages. First, it provides a standardized data classification method, ensuring that data with the same score can receive consistent level judgments. Second, through the preset evaluation matrix, the system can flexibly adjust the division of scoring intervals and the corresponding level definitions to adapt to the classification requirements of different scenarios. In addition, this method is also convenient for the statistics and analysis of evaluation results, helping to understand the overall sensitivity distribution of enterprise data. In practical applications, this processing method demonstrates good practicality and maintainability. The system can optimize the division of scoring intervals according to the actual data distribution, making the level division more reasonable. At the same time, by recording detailed information about the matching process, including the score value, the belonging interval, and the final level, etc., the system can provide managers with a complete trace of the evaluation process. These records can also be used for subsequent audits and optimization of evaluation criteria.

[0064] Step 104: Based on the data feature identifiers and data level information of each data node, associate with the reference index table of the central management node to obtain a hierarchical index table, and assign corresponding access paths to each data node in the hierarchical index table.

[0065] Among them, the hierarchical index table is an enhanced index structure formed on the basis of the reference index table of the central management node by associating the data feature identifiers and data level information of the data nodes. This index structure supports hierarchical management and access control based on data sensitivity, enabling the system to implement differential management strategies according to the sensitive levels of data.

[0066] An access path refers to the data access method and routing information assigned by the system to data nodes of different levels. It contains the specific technical paths required to access the data node, such as network access protocols, authentication requirements, encrypted transmission channels, etc. For highly sensitive data, its access path may include security mechanisms such as multiple authentication requirements, mandatory encrypted transmission, access log recording, etc.; for low-sensitivity data, a relatively simplified access method may be adopted. The design of the access path takes into account multiple factors such as the sensitivity level of the data, the security level of the storage location, and the network topology structure to ensure that data access is both secure and controllable and convenient to use.

[0067] Specifically, the system first associates the data feature identifier and data level information of the data node with the records in the reference index table. The reference index table contains the basic information of the data node, such as node identifier, storage location, data type, etc. The system appends the data feature identifier and level information to the corresponding index record through the unique identifier of the data node to form a more complete data description. This association operation not only retains the original index information but also adds the sensitivity attribute of the data, enabling the index table to support sensitivity-based data management. After the association is completed, the system assigns corresponding access paths to each data node according to the data level information and physical storage location. The design of the access path needs to consider multiple factors, including the sensitivity level of the data, the security level of the storage location, and the network topology structure. For highly sensitive data, the system may assign an access path with multiple layers of authentication or perform access control through a security gateway; for low-sensitivity data, a relatively simplified access method can be used. This differential access path assignment ensures that data of different levels can be protected at the corresponding level. This index management method has multiple advantages. First, by integrating data feature and level information into the index table, the system can achieve unified data management and access control. Second, the access path assignment based on level information enables the system to provide corresponding protection mechanisms for data of different sensitivities. In addition, the structure design of the hierarchical index table supports flexible query and management operations, facilitating the implementation of various data management functions. In practical applications, this processing method exhibits good scalability and maintainability. The system can dynamically update the structure of the hierarchical index table according to requirements, add new attribute fields or optimize the configuration of the access path. At the same time, by recording detailed index information and access paths, the system can provide managers with a complete view of data assets to help understand the distribution and usage of data. The system also regularly checks the validity of the access path to ensure the security and availability of data access.

[0068] Based on the above embodiments, as an alternative embodiment, in step 104: associating the reference index table of the central management node based on the data feature identifiers and data level information of each data node to obtain a hierarchical index table, this step may further include the following steps:

[0069] Step 401: Divide the reference index table into initial sub-index tables corresponding to multiple data level information; for each data node, determine the corresponding sub-index table in each initial sub-index table according to the data level information.

[0070] Specifically, the system first divides the reference index table into multiple initial sub-index tables according to the preset data level information. Each initial sub-index table corresponds to a specific data level, such as highly sensitive, relatively sensitive, and generally sensitive levels. This division retains the basic field structure in the reference index table but establishes independent data storage spaces, enabling logically complete separation of data at different levels. The creation process of the initial sub-index table includes operations such as table structure replication and index reconstruction to ensure that each sub-index table has complete management functions. After the creation of the initial sub-index tables is completed, the system assigns each data node to the corresponding sub-index table according to its data level information. This process is completed by comparing the level information of the data node with the level attributes of the sub-index table. When the sub-index table to which the data node should belong is determined, the system migrates or copies the complete index information of the node to the corresponding sub-index table. This assignment mechanism ensures that each data node can be correctly classified into the index table of the corresponding level. The hierarchical index structure formed in this way can effectively support the hierarchical management of data. Especially when dealing with a large amount of heterogeneous data, this hierarchical index method can provide more efficient data organization and query capabilities. The system will regularly check the data distribution of the sub-index tables and perform structure optimization or load balancing if necessary to ensure management efficiency.

[0071] Step 402: Establish an association relationship between the data feature identifiers of each data node and the corresponding sub-index table, and assign corresponding access permissions to each sub-index table to obtain the target sub-index table of each data node; integrate each target sub-index table to obtain a hierarchical index table.

[0072] Specifically, the system first establishes an association relationship between the data feature identifiers of the data nodes and their corresponding sub-index tables. This association is achieved by adding a feature association field to the sub-index table, which records the feature information of the data nodes, such as data type, business attributes, sensitive features, etc. The system converts the data feature identifiers into standardized index attributes through a feature mapping algorithm and writes these attribute information into the corresponding sub-index tables. This association mechanism enables the system to quickly locate and retrieve relevant data nodes based on data features. While establishing the feature association relationship, the system assigns corresponding access permissions to each sub-index table. The setting of access permissions is based on data level information, covering permission control rules for different operation types such as reading, modifying, and deleting. For the sub-index tables corresponding to high-sensitivity data, the system will set more stringent access restrictions, such as requiring multi-factor authentication, operation review, etc.; for the sub-index tables of low-sensitivity data, a relatively loose access control policy is adopted. This differential permission setting ensures that data at different levels can be protected at the corresponding level. After completing the feature association and permission assignment, each sub-index table becomes a target sub-index table containing complete management information. These target sub-index tables contain complete management metadata such as the basic information, feature attributes, and access permissions of the data nodes. The system combines all the target sub-index tables into a unified hierarchical index table through a unified index integration mechanism. This integration process preserves the independence of each sub-index table while establishing a unified retrieval and management interface. The hierarchical index table constructed in this way can effectively support the data governance requirements of enterprises. Especially when dealing with complex data access scenarios, this multi-level index structure can provide precise access control and efficient data retrieval services.

[0073] Based on the above embodiments, as an optional embodiment, in step 104: Assigning corresponding access paths to each data node in the hierarchical index table, this step may further include the following steps:

[0074] Step 403: Obtain the access permissions of each data node; generate access path identifiers for each data node based on a preset access path template and each access permission.

[0075] Specifically, the system first obtains the access permission information of each data node from the hierarchical index table. These access permission information include security control parameters such as the access level of the data node, operation restrictions, authentication requirements, etc. The system will parse these permission information and extract the key parameters that affect the generation of the access path, such as whether encrypted transmission is required, whether multi-factor authentication is required, whether access auditing is required, etc. These parameters will directly affect the structure and characteristics of the finally generated access path. After obtaining the access permission information, the system uses a preset access path template as the basic framework to generate the access path identifier. The access path template is a standardized path description structure that includes multiple configurable components such as protocol type, authentication method, transmission encryption, access control, etc. The system selects the appropriate component combination from the access path template according to the access permission of the data node and fills in the specific parameter values. For example, for highly sensitive data nodes, the system will select path components that include forced encrypted transmission and multi-factor authentication; while for low-sensitivity data nodes, relatively simplified access path components may be selected.

[0076] Based on the above embodiments, as an optional embodiment, in step 403: generating the access path identifier of each data node based on the preset access path template and each access permission, this step may further include the following steps:

[0077] Step 413: Determine the access protocol type of the security level corresponding to each access permission according to each access permission; bind each access protocol type to the preset access path template respectively to obtain the access path identifier of each data node.

[0078] Specifically, in the process of data access management, to ensure the security and standardization of data access, the system needs to establish a corresponding access protocol mechanism based on access permissions and effectively bind it to the access path template. This process is of great significance for implementing differential security access control. The system first analyzes the access permission information of each data node, extracts the security level requirements from it, and determines the corresponding access protocol type according to the preset security policy. The determination process of this protocol type takes into account multiple security factors, including data sensitivity, access authentication requirements, transmission security requirements, etc. For access permissions with a high security level, the system will select access protocols with strong authentication mechanisms and high security, such as using encrypted transmission protocols, multi-factor authentication protocols, etc.; for access permissions with a low security level, relatively simplified access protocols may be selected, such as basic authentication protocols. After determining the access protocol type, the system binds these protocol types to the preset access path template. The access path template, as a standardized access structure framework, contains multiple components such as protocol parameter configuration, authentication method setting, transmission control, etc. The system uses a protocol mapping mechanism to associate the determined access protocol type with the corresponding components in the template and configures specific protocol parameters according to security requirements. This binding process ensures that each data node can obtain an access control mechanism that matches its security level. This processing method realizes the standardized management of access protocols, ensuring that all data accesses follow the predetermined security specifications. Secondly, through the binding of the protocol type and the template, the system can flexibly adjust access control parameters and quickly adapt to different security requirements. In addition, the standardized access path structure facilitates the system to conduct unified access management and security audits.

[0079] Step 404: Establish a mapping relationship between the physical storage addresses of each data node and the access path identifiers, and encrypt each access path identifier according to the preset encryption rules to obtain the access paths corresponding to each data node.

[0080] Specifically, the system first establishes a mapping relationship between the physical storage address and the access path identifier for each data node. This mapping is achieved by creating an association table that contains information such as the unique identifier of the data node, the physical storage address, and the access path identifier. The system adopts a unified mapping rule to bind the storage address information with the access path identifier, forming a complete description of the access link. This mapping mechanism ensures that the system can accurately direct access requests to the correct data storage location while maintaining the integrity of access control. After establishing the mapping relationship, the system encrypts the access path identifier according to the preset encryption rules. The encryption rules include multiple security parameters such as encryption algorithm selection, key management, and encryption strength setting. The system first normalizes the access path identifier to ensure that its format meets the encryption requirements. Then, according to the sensitivity level of the data node, the corresponding encryption algorithm and key strength are selected to perform an encryption transformation on the access path identifier. For highly sensitive data, the system will adopt a stronger encryption algorithm and a more complex key system; for low-sensitivity data, a relatively simplified encryption scheme may be used. In practical applications, this access path processing mechanism exhibits good security and maintainability. The system can update the encryption key regularly to achieve dynamic changes in the access path and increase security. At the same time, by recording detailed mapping and encryption operation logs, the system can accurately track the generation and usage of the access path, facilitating security audits and problem troubleshooting. The access path generated in this way can effectively support secure data access requirements. Especially when dealing with sensitive data access, this multi-layer protection mechanism can provide reliable access control and security guarantees. The system will regularly evaluate the security of the access path and update the encryption parameters or adjust the mapping rules when necessary to ensure the effectiveness of the access control measures.

[0081] Referring to Figure 2 , a data classification and grading system based on big data provided by an embodiment of the present application. The system includes: a data acquisition module, a feature identifier determination module, a level information determination module, and a data grading module, where:

[0082] The data acquisition module is used to acquire the data to be graded of multiple data nodes and the reference index table of the central management node;

[0083] The feature identifier determination module is used to extract the content features and business attributes of each data to be graded, and generate corresponding data feature identifiers according to each content feature and each business attribute;

[0084] The level information determination module is used to scan each data feature identifier using a preset sensitive word library, generate corresponding data sensitivity scores, and evaluate each data sensitivity score through a preset data evaluation matrix to obtain corresponding data level information;

[0085] A data classification module, which is used to associate the reference index table of the central management node based on the data feature identifiers and data level information of each data node, obtain a classified index table, and allocate corresponding access paths to each data node in the classified index table.

[0086] Based on the above embodiments, the feature identifier determination module is further configured to, for each data to be classified, perform text tokenization on the data to be classified to obtain a keyword set; count the occurrence frequencies of each keyword in the keyword set, and select a preset number of high-frequency keywords as content features, where the high-frequency keywords are keywords whose occurrence frequencies are greater than the frequency threshold; obtain the creation time, business scope, and using department of the data to be classified, and determine the business attribute based on a preset business type rule template, where the preset business type rule template includes business attributes corresponding to various combinations of different creation times, different business scopes, and different using departments; perform feature vectorization processing on the content features and business attributes to generate a data feature identifier.

[0087] Based on the above embodiments, the level information determination module is further configured to, according to each data feature identifier, determine the corresponding text scanning range; for each data feature identifier, retrieve multiple sensitive word entries corresponding to the text scanning range from a preset sensitive word library, where the sensitive word entries include sensitive word texts, sensitive word weights, and sensitive word categories; match the text within the text scanning range with the sensitive word texts in the sensitive word entries to obtain sensitive word hit records; multiply the sensitive word weights and sensitive word hit times of each target sensitive word entry in the sensitive word hit records to obtain a sub-sensitivity score; based on the sensitive word categories of each target sensitive word entry, determine the corresponding category coefficients, and perform weighted summation on each sub-sensitivity score with the category coefficients as weights to obtain the sensitivity score corresponding to the data feature identifier.

[0088] Based on the above embodiments, the level information determination module is further configured to obtain a preset data evaluation matrix, where the preset data evaluation matrix includes multiple score intervals and the data level information corresponding to each score interval; match each data sensitivity score with each score interval in the data evaluation matrix to obtain the data level information corresponding to the score interval where each data sensitivity score is located.

[0089] Based on the above embodiments, the data classification module is further configured to divide the reference index table into initial sub-index tables corresponding to multiple data level information; for each data node, determine the corresponding sub-index table in each initial sub-index table according to the data level information; establish an association relationship between the data feature identifier of each data node and the corresponding sub-index table, and allocate corresponding access permissions to each sub-index table to obtain the target sub-index table of each data node; integrate each target sub-index table to obtain a classified index table.

[0090] Based on the above embodiments, the data classification module is further configured to obtain the access permissions of each data node; generate the access path identifiers of each data node based on a preset access path template and each access permission; establish a mapping relationship between the physical storage address of each data node and the access path identifier, and perform encryption processing on each access path identifier according to a preset encryption rule to obtain the access path corresponding to each data node.

[0091] Based on the above embodiments, the data classification module is further configured to determine the access protocol type of the security level corresponding to each access permission according to each access permission; bind each access protocol type to a preset access path template respectively to obtain the access path identifier of each data node.

[0092] It should be noted that: when the device provided in the above embodiments implements its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be elaborated here.

[0093] This application also discloses an electronic device. Refer to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.

[0094] Among them, the communication bus 302 is used to realize the connection and communication between these components.

[0095] Among them, the user interface 303 may include a display interface and a camera interface. Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0096] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0097] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface graphics, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0098] Among them, the memory 305 may include random access memory (RAM), and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. Refer to Figure 3 , in the memory 305 as a computer storage medium, there may be included an operating system, a network communication module, a user interface module, and an application program of a data classification and grading method based on big data.

[0099] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user. The processor 301 can be used to call an application program stored in the memory 305 for a data classification and grading method based on big data. When executed by one or more processors 301, the electronic device 300 is caused to execute the method(s) of one or more of the above embodiments. It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0100] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0101] In several implementation manners provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0102] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0103] In addition, in each embodiment of this application, the functional units can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0104] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0105] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation schemes of the present disclosure after considering the specification and the practice of the disclosure.

[0106] This application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include well-known common knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary.

Claims

1. A data classification and grading method based on big data, characterized in that: include: Obtaining the data to be classified of multiple data nodes and the benchmark index table of the central management node; Extracting content features and business attributes of each of the data to be classified, and generating corresponding data feature identifiers according to each of the content features and each of the business attributes; Scanning and processing each of the data feature identifiers using a preset sensitive word library to generate a corresponding data sensitivity score, and evaluating each of the data sensitivity scores using a preset data evaluation matrix to obtain corresponding data level information; Associating the reference index table of the central management node based on the data feature identifier and data level information of each data node to obtain a hierarchical index table, and allocating a corresponding access path to each data node in the hierarchical index table; The allocating a corresponding access path to each of the data nodes in the hierarchical index table includes: Obtaining access rights to each of the data nodes; Based on a preset access path template and each of the access permissions, generating an access path identifier for each of the data nodes; Establishing a mapping relationship between the physical storage address of each data node and the access path identifier, and encrypting each access path identifier according to a preset encryption rule to obtain an access path corresponding to each data node; The step of generating an access path identifier for each data node based on a preset access path template and each access permission includes: Determine, based on each of the access rights, the type of access protocol corresponding to the security level of each of the access rights; Each of the access protocol types is respectively bound to the preset access path template to obtain an access path identifier of each of the data nodes.

2. The data classification and grading method based on big data according to claim 1 is characterized in that: The extracting of content features and service attributes of each of the to-be-classified data, and generating corresponding data feature identifiers according to each of the content features and each of the service attributes, includes: For each of the to-be-classified data, performing text segmentation processing on the to-be-classified data to obtain a keyword set; Counting the frequency of occurrence of each keyword in the keyword set, and selecting a preset number of high-frequency keywords as content features, wherein the high-frequency keywords are keywords whose frequency of occurrence is greater than a frequency threshold; Acquire the creation time, business scope and user department of the data to be classified, and determine the business attributes based on a preset business type rule template, wherein the preset business type rule template includes business attributes corresponding to a combination of multiple different creation times, different business scopes and different user departments; The content features and service attributes are subjected to feature vectorization processing to generate data feature identifiers.

3. The data classification and grading method based on big data according to claim 1 is characterized in that: The method of using a preset sensitive word library to scan and process each of the data feature identifiers to generate a corresponding data sensitivity score includes: Determine the corresponding text scanning range according to each of the data feature identifiers; For each of the data feature identifiers, a plurality of sensitive word entries corresponding to the text scanning range are retrieved from the preset sensitive word library, wherein the sensitive word entries include sensitive word text, sensitive word weight, and category to which the sensitive word belongs; Matching the text within the text scanning range with the sensitive word text in the sensitive word entry to obtain a sensitive word hit record; Multiplying the sensitive word weight and the number of sensitive word hits of each target sensitive word entry in the sensitive word hit record to obtain a sub-sensitivity score; Based on the category to which the sensitive word of each target sensitive word entry belongs, a corresponding category coefficient is determined, and a weighted sum is performed on each sub-sensitivity score using each category coefficient as a weight to obtain a sensitivity score corresponding to the data feature identifier.

4. The data classification and grading method based on big data according to claim 1 is characterized in that: The data sensitivity scores are evaluated by a preset data evaluation matrix to obtain corresponding data level information, including: Obtaining a preset data evaluation matrix, wherein the preset data evaluation matrix includes a plurality of scoring intervals and data level information corresponding to each of the scoring intervals; Each of the data sensitivity scores is matched with each scoring interval in the data evaluation matrix to obtain data level information corresponding to the scoring interval in which each of the data sensitivity scores is located.

5. The data classification and grading method based on big data according to claim 1 is characterized in that: The step of associating the reference index table of the central management node based on the data feature identifier and the data level information of each data node to obtain a hierarchical index table includes: Dividing the reference index table into a plurality of initial sub-index tables corresponding to data level information; For each of the data nodes, determining a corresponding sub-index table in each of the initial sub-index tables according to the data level information; Establishing an association relationship between the data feature identifier of each data node and the corresponding sub-index table, and assigning corresponding access rights to each sub-index table, to obtain a target sub-index table of each data node; The target sub-index tables are integrated to obtain a hierarchical index table.

6. A data classification and grading system based on big data, characterized in that: The system comprises: A data acquisition module, used to acquire the data to be classified of multiple data nodes and the benchmark index table of the central management node; A feature identification determination module, used to extract the content features and business attributes of each of the data to be classified, and generate a corresponding data feature identification according to each of the content features and each of the business attributes; A level information determination module, used to scan and process each of the data feature identifiers using a preset sensitive word library to generate a corresponding data sensitivity score, and evaluate each of the data sensitivity scores using a preset data evaluation matrix to obtain corresponding data level information; A data classification module, used to associate the reference index table of the central management node based on the data feature identifier and data level information of each data node to obtain a hierarchical index table, and allocate a corresponding access path to each data node in the hierarchical index table; The allocating a corresponding access path to each of the data nodes in the hierarchical index table includes: Obtaining access rights to each of the data nodes; Based on a preset access path template and each of the access permissions, generating an access path identifier for each of the data nodes; Establishing a mapping relationship between the physical storage address of each data node and the access path identifier, and encrypting each access path identifier according to a preset encryption rule to obtain an access path corresponding to each data node; The step of generating an access path identifier for each data node based on a preset access path template and each access permission includes: Determine, based on each of the access rights, the type of access protocol corresponding to the security level of each of the access rights; Each of the access protocol types is respectively bound to the preset access path template to obtain an access path identifier of each of the data nodes.

7. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the data classification and grading method based on big data as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the data classification and grading method based on big data as described in any one of claims 1 to 5 is executed.

Citation Information

Patent Citations

  • Cryptographic technology security control method and system for core table data

    CN112182616A