Method and device for optimizing data classification and grading scanning performance and electronic equipment
Through the preprocessing and cache optimization method combined with the decision tree algorithm and the CN2 algorithm, the problem of traditional data classification and grading methods taking too long to process large data volumes and server resource occupancy is solved, and efficient data processing and resource utilization is achieved.
Patent Information
- Application Number
- CN202510252658.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-06
AI Technical Summary
When traditional data classification and grading methods deal with large amounts of data in complex scenarios, preset identification rules need to match each piece of data, resulting in too long scanning and identification, occupying server resources for a long time, affecting the normal use of other functions.
The decision tree algorithm is used for preprocessing, a first-level sensitive data recognition model is established, the data is classified and graded, and sample data is extracted through the cache device. The CN2 algorithm is used to set the second-level sensitive data recognition rules to match the sample data, and the amount of scanned data is reduced.
Through preprocessing and cache optimization, the amount of scanned data is significantly reduced, data processing efficiency is improved, server resource consumption is reduced, server resource consumption is avoided, and system efficiency and stability is improved.
Smart Images

Figure CN119939356A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of data classification and grading, and in particular, relates to a method, device, computer-readable storage medium, and electronic device for optimizing data classification and grading scanning performance. Background Art
[0002] Data classification and grading technology has developed rapidly in the past few years. With the surge in information volume, the strengthening of privacy protection requirements and the increase in compliance needs, data classification and grading technology has gradually become more intelligent, automated, and able to cope with increasingly complex data management challenges.
[0003] With the introduction of privacy protection regulations such as GDPR (General Data Protection Regulation), companies are becoming more and more strict in protecting sensitive data. The development of data classification and grading technology is not only to improve data management efficiency, but more importantly to ensure the compliance and privacy protection of sensitive data. The traditional data classification and grading method relies on preset rules, scanning data, matching rules, marking and labeling data. Sensitive data is usually labeled with specific labels (such as "sensitive", "confidential", "public", etc.), and processed and stored accordingly according to the labels. In a cloud environment, the data security management platform is used to identify sensitive data in the database and label it. The data objects faced are very large, often with complex scenarios, complex table structures, and huge amounts of table data. Based on preset rule recognition, the classification and grading of data has the following defects and deficiencies: 1. In complex scenarios with large amounts of data, the preset recognition rules need to match each piece of data. The time spent on scanning and recognition often increases by multiples of the amount of data, causing the process to take a very long time.
[0004] 2. During this period, functional services will occupy a large amount of server resources for a long time, and the server's CPU, memory and other resources will be under high load for a long time, which will affect the normal use of other functions. Summary of the invention
[0005] In order to address the above problems, the present application proposes a new method for optimizing data classification and grading scanning performance. This method aims at the scenario where a large amount of data needs to match preset rules for data classification and grading processing. A preprocessing mechanism is added to the processing method, and the preprocessed data is stored in a temporary cache device. Then, the matching and identification processing of the preset rules is performed again on the data stored in the cache device.
[0006] In order to achieve the above objectives, this application provides the following technical solutions: A first aspect of the present application provides a method for optimizing data classification and grading scanning performance, the method comprising: Preprocessing step: Use decision tree algorithm to extract data features from industry data, establish a primary sensitive data identification model, and use the model to perform primary classification and grading on the scanned data to obtain table structure information and its classification and grading labels; Cache step: obtaining sample data from the data table according to a preset percentage, and storing the sample data in a cache device; Secondary processing steps: Use the CN2 algorithm to set secondary sensitive data identification rules, match the sample data in the cache device, obtain the classification and grading labels of the sample data, and mark, label and store them; Final processing step: Based on the labeling of the sample data, label and store each piece of information data in the table structure to complete the classification, grading and labeling of the data.
[0007] Optionally, in the method of the present application, the pretreatment step comprises: According to the characteristics of industry data, a decision tree algorithm is used to extract data features, and a primary sensitive data identification model is established based on the extracted data features; Get all scan table structure information of the scan library; Use the first-level sensitive data identification model to scan and match table structure information; Get the first-level classification and grading label information of each information in the table structure.
[0008] Optionally, in the method of the present application, the characteristics of the industry data include but are not limited to: numerical type, categorical type, Boolean type, text type, normalized characteristics, and standardized characteristics.
[0009] Optionally, in the method of the present application, the caching step includes: According to the percentage preset in the cache device, sample data of each information of the table structure is obtained; The sample data is stored in a cache device.
[0010] Optionally, in the method of the present application, the secondary processing step includes: According to the characteristics of business data, the CN2 algorithm is used to set secondary sensitive data identification rules; Obtaining sample data in a cache device; Match each piece of sample data with the secondary sensitive data identification rules; The sample data is classified, graded, labeled and stored.
[0011] Optionally, in the method of the present application, the characteristics of the business data include but are not limited to: industry, key business indicators, data life cycle, and data relevance.
[0012] Optionally, in the method of the present application, the final processing step includes: According to the labeling of sample data, each piece of information data in the table structure is labeled and stored; Complete the classification, grading and labeling of all data in different tables in the database.
[0013] A second aspect of the present application provides a device for optimizing data classification and grading scanning performance, the device comprising: The preprocessing module is used to extract data features from industry data using a decision tree algorithm, establish a primary sensitive data identification model, and use the model to perform primary classification and grading processing on the scanned data to obtain table structure information and its classification and grading labels; A cache module, used to obtain sample data from a data table according to a preset percentage and store the sample data in a cache device; The secondary processing module is used to set secondary sensitive data identification rules using the CN2 algorithm, match the sample data in the cache device, obtain the classification and grading labels of the sample data, and mark, label and store them; The final processing module is used to label and store each piece of information data in the table structure according to the labeling of the sample data, and complete the classification and grading of the data.
[0014] The device implements the steps of the aforementioned method for optimizing data classification and grading scanning performance when running.
[0015] A third aspect of the present application provides an electronic device, comprising: a memory and a processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the aforementioned method for optimizing data classification and grading scanning performance.
[0016] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the aforementioned method for optimizing data classification and grading scanning performance are implemented.
[0017] To sum up, the present application scheme optimizes the data classification and grading process through the preprocessing mechanism, optimization of the cache device, and asynchronous processing of the classification and grading service, improves the efficiency of data processing, reduces the consumption of server resources, and has significant technical advantages and practical value.
[0018] Compared with the existing technology, this solution has the following advantages: (1) Application of preprocessing mechanism: This solution introduces a preprocessing mechanism that can preset a limited number of first-level sensitive data identification models based on the data characteristics of a specific industry. These models are specifically used to identify the structural information of each table in the database. During the preprocessing process, it is only necessary to scan the structural information of each table to directly obtain the label status of each field in the table structure. Since the number of preset industry rules and table structure identification rules is small, the number of table structure fields that need to be identified is also reduced accordingly, which significantly reduces the amount of data scanned, so that the classification and grading of the table structure can be quickly identified and marked.
[0019] (2) Optimization of the cache device: The cache device in this solution designs a method of extracting sample data by percentage, and sets a secondary sensitive identification model according to the characteristics of the business data. According to the preset sample extraction rate, only a small amount of sample data needs to be extracted, and these sample data are scanned, matched and identified through the preset secondary sensitive data identification rules. In this way, the amount of data scanned is no longer all the data in the entire table, but is reduced to sample data, so that the classification and grading label information of the sample data can be quickly obtained.
[0020] (3) Asynchronous processing of classification and grading services: This solution divides the classification and grading service into two independent services to process different types of data asynchronously. This design can alleviate the high load requirements on server resources in the case of single-task scanning, solve the problem of long-term consumption and occupation of server resources, and improve the efficiency and stability of the overall system.
[0021] Other features and advantages of the present application will be described in detail in the subsequent description, or can be understood by implementing the relevant technical solutions of the present application. The purpose and other advantages of the present application can be achieved through the technical features and technical means clearly indicated in the description, claims and drawings, and obtained through the implementation process of these technical contents. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings involved in the description of the embodiments. It should be noted that the drawings only show some embodiments of the present application. For those skilled in the art, other related drawings can be derived from these drawings without creative work.
[0023] Figure 1 An overall implementation flowchart of the method for optimizing data classification and grading scanning performance for this application.
[0024] Figure 2 This is the overall design architecture diagram of this application method.
[0025] Figure 3This is a flow chart of the pre-processing steps in the method of this application.
[0026] Figure 4 This is a flow chart of the cache device processing steps in the method of this application.
[0027] Figure 5 A structural diagram of the device for optimizing data classification and grading scanning performance for this application.
[0028] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be clear that the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present application.
[0030] In this document, the term "including" and any of its variations (such as "including", "including", etc.) are open expressions and should be understood as "including but not limited to", that is, the listed contents are not exhaustive and may also include other contents not explicitly mentioned. The term "based on" should be understood as "based at least in part", that is, the basis or condition referred to may not be the only factor, and other relevant factors may also be involved. The term "one embodiment" should be understood as "at least one embodiment", that is, the described embodiment is not the only possible implementation method, and there may be other similar embodiments.
[0031] In this application, when the terms "one" and "multiple" are used to modify related elements or features, their expressions are illustrative rather than restrictive. Unless otherwise clearly stated in the context, "one" should be understood as "at least one" and "multiple" should be understood as "at least two". Those skilled in the art should reasonably interpret these terms based on the semantics and logical relationship of the context to ensure that they cover the possibility of "one or more".
[0032] Terminology explanation: Decision tree algorithm: It is a classic supervised learning algorithm, widely used in classification and regression tasks. It makes decisions by building a tree model, gradually dividing the data set according to the features, and finally forming a series of rules for predicting the value of the target variable.
[0033] CN2 (Classification based on 2 Rules algorithm) algorithm: is a rule-based classification algorithm widely used in data mining and machine learning.
[0034] Figure 1 The overall implementation process of the method for optimizing data classification and grading scanning performance provided by the present application is shown, including the following steps: S1. Preprocessing step: Use the decision tree algorithm to extract data features from industry data, establish a primary sensitive data identification model, and use the model to perform primary classification and grading on the scanned data to obtain table structure information and its classification and grading labels; S2 caching step: according to a preset percentage, obtain sample data from the data table, and store the sample data in a cache device; S3. Secondary processing step: Use CN2 algorithm to set secondary sensitive data identification rules, match the sample data in the cache device, obtain the classification and grading labels of the sample data, and mark, label and store them; S4. Final processing step: According to the labeling of the sample data, label and store each piece of information data in the table structure to complete the classification and grading of the data.
[0035] In order to more clearly illustrate the technical solution of the present application, further explanation will be given below through embodiments of specific scenarios.
[0036] This method aims at the scenario where a large amount of data needs to match preset rules for data classification and grading. A preprocessing mechanism is added to the processing method, and the preprocessed data is stored in a temporary cache device. The data stored in the cache device is then matched and identified again according to the preset rules.
[0037] like Figure 2 As shown, the implementation process of this method is as follows: 1. Use the decision tree algorithm to extract data feature attributes from industry data, develop a first-level sensitive data identification model, obtain scanned data, and divide the data into first-level classification and grading, such as Figure 3 As shown, the specific steps are as follows: (1) Based on the numerical, categorical, Boolean, textual, normalized, and standardized characteristics of industry data, a decision tree algorithm is used to obtain data features and a first-level sensitive data identification model is preset; (2) Obtain all scan table structure information of the scan library; (3) The first-level sensitive data identification model scans and matches the table structure information; (4) Obtain the first-level classification and grading label information for each information in the table structure.
[0038] 2. Obtain sample data from the data table based on the percentage and store it in the cache device. The specific steps are as follows: (1) Obtaining sample data for each information in the table structure according to a preset percentage in the cache device; (2) Store the sample data in a cache device.
[0039] 3. Use the CN2 algorithm to set the secondary sensitive data identification rules. For the data stored in the cache device, each sample data is matched with the preset secondary sensitive data identification rules (model), and the classification and grading labels of the sample data are obtained, and data marking, labeling processing and storage are performed, such as Figure 4 As shown, the specific steps are as follows: (1) Based on the industry, key business indicators, data life cycle, data relevance and other characteristics of the business data, the CN2 algorithm is used to preset secondary sensitive data identification rules; (2) Obtaining sample data from the cache device; (3) Each piece of data matches the secondary sensitive data identification rule once; (4) Classify, grade, label and store sample data.
[0040] 4. According to the labeling of the sample data, label processing and storage are performed for each piece of information data in the table structure, thereby completing the classification and grading of all data in different tables in the database.
[0041] Figure 5 The present invention provides a device for optimizing data classification and grading scanning performance, the device comprising: The preprocessing module is used to extract data features from industry data using a decision tree algorithm, establish a primary sensitive data identification model, and use the model to perform primary classification and grading processing on the scanned data to obtain table structure information and its classification and grading labels; A cache module, used to obtain sample data from a data table according to a preset percentage and store the sample data in a cache device; The secondary processing module is used to set secondary sensitive data identification rules using the CN2 algorithm, match the sample data in the cache device, obtain the classification and grading labels of the sample data, and mark, label and store them; The final processing module is used to label and store each piece of information data in the table structure according to the labeling of the sample data, and complete the classification and grading of the data.
[0042] When the above device is running, the steps of the method for optimizing data classification and grading scanning performance disclosed in this application are implemented.
[0043] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementations of the apparatus, methods, and computer program products according to various embodiments of the present application, including architecture, functions, and operations. In these figures, each box may represent a module, a program segment, or a portion of a code, which contains one or more executable instructions for implementing a specified logical function. It should be noted that each box in the block diagram and / or flowchart, as well as the combination of these boxes, can be implemented using a dedicated hardware-based system to implement the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0044] like Figure 6 As shown, the embodiment of the present application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing a computer program executable by the processor, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 runs the executable computer program to implement the steps of the above-mentioned method for optimizing data classification and grading scanning performance.
[0045] It is understandable that, in addition to the memory and the processor, the electronic device may also include an input device (such as a keyboard), an output device (such as a display) and other communication modules. These input devices, output devices and other communication modules communicate with the processor through an I / O interface (i.e., an input / output interface).
[0046] The operation of the present application can be implemented by writing computer program codes using one or more programming languages or a combination thereof. The programming languages include but are not limited to the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc.; A conventional procedural programming language, such as "C" or a similar programming language.
[0047] The execution methods of program code include but are not limited to: Executes entirely on the user's computer; Partial execution on the user's computer and part execution on a remote computer; Implemented as a standalone package; Executes entirely on the remote computer or server.
[0048] In scenarios involving remote computers, the remote computer can be connected to the user's computer through any type of network, including but not limited to a local area network (LAN) or a wide area network (WAN). In addition, the remote computer can also be connected to an external computer through an Internet service provider, such as using the Internet.
[0049] Furthermore, the present application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the various steps of the method for optimizing data classification and grading scanning performance disclosed in the present application.
[0050] In the context of this application, computer-readable storage media refers to tangible media that can store computer program code and related data. Specific examples include, but are not limited to, the following: (1) Portable computer disk: A removable magnetic storage medium such as a floppy disk.
[0051] (2) Hard disk: includes fixed storage devices such as mechanical hard disks and solid-state hard disks.
[0052] (3) Random Access Memory (RAM): Volatile storage medium used for temporary storage of data and program code.
[0053] (4) Read-only memory (ROM): A non-volatile storage medium used to store fixed programs and data.
[0054] (5) Erasable Programmable Read-Only Memory (EPROM) or Flash Memory: A non-volatile storage medium that supports multiple erasing and programming.
[0055] (6) Fiber optic storage device: storage medium based on fiber optic technology.
[0056] (7) CD-ROM: A read-only medium that stores data in the form of an optical disc.
[0057] (8) Optical storage devices: storage media based on optical principles, such as DVDs and Blu-ray discs.
[0058] (9) Magnetic storage devices: storage media based on magnetic principles, such as magnetic tapes and disks.
[0059] (10) Any suitable combination of the above: for example, combining multiple storage media to meet different storage requirements.
[0060] These computer-readable storage media can be used to store the program code and related data described in this application to support the operation of the program and the persistent storage of data.
[0061] In particular, according to an embodiment of the present application, the process described in the flowchart can be implemented as a computer software program. For example, an embodiment of the present application relates to a computer program product, which includes a computer program carried on a non-transitory computer-readable medium. The computer program contains program code for executing the method for optimizing data classification and grading scanning performance disclosed in the present application. When the computer program is executed by a processing device, the above functions defined in the embodiments of the present application can be implemented.
[0062] Although the above discussion contains some specific implementation details, these details should not be interpreted as limiting the scope of this application. The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the disclosure scope involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features. At the same time, this application should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above public concept.
[0063] Those skilled in the art should also understand that they can modify the technical solutions described in the above embodiments without departing from the spirit and scope of the technical solutions of the embodiments of the present application, or replace some of the technical features therein by equivalents. These modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for optimizing data classification and grading scanning performance, characterized in that: The following steps are involved: Preprocessing step: Use decision tree algorithm to extract data features from industry data, establish a primary sensitive data identification model, and use the model to perform primary classification and grading on the scanned data to obtain table structure information and its classification and grading labels; Cache step: obtaining sample data from the data table according to a preset percentage, and storing the sample data in a cache device; Secondary processing steps: Use the CN2 algorithm to set secondary sensitive data identification rules, match the sample data in the cache device, obtain the classification and grading labels of the sample data, and mark, label and store them; Final processing step: Based on the labeling of the sample data, label and store each piece of information data in the table structure to complete the classification, grading and labeling of the data.
2. The method according to claim 1, characterized in that: The pre-processing step comprises: According to the characteristics of industry data, a decision tree algorithm is used to extract data features, and a primary sensitive data identification model is established based on the extracted data features; Get all scan table structure information of the scan library; Use the first-level sensitive data identification model to scan and match table structure information; Get the first-level classification and grading label information of each information in the table structure.
3. The method according to claim 2, characterized in that The characteristics of the industry data include: numerical type, categorical type, Boolean type, text type, normalized characteristics, and standardized characteristics.
4. The method according to claim 1, characterized in that: The caching step includes: According to the percentage preset in the cache device, sample data of each information of the table structure is obtained; The sample data is stored in a cache device.
5. The method according to claim 1, characterized in that The secondary processing steps include: According to the characteristics of business data, the CN2 algorithm is used to set secondary sensitive data identification rules; Obtaining sample data in a cache device; Match each piece of sample data with the secondary sensitive data identification rules; The sample data is classified, graded, labeled and stored.
6. The method according to claim 5, characterized in that The characteristics of the business data include: industry, key business indicators, data life cycle, and data relevance.
7. The method according to claim 1, characterized in that The final processing steps include: According to the labeling of sample data, each piece of information data in the table structure is labeled and stored; Complete the classification, grading and labeling of all data in different tables in the database.
8. A device for optimizing data classification and grading scanning performance, characterized in that: The device comprises: The preprocessing module is used to extract data features from industry data using a decision tree algorithm, establish a primary sensitive data identification model, and use the model to perform primary classification and grading processing on the scanned data to obtain table structure information and its classification and grading labels; A cache module, used to obtain sample data from a data table according to a preset percentage and store the sample data in a cache device; The secondary processing module is used to set secondary sensitive data identification rules using the CN2 algorithm, match the sample data in the cache device, obtain the classification and grading labels of the sample data, and mark, label and store them; The final processing module is used to label and store each piece of information data in the table structure according to the labeling of the sample data, and complete the classification and grading of the data.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for optimizing data classification and grading scanning performance as described in any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: Memory and processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the method for optimizing data classification and grading scanning performance as described in any one of claims 1 to 7.
Citation Information
Cited By
Health risk grading evaluation method and system based on dynamic feature recognition
CN120544915A