Kettle-based distributed database desensitization system, method and equipment
By designing a distributed database desensitization system based on kettle and using a distributed parallel processing mechanism, the traditional data desensitization method has solved the problems of long development cycle, high maintenance cost and poor flexibility in the development process, and achieved efficient and flexible massive data desensitization processing.
Patent Information
- Application Number
- CN202510518602.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing traditional data desensitization methods have long development cycles, high maintenance costs and poor flexibility, making it difficult to effectively process massive sensitive data in distributed databases.
Design a distributed database desensitization system based on kettle, including data acquisition module, desensitization rule management module, kettle distributed execution engine, distributed database cluster and result storage module, and realize large-scale data desensitization through distributed parallel processing mechanism.
It significantly shortens the data desensitization processing time, improves the desensitization efficiency, can efficiently process massive data, and meets the data processing needs in a distributed database environment.
Smart Images

Figure CN120030603A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database security technology, and specifically to a kettle-based distributed database desensitization system, method and device. Background Art
[0002] With the advent of the big data era, the amount of data has exploded. Distributed databases are widely used in various industries due to their high scalability and high availability. However, distributed databases store a large amount of sensitive data, such as personal privacy information, commercial secrets, etc., which will have serious consequences once leaked. Therefore, it is very important to desensitize sensitive data in distributed databases.
[0003] However, the current traditional data desensitization methods usually use scripts or programs to process data, which has problems such as long development cycle, high maintenance cost, and poor flexibility. Summary of the invention
[0004] In view of the deficiencies in the prior art, the object of the present invention is to provide a distributed database desensitization system, method and device based on kettle.
[0005] To achieve the above object, the present invention provides the following technical solution: a distributed database desensitization system based on kettle, comprising a data acquisition module, a desensitization rule management module, a kettle distributed execution engine, a distributed database cluster and a result storage module. Distributed database cluster, used to store raw information data; The data collection module is used to collect data from each node in the distributed database cluster, and use the adapted driver and interface to read data for different types of distributed databases in the distributed database cluster, and transmit the collected data to the kettle distributed execution engine in the form of data stream; The desensitization rule management module is used to define and manage desensitization rules and store the desensitization rules in the rule library; Kettle distributed execution engine, used to receive data transmitted by the data acquisition module and the desensitization rules provided by the desensitization rule management module, and desensitize the data in a distributed environment; The result storage module is used to store the anonymized data processed by the kettle distributed execution engine.
[0006] In some of the embodiments, the desensitization rules include a method for identifying sensitive data, a selection of a desensitization algorithm, and a scope of application of the desensitization rules.
[0007] In some of the embodiments, the sensitive data identification method includes matching sensitive fields through regular expressions or identifying sensitive data based on a data dictionary; The selection of the desensitization algorithm includes one or more combinations of replacement, masking, and encryption; The applicable scope of the rules includes database tables, fields or data subsets.
[0008] In some of the embodiments, the method for processing data desensitization by the kettle distributed execution engine is: Based on the Apache Kettle big data processing platform, using Kettle's conversion and operation mechanism, the data collected by the data collection module is extracted, converted and loaded based on desensitizing rules.
[0009] In some of the embodiments, the method of desensitizing data in a distributed environment is as follows: By configuring multiple Kettle execution nodes, data processing tasks are distributed to each node in parallel for execution. Each Kettle execution node is responsible for processing a part of the data. Each node performs desensitization operations at the same time, and finally summarizes the processing results.
[0010] In some of the embodiments, the distributed database cluster is a distributed database system that stores original information data. The distributed database cluster includes multiple data nodes and management nodes. The data nodes are used to store and manage data, and the management nodes are used to coordinate data storage, query and update operations.
[0011] In some of the embodiments, the distributed database cluster includes multiple types of distributed databases.
[0012] To achieve the above object, the present invention also provides the following technical solution: a method for using a distributed database desensitization system based on kettle, using the distributed database desensitization system based on kettle, the steps are: (1) Start the data acquisition module, which connects to each node of the distributed database cluster according to the pre-configured data source information; (2) For different types of databases, use corresponding connection methods and query statements to obtain data; (3) The collected data is packaged according to the preset format requirements and transmitted to the kettle distributed execution engine through the network; (4) After receiving the data, the kettle distributed execution engine requests the applicable desensitization rules from the desensitization rule management module; (5) The desensitization rule management module retrieves the corresponding desensitization rule from the rule library according to the data source and the pre-set rule priority, and sends the desensitization rule to the kettle distributed execution engine; (6) The kettle distributed execution engine desensitizes the data according to the received desensitization rules; (7) According to step (6), for each data record, determine whether sensitive data exists according to the sensitive data identification method defined in the desensitization rule. If so, process it according to the selected desensitization algorithm; (8) In a distributed environment, data shards are allocated to each Kettle execution node. Multiple Kettle execution nodes process data in parallel. Each node processes data according to the allocated data subset and desensitization rules, and summarizes the processed results. (9) After the desensitization process is completed, the kettle distributed execution engine sends the processing results to the result storage module; (10) The result storage module stores the desensitized data to the specified target location according to the configured storage strategy.
[0013] In some of the embodiments, according to step (5), the rule priority management logic is: Define a serial number field for each rule, and set the smaller the serial number, the higher the priority. When the kettle distributed execution engine receives the desensitization rules, it first sorts them from small to high according to the serial number, and then desensitizes the data in the order of sorting.
[0014] To achieve the above purpose, the present invention also provides the following technical solution: a desensitizing device based on a distributed database of a kettle, comprising: Memory for storing computer programs; A processor is used to implement the steps of the method for using a kettle-based distributed database desensitization system when executing the computer program.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. By building a system architecture that is coordinated by five core modules: data collection, desensitization rule management, kettle distributed execution engine, distributed database cluster and result storage. Among them, the kettle distributed execution engine is deeply expanded based on the Apache Kettle platform, breaking through the limitations of traditional single-machine processing. With the help of distributed parallel processing mechanism, large-scale data processing tasks are disassembled and distributed to multiple execution nodes for simultaneous execution. For example, when processing ultra-large-scale distributed database tables with billions of records, compared with traditional single-machine desensitization methods, this system can significantly shorten the processing time by several times or even dozens of times, greatly improving the desensitization efficiency and effectively meeting the stringent requirements of massive data processing.
[0016] 2. The design of the desensitization rule management module greatly simplifies the rule definition and management process. In the sensitive data identification process, it supports regular expression matching and data dictionary-based identification methods, which greatly improves the accuracy and flexibility of identification. The selection of desensitization algorithms is rich and varied, covering a variety of mature technologies such as replacement, masking, encryption, etc., and allows customized desensitization rules for different database tables, fields and even specific data subsets. In addition, the module also has the function of setting rule priorities to ensure that desensitization rules can be executed in an orderly and efficient manner in complex business scenarios.
[0017] 3. The distributed execution engine of Kettle receives the data transmitted by the data acquisition module and the rules provided by the desensitization rule management module, and reasonably distributes the data processing tasks to each Kettle execution node. Each node carries out desensitization processing in parallel based on the assigned data subset and the corresponding desensitization rules. This distributed execution mode fully taps the potential of cluster computing resources, significantly improves the overall processing capacity and efficiency of the system, and effectively responds to data processing challenges in a distributed database environment.
[0018] 4. Through the data collection module's multi-type distributed database support capabilities, the system can closely interact with the distributed database management nodes to accurately obtain data distribution details and execution permissions, effectively ensuring the accuracy and reliability of data collection and processing, and fully adapting to the complex and diverse database architecture of the enterprise.
[0019] Details of one or more embodiments of the present application are presented in the following drawings and descriptions to make other features, purposes and advantages of the present application more concise and easy to understand. The present application is fully described and understood through the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a diagram of the architecture of the distributed database desensitization system based on Kettle of the present invention; Figure 2It is a schematic diagram of the structural principle of the acquisition module of the present invention; Figure 3 It is a configuration desensitization label interactive interface of the desensitization rule management module in the embodiment; Figure 4 It is the configuration label algorithm interactive interface of the desensitization rule management module in the embodiment; Figure 5 It is an interactive interface for configuring desensitization algorithm details of the desensitization rule management module in the embodiment; Figure 6 It is a flowchart of desensitization processing of the kettle distributed execution engine in a distributed environment in an embodiment; Figure 7 for Figure 1 Data desensitization processing flow chart of the system; Figure 8 It is a flow chart of the implementation method of the data acquisition module in the embodiment; Fig. 9 This is an example diagram of the desensitization rule algorithm structure of the desensitization rule management module in the embodiment; Fig.10 This is an example diagram of the mask rule structure of the desensitization rule management module in the embodiment; Fig.11 It is the desensitization algorithm editing interactive interface of the desensitization rule management module in the embodiment; Fig.12 It is a flow chart of the method for using the distributed database desensitization system based on kettle of the present invention; Fig.13 This is a system module topology diagram of the present invention. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] See also Figure 1 , 2 , 7 and 13, the present invention provides a technical solution: a distributed database desensitization system based on kettle, including a data acquisition module, a desensitization rule management module, a kettle distributed execution engine, a distributed database cluster and a result storage module, Distributed database cluster, used to store raw information data; The data collection module is used to collect data from each node in the distributed database cluster, and use the adapted driver and interface to read data for different types of distributed databases in the distributed database cluster, and transmit the collected data to the kettle distributed execution engine in the form of data stream; The desensitization rule management module is used to define and manage desensitization rules and store the desensitization rules in the rule library; Kettle distributed execution engine, used to receive data transmitted by the data acquisition module and the desensitization rules provided by the desensitization rule management module, and desensitize the data in a distributed environment; The result storage module is used to store the anonymized data processed by the kettle distributed execution engine.
[0023] Based on the above solution, as a preferred embodiment, Figure 2 and 8 As shown in the figure, its data acquisition module uses adapted drivers and interfaces to read data for different types of distributed databases (such as Hadoop Distributed File System (HDFS) combined with Hive database, Cassandra distributed database, etc.). For example, for Hive database, use Hive JDBC driver to connect to Hive server and obtain data of specified table or partition by executing SQL query statement; for Cassandra database, use Cassandra Java driver to read corresponding data according to keyspace and table name. The collected data is transmitted to kettle distributed execution engine in the form of data stream.
[0024] In practical applications, for example, for a distributed database based on Hadoop, the database table structure and data distribution information are obtained through the metadata service of Hive, and then the data blocks are read from HDFS using MapReduce jobs or Spark tasks. The collected data is packaged in a certain format (such as CSV, Parquet, etc.) and transmitted to the kettle distributed execution engine through the network.
[0025] At the same time, a fault-tolerance mechanism is set up: when the primary node (Leader) in the ZooKeeper cluster fails or the cluster starts, the system will enter crash recovery mode. In this stage, the cluster will elect a new Leader node, and each node will synchronize the latest data from the Leader node to ensure data consistency.
[0026] like Figure 3 , 4As shown in Figures 5, 9, 10 and 11, as a preferred embodiment, the desensitization rules include a method for identifying sensitive data, a selection of a desensitization algorithm, and an applicable scope of the desensitization rules.
[0027] The sensitive data identification method includes matching sensitive fields (such as ID card number, mobile phone number, etc.) through regular expressions or identifying sensitive data based on a data dictionary; The selection of the desensitization algorithm includes one or more combinations of replacement, masking, encryption, etc., such as Fig.11 As shown, serverless dynamic programming is used to implement desensitization algorithm editing; The applicable scope of the rules includes database tables, fields or data subsets.
[0028] The configured desensitization rules are stored in the rule base in JSON format for easy management and query. At the same time, the module provides a visual interface through which administrators can add, modify, delete, and set the priority of rules.
[0029] like Figure 6 As shown, as a preferred embodiment, the method for processing data desensitization by the kettle distributed execution engine is: Based on the Apache Kettle big data processing platform, using Kettle's conversion and operation mechanism, the data collected by the data collection module is extracted, converted and loaded based on desensitizing rules.
[0030] The method of desensitizing data in a distributed environment is as follows: By configuring multiple Kettle execution nodes, data processing tasks are distributed to each node in parallel for execution. Each Kettle execution node is responsible for processing a part of the data. Each node performs desensitization operations at the same time, and finally summarizes the processing results.
[0031] As a preferred embodiment, the distributed database cluster is a distributed database system for storing original information data. The distributed database cluster includes multiple data nodes and management nodes. The data nodes are used to store and manage data, and the management nodes are used to coordinate data storage, query and update operations.
[0032] The distributed database cluster includes various types of distributed databases, such as Hadoop-based distributed file systems and databases (HDFS + Hive, HBase, etc.), NoSQL distributed databases (Cassandra, MongoDB, etc.) and distributed relational databases (such as MySQL Cluster, etc.). When performing desensitization processing, the system interacts with the management node of the distributed database to obtain data distribution information and execution permissions to ensure accurate data collection and processing.
[0033] As a preferred embodiment, the result storage module can choose to store the desensitized data back to the specified table or partition in the original distributed database cluster (if the target is the original distributed database cluster, the data is written back through the database write interface (such as Hive's INSERT INTO statement, Cassandra's write operation, etc.)), or it can be stored in other target databases or file systems (if the target is other databases or file systems, use the corresponding tools and interfaces for storage).
[0034] During the storage process, the data is reorganized and formatted according to the structure and desensitization rules of the original data to ensure data integrity and availability (data storage in this system only needs to focus on data integrity, and there is no data consistency problem, because this system divides the data and processes it in batches, and there is no problem of the same data set being repeatedly processed by multiple threads). At the same time, the result storage module records the result information of data processing, such as the amount of data processed, processing time, success, etc., and feeds back to the relevant modules for monitoring and management.
[0035] Through the technical solution of this application, a full-process data processing guarantee mechanism is designed: from the beginning of data collection, exclusive drivers and interfaces are used for different types of databases to ensure accurate data collection; in the desensitization processing stage, sensitive data is effectively desensitized according to detailed rules; in the result storage link, it supports flexible selection of storage targets, which can be restored to the specified location of the original distributed database cluster, or stored in other target databases or file systems. At the same time, the result storage module is also responsible for recording key information of data processing (such as the amount of data processed, time consumption, success or failure, etc.), and timely feedback to related modules, providing strong data support for system monitoring and management, and realizing closed-loop management and quality assurance of the entire data processing process.
[0036] The performance of the technical solution of this application is compared with the traditional processing system: 1. Quantitative analysis of performance indicators and comparison with existing technologies are shown in Table 1: Table 1 Processing speed under different scale data
[0037] It can be clearly seen from the data in the above table that as the data scale increases, the processing speed advantage of this system in a distributed environment becomes more and more obvious, and the resource utilization rate is relatively low, which can make more efficient use of system resources.
[0038] The tests were conducted in different distributed cluster environments as shown in Table 2, including clusters with 3 nodes, 5 nodes, and 10 nodes.
[0039] Table 2 Performance in different distributed environments
[0040] It can be seen from Table 2 that as the number of cluster nodes increases, the processing speed is significantly improved, while the resource utilization rate is further reduced, indicating that the system has good scalability in a distributed environment.
[0041] Performance optimization measures Optimize the conversion and execution process of Kettle jobs: By fine-tuning the conversion steps in Kettle jobs, reducing unnecessary intermediate data storage and transmission, and merging and optimizing the data processing logic, the execution of each step is more efficient. For example, operations that originally required multiple reads and writes of temporary files can now be directly transferred in memory after optimization, and the time required to process 100 million data records has been reduced from 70 minutes to 60 minutes.
[0042] Adopt a more efficient distributed computing framework: Introduce the Spark distributed computing framework, use its memory computing and efficient scheduling mechanism to replace some of the less efficient distributed processing links in Kettle. When processing 1 billion pieces of data, combined with the Spark framework, the processing time was shortened from 700 minutes to 600 minutes.
[0043] 2. Comparative experimental data with traditional single-machine desensitization methods In the experiment of processing 1 billion data, the traditional single-machine desensitization method needs to process data line by line, and is limited by the computing power and memory size of a single machine. The entire processing process takes up to 4,000 minutes. In the process, the CPU is in full load for a long time, and the memory overflows. Additional processing steps are required to solve the problem of insufficient memory. This system uses a distributed processing method to disperse data to multiple nodes for parallel processing, which only takes 600 minutes, greatly improving processing efficiency. At the same time, resource utilization is relatively low, and the system operation is more stable.
[0044] 3. Compatibility test results Hive database test case Test environment: Hive 3.1.2 version, running on a 5-node Hadoop cluster.
[0045] Test data: A data table containing 100 million user information items, including sensitive information such as name, ID number, and mobile phone number.
[0046] Testing process: Integrate this system with the Hive database, configure the corresponding connection parameters and desensitization rules, and desensitize the data table.
[0047] Test results: The system can successfully connect to the Hive database, accurately read data and process it according to the preset desensitization rules, and the processed results are correctly written into the Hive table. It takes 70 minutes to process 100 million data items, and the data accuracy reaches more than 99.9%, proving that the system has good compatibility with the Hive database.
[0048] Cassandra database test case Test environment: Cassandra version 4.0.7, running on a 3-node cluster.
[0049] Test data: A data table containing 50 million product information, including sensitive information such as product price and inventory.
[0050] Testing process: Integrate this system with the Cassandra database, set appropriate connection and desensitization rules, and perform desensitization operations on the data table.
[0051] Test results: The system has a stable connection with the Cassandra database, can efficiently read data from the database and perform desensitization processing, and the processed results are accurately written back to the Cassandra table. It took 30 minutes to process 50 million data items, and data consistency was guaranteed, which fully demonstrated that the system is well adapted to the Cassandra database.
[0052] like Figure 7 and 12 As shown, a method for using a distributed database desensitization system based on kettle, using the distributed database desensitization system based on kettle, the steps are: (1) Start the data acquisition module, which connects to each node of the distributed database cluster according to the pre-configured data source information; (2) For different types of databases, use corresponding connection methods and query statements to obtain data; (3) The collected data is packaged according to the preset format requirements and transmitted to the kettle distributed execution engine through the network; (4) After receiving the data, the kettle distributed execution engine requests the applicable desensitization rules from the desensitization rule management module; (5) The desensitization rule management module retrieves the corresponding desensitization rule from the rule library according to the data source and the pre-set rule priority, and sends the desensitization rule to the kettle distributed execution engine; (6) The kettle distributed execution engine desensitizes the data according to the received desensitization rules; (7) According to step (6), for each data record, determine whether sensitive data exists according to the sensitive data identification method defined in the desensitization rule. If so, process it according to the selected desensitization algorithm (for example, for the ID card number field, use the masking algorithm to replace the middle digits with asterisks; for the mobile phone number field, use the replacement algorithm to replace the mobile phone number with a virtual but correctly formatted number); (8) In a distributed environment, data shards are allocated to each Kettle execution node. Multiple Kettle execution nodes process data in parallel. Each node processes data according to the allocated data subset and desensitization rules, and summarizes the processed results. (9) After the desensitization process is completed, the kettle distributed execution engine sends the processing results to the result storage module; (10) The result storage module stores the desensitized data to the specified target location according to the configured storage strategy.
[0053] According to the above technical solution, the rule priority management logic in step (5) is: Define a serial number field for each rule, and set the smaller the serial number, the higher the priority. When the kettle distributed execution engine receives the desensitization rules, it first sorts them from small to high according to the serial number, and then desensitizes the data in the order of sorting, covering the rule: the smaller the granularity, the higher the priority, library < table < field.
[0054] As a preferred embodiment, it also includes a data desensitization device based on a distributed database of kettle, which includes Memory for storing computer programs; A processor is used to implement the steps of the method for using the kettle-based distributed database desensitization system as mentioned in the above embodiment when executing the computer program.
[0055] The data desensitization device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0056] The processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), or a programmable logic array (PLA).
[0057] The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state.
[0058] In some embodiments, the processor may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen.
[0059] In some embodiments, the processor may also include an artificial intelligence (AI) processor for processing computing operations related to machine learning.
[0060] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In this embodiment, the memory is used to store at least the following computer programs: Among them, after the computer program is loaded and executed by the processor, it can implement the steps related to the method of using the kettle-based distributed database desensitization system disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to the data involved in the data desensitization method.
[0061] In some embodiments, the data desensitizing device may also include a display screen, an input and output interface, a communication interface, a power supply, and a communication bus.
[0062] In this embodiment, the data desensitization device includes a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the steps of the method for using the distributed database desensitization system based on kettle as mentioned in the above embodiment when executing the computer program.
[0063] As a preferred embodiment, it also includes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for using the kettle-based distributed database desensitization system mentioned in the above embodiment are implemented.
[0064] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0065] In this embodiment, a computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.
[0066] The above is a detailed introduction to a kettle-based distributed database desensitization system, method, device and medium provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referenced to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
[0067] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A distributed database desensitization system based on kettle, characterized by: It includes data collection module, desensitization rule management module, kettle distributed execution engine, distributed database cluster and result storage module. Distributed database cluster, used to store raw information data; The data collection module is used to collect data from each node in the distributed database cluster, and use the adapted driver and interface to read data for different types of distributed databases in the distributed database cluster, and transmit the collected data to the kettle distributed execution engine in the form of data stream; The desensitization rule management module is used to define and manage desensitization rules and store the desensitization rules in the rule library; Kettle distributed execution engine, used to receive data transmitted by the data acquisition module and the desensitization rules provided by the desensitization rule management module, and desensitize the data in a distributed environment; The result storage module is used to store the anonymized data processed by the kettle distributed execution engine.
2. A distributed database desensitization system based on kettle according to claim 1, characterized in that: The desensitization rules include the identification method of sensitive data, the selection of desensitization algorithms, and the scope of application of the desensitization rules.
3. A distributed database desensitization system based on kettle according to claim 2, characterized in that: The sensitive data identification method includes matching sensitive fields through regular expressions or identifying sensitive data based on a data dictionary; The selection of the desensitization algorithm includes one or more combinations of replacement, masking, and encryption; The applicable scope of the rules includes database tables, fields or data subsets.
4. The distributed database desensitization system based on kettle according to claim 1 is characterized by: The method for processing data desensitization by the kettle distributed execution engine is: Based on the Apache Kettle big data processing platform, using Kettle's conversion and operation mechanism, the data collected by the data collection module is extracted, converted and loaded based on desensitizing rules.
5. A distributed database desensitization system based on kettle according to claim 4, characterized in that: The method of desensitizing data in a distributed environment is as follows: By configuring multiple Kettle execution nodes, data processing tasks are distributed to each node in parallel for execution. Each Kettle execution node is responsible for processing a part of the data. Each node performs desensitization operations at the same time, and finally summarizes the processing results.
6. The distributed database desensitization system based on kettle according to claim 1 is characterized by: The distributed database cluster is a distributed database system for storing original information data. The distributed database cluster includes multiple data nodes and management nodes. The data nodes are used to store and manage data, and the management nodes are used to coordinate data storage, query and update operations.
7. A distributed database desensitization system based on kettle according to claim 6, characterized in that: The distributed database cluster includes multiple types of distributed databases.
8. A method for using a distributed database desensitization system based on kettle, characterized by: The distributed database desensitization system based on kettle as described in any one of claims 1 to 7 is adopted, and the steps are as follows: (1) Start the data acquisition module, which connects to each node of the distributed database cluster according to the pre-configured data source information; (2) For different types of databases, use corresponding connection methods and query statements to obtain data; (3) The collected data is packaged according to the preset format requirements and transmitted to the kettle distributed execution engine through the network; (4) After receiving the data, the kettle distributed execution engine requests the applicable desensitization rules from the desensitization rule management module; (5) The desensitization rule management module retrieves the corresponding desensitization rule from the rule library according to the data source and the pre-set rule priority, and sends the desensitization rule to the kettle distributed execution engine; (6) The kettle distributed execution engine desensitizes the data according to the received desensitization rules; (7) According to step (6), for each data record, determine whether sensitive data exists according to the sensitive data identification method defined in the desensitization rule. If so, process it according to the selected desensitization algorithm; (8) In a distributed environment, data shards are allocated to each Kettle execution node. Multiple Kettle execution nodes process data in parallel. Each node processes data according to the allocated data subset and desensitization rules, and summarizes the processed results. (9) After the desensitization process is completed, the kettle distributed execution engine sends the processing results to the result storage module; (10) The result storage module stores the desensitized data to the specified target location according to the configured storage strategy.
9. The method for using the kettle-based distributed database desensitization system according to claim 8, characterized in that: According to step (5), the rule priority management logic is: Define a serial number field for each rule, and set the smaller the serial number, the higher the priority. When the kettle distributed execution engine receives the desensitization rules, it first sorts them from small to high according to the serial number, and then desensitizes the data in the order of sorting.
10. A data desensitization device based on a distributed database of a kettle, characterized in that: include Memory for storing computer programs; A processor, used to implement the steps of a method for using a kettle-based distributed database desensitization system as described in claim 8 or 9 when executing the computer program.
Citation Information
Patent Citations
Database masking system and method based on big data
CN106599713A
Data desensitization system and method
CN107766741A
Data desensitization method and device and desensitization service platform
CN111274610A
Data desensitization method and device
CN114357498A
Data desensitization method and device, storage medium and electronic device
CN115935409A