Distributed real-time data exchange system based on Kafka
Through the management center, the data exchange channel and task executor are configured, combined with the Kafka cluster partition storage mechanism, the unified management problem of data exchange among various systems in the existing technology is solved, flexible configuration, security protection and efficient data exchange are realized, and system management and data security are improved.
Patent Information
- Application Number
- CN202510745996.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-18
AI Technical Summary
The existing real-time data exchange between various systems lacks a centralized and unified management architecture, and the exchange process is difficult to coordinate monitoring and flexible configuration. The large differences in data structures lead to high conversion and adaptation costs, insufficient data security, and susceptible to threats.
The management center is used to configure the data exchange channel and generate tasks. The task executor in the execution area is used to respond to the data exchange task. It combines the Kafka cluster to perform partition storage according to data concealment, and sets partition configuration rules and extraction rules to realize centralized management of the data exchange process, flexible processing of data structures, and security hierarchical protection.
A real-time data interaction environment that is efficient, secure and easy to manage is built, which improves the efficiency and convenience of system management, reduces development costs, and enhances data security and reliability.
Smart Images

Figure CN120342829A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data exchange, and specifically, to a distributed real-time data exchange system based on Kafka. Background Art
[0002] With the rapid development of information technology, various industries have applied a large number of information systems to meet different business requirements. For example, in enterprise operations, there are both CRM systems responsible for customer relationship management, ERP systems for enterprise resource planning management, and monitoring systems for production safety, etc. These systems are numerous, operate independently, and there is a large demand for real-time data exchange among them for purposes such as business collaboration. However, different systems are developed by different teams or even different manufacturers, and there are both internal and external systems. The existing real-time data exchange between systems is point-to-point, presenting a mesh structure, lacking a centralized unified management architecture like a management center, resulting in difficult overall monitoring and flexible configuration of the exchange process. In addition, the data structures of each system vary greatly. Different systems adopt different data structures to store and organize data according to their own business characteristics and design concepts. When these systems perform real-time data exchange, a large amount of time and effort are required for data structure conversion and adaptation, greatly increasing the development cost of real-time data exchange. At the same time, the performance of real-time data exchange between each system is uneven, and various problems are often brought to the normal business operation due to data latency. Finally, the issue of data security cannot be ignored. During the real-time data exchange process, due to the lack of a perfect security mechanism, data is vulnerable to threats such as theft and tampering, which may also lead to data abuse.
[0003] In summary, the current situation of real-time data exchange between existing systems urgently needs to be improved, and there is an urgent need for an efficient, secure and easy-to-manage data exchange solution.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present application, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The object of the present invention is to address the problem of difficult unified management of data caused by the complex real-time data exchange environment among existing systems. A distributed real-time data exchange system based on Kafka is proposed. By configuring data exchange channels and generating data exchange tasks through the management center, overall management of the entire data exchange process is achieved. Using multiple task executors in the execution area, data exchange tasks are responded to according to the data exchange channels, realizing real-time access, processing, pushing, and status monitoring of data. With the help of the Kafka cluster, data is partitioned and stored according to the concealment of the data, and data production requests and data consumption requests issued by the task executors are responded to. By setting partition configuration rules and extraction rules for different privacy data, centralized management of the data exchange process, flexible processing of data structures, optimization of data storage and exchange performance, and hierarchical protection of data security are realized, thereby constructing an efficient, secure, and easy-to-manage real-time data interaction environment.
[0006] A technical solution provided in an embodiment of the present invention is: a distributed real-time data exchange system based on Kafka, including: a management center, an execution area, and a Kafka cluster; The management center configures data exchange channels and generates data exchange tasks; The execution area includes multiple task executors, which respond to data exchange tasks according to the data exchange channels; The Kafka cluster partitions and stores data according to the concealment of the data, and responds to data production requests and data consumption requests issued by the task executors.
[0007] Preferably, the management center includes a data bus management module, a data type management module, an execution node management module, an application management module, a monitoring and alarm management module, and a data source management module; The bus management module is used to configure the basic information of the Kafka cluster and monitor the status of the Kafka cluster; The data type management module is used to determine the definition information of the exchanged or to-be-exchanged information, and the definition information at least includes data flow direction, data name, and Kafka topic; The application management module is used to configure data exchange channels and generate data exchange tasks, and the data exchange channels include a data collection channel and a data push channel; The execution node management module is used to monitor the task executors and distribute the execution data exchange tasks according to the running status of the task executors; The monitoring and alarm management module is used to obtain data exchange process data and respond to the alarm signals of the task executors; The data source management module includes an interactive protocol for configuring different data sources.
[0008] Preferably, the task executor includes a heartbeat monitoring module, a data access module, a data push module, a data processing module, and a monitoring and alarming module; The heartbeat monitoring module sends heartbeat information to the management center according to a preset period; The data access module matches the corresponding upstream application terminal and the data partition of the Kafka cluster according to the data exchange channel and the data exchange task; The data push module matches the corresponding downstream application terminal and the data partition of the Kafka cluster according to the data exchange channel and the data exchange task; The data processing module builds a data processing link relying on the Groovy dynamic scripting language; The monitoring and alarming module monitors the abnormal state of the data exchange process in real time and feeds it back to the management center.
[0009] Preferably, the heartbeat information includes at least CPU usage rate, memory usage rate, and running status information.
[0010] Preferably, the data access module includes a data production sub-module, and the data production sub-module stores data into the data partition of the corresponding Kafka topic according to the data production request issued by the task executor.
[0011] Preferably, the data push module includes a data consumption sub-module, and the data consumption sub-module extracts the data partition of the corresponding Kafka topic according to the data consumption request issued by the task executor.
[0012] Preferably, the Kafka cluster includes a number of data storage units and marks the Kafka topics of the corresponding data storage units according to the data production request; determines the corresponding partition configuration rules according to the data privacy of the Kafka topic; The Kafka cluster associates the data partition of the corresponding data storage unit according to the data consumption request, and extracts the data of the corresponding data partition according to the partition extraction rule.
[0013] Preferably, determining the corresponding partition configuration rules according to the data privacy of the Kafka topic includes the following steps: Associate a data privacy label for each Kafka topic through the data type management module and preset partition configuration rules for different privacy data. The data privacy labels include public level, sensitive level, and confidential level.
[0014] Preferably, the partition configuration rules include: When the data privacy label is at the public level, the public-level data is stored in the data partition in a partition polling manner; When the data privacy label is at the sensitive level, sensitive fields in the sensitive-level data are extracted through a Groovy script, a hash operation is performed on the sensitive fields to generate a partition key, and the sensitive-level data with the same partition key is stored in the same data partition; When the data privacy label is at the confidential level, the partition whitelist in the data bus management module is retrieved and a private channel is configured, and the confidential-level data is sent through the private channel to an independent physical storage unit, and the independent physical storage unit is statically encrypted.
[0015] Preferably, the partition extraction rule includes: When the data privacy label is at the public level, the task executor retrieves the data of the corresponding data partition according to the Kafka topic corresponding to the data consumption request; When the data privacy label is at the sensitive level, sensitive fields in the sensitive-level data are extracted through a Groovy script, a hash operation is performed on the sensitive fields to generate a partition key, and the task executor retrieves the data of the corresponding data partition according to the discrimination key; When the data privacy label is at the confidential level, authentication information for the confidential-level data is generated according to the partition whitelist, access rights and keys are determined according to the authentication information, and the task executor retrieves the encrypted data of the independent physical storage unit and decrypts it with the key.
[0016] Advantages of the present invention: (1) Aiming at the problems that the existing real-time data exchange between systems lacks a centralized unified management architecture and the exchange process is difficult to overall monitor and flexibly configure, in this application, the application management module of the management center configures the data exchange channel and generates a data exchange task, and the execution node management module monitors the task executor and distributes tasks according to its running state, realizing the overall management of the entire data exchange process, constructing a real-time data interaction environment convenient for management, and through the centralized management architecture, realizing the flexible configuration and comprehensive monitoring of the exchange process, and improving the efficiency and convenience of system management; (2) Aiming at the problems of large differences in data structures between different systems and high costs for data structure conversion and adaptation, in this application, the data processing module of the task executor relies on the Groovy dynamic scripting language to construct a data processing link, combined with the mechanism of Kafka cluster for partition storage according to data concealment, realizing flexible processing of data structures and reducing the development cost of real-time data exchange; through the coordination of the dynamic script to construct the processing link and the partition storage mechanism, efficiently solving the problem of heterogeneous data adaptation and improving the flexibility and development efficiency of data processing; (3) Aiming at the problem that there is a lack of a perfect security mechanism in the real-time data exchange process and the data is vulnerable to threats, this application sets different partition configuration rules and extraction rules according to the data privacy labels (public level, sensitive level, confidential level) through the Kafka cluster. For example, it performs a hash operation on sensitive-level data to generate a partition key for storage, and configures a private channel for confidential-level data and statically encrypts an independent physical storage unit, realizing hierarchical protection of data security and ensuring the security of data during the exchange process. By combining the hierarchical protection mechanism with the partition storage technology, a multi-level data security protection system is constructed, effectively resisting threats such as data theft and tampering, and improving the scientificity and reliability of data security protection.
[0017] The above invention content is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0019] Figure 1 It is a structural block diagram of the Kafka-based distributed real-time data exchange system of the present invention.
[0020] Figure 2 It is a schematic block diagram of data flow of the Kafka-based distributed real-time data exchange system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described here are only the best embodiments of the present invention, only for explaining the present invention, and do not limit the protection scope of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.
[0022] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the operations (or steps) as sequential processes, many of the operations (or steps) can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings; the process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0023] Embodiment 1: Figure 1 As shown, a Kafka-based distributed real-time data exchange system includes: a management center, an execution area, and a Kafka cluster; The management center configures data exchange channels and generates data exchange tasks; The execution area contains multiple task executors that respond to data exchange tasks according to the data exchange channels; The Kafka cluster stores data in partitions according to the concealment of the data and responds to the data production requests and data consumption requests issued by the task executors.
[0024] In this embodiment, a collaborative mechanism is adopted in which the management center centrally configures data exchange channels and generates tasks, multiple task executors in the execution area respond to tasks distributively based on the data exchange channels, the Kafka cluster stores data dynamically in partitions according to data concealment and responds to data production and consumption requests, realizing the overall management of the data exchange process, the real-time access, processing, and pushing of data and status monitoring, as well as the optimization of data storage and exchange performance and security hierarchical protection, and constructing an efficient, secure, and easy-to-manage real-time data interaction environment.
[0025] As an alternative embodiment, as Figure 2 shown, the management center consists of a data bus management module, a data type management module, an execution node management module, an application management module, a monitoring and alarm management module, and a data source management module; the bus management module is used to configure the basic information of the Kafka cluster and monitor the status of the Kafka cluster; the data type management module is used to determine the definition information of the exchanged or to-be-exchanged information, and the definition information at least includes the data flow direction, data name, and Kafka topic; the application management module is used to configure data exchange channels and generate data exchange tasks, and the data exchange channels include a data collection channel and a data push channel; the execution node management module is used to monitor the task executors and distribute the execution data exchange tasks according to the running status of the task executors; the monitoring and alarm management module is used to obtain data exchange process data and respond to the alarm signals of the task executors; the data source management module includes an interactive protocol for configuring different data sources.
[0026] It is understandable that in this embodiment, a multi-level modular collaborative mechanism is adopted in which the data bus management module in the management center configures and monitors the Kafka cluster, the data type management module defines the data flow direction and topics, the application management module configures channel generation tasks, the execution node management module distributes tasks according to the executor status, the monitoring and alarm management module obtains data and responds to alarms, and the data source management module configures multi-source interaction protocols. This realizes the centralized control of the basic information and running status of the Kafka cluster, the standardized definition and flexible scheduling of the entire data exchange process, the adaptive adaptation of multi-source heterogeneous system interaction protocols, and the real-time perception and response to abnormal states during the data exchange process, constructing an intelligent management system covering the entire life cycle of data exchange, effectively improving the collaborative efficiency of multi-system data exchange and the system running stability in complex business scenarios.
[0027] As an alternative embodiment, the task executor consists of a heartbeat monitoring module, a data access module, a data push module, a data processing module, and a monitoring and alarm module; the heartbeat monitoring module sends heartbeat information to the management center according to a preset period; the data access module matches the corresponding upstream application terminal and the data partition of the Kafka cluster according to the data exchange channel and the data exchange task; the data push module matches the corresponding downstream application terminal and the data partition of the Kafka cluster according to the data exchange channel and the data exchange task; the data processing module constructs a data processing link relying on the Groovy dynamic scripting language; the monitoring and alarm module monitors the abnormal state of the data exchange process in real time and feeds it back to the management center.
[0028] It is understandable that in this embodiment, a multi-module collaborative mechanism is adopted in which the heartbeat monitoring module in the task executor feedbacks operation status information such as CPU usage rate and memory usage rate to the management center according to a preset period to ensure system stability, the data access and push modules dynamically match the upstream and downstream application terminals and Kafka data partitions based on the data exchange channel and tasks to achieve real-time data transfer, the data processing module flexibly constructs a data processing link relying on the Groovy dynamic scripting language to adapt to multi-source heterogeneous data structures, and the monitoring and alarm module captures data exchange abnormalities in real time and feeds them back to the management center. This realizes the full-process automatic control of data from access, processing to pushing, the real-time perception of the running status of execution nodes, and the dynamic adaptation of heterogeneous data structures, constructing an intelligent data exchange execution system with self-monitoring and self-adaptation capabilities, effectively improving the real-time performance, flexibility and system robustness of data exchange in complex business scenarios.
[0029] As an alternative embodiment, the heartbeat information includes at least CPU usage rate, memory usage rate, and running status information.
[0030] It can be understood that in this embodiment, the heartbeat monitoring module of the task executor is used to send heartbeat data including CPU usage rate, memory usage rate, and running status information to the management center at a preset period, achieving real-time and accurate perception of the resource load (such as computing resource occupancy) and health status of the execution node by the management center; by digitalizing and visualizing the underlying hardware operation parameters and system operation status, a dynamic monitoring system for the execution node status is constructed, enabling the management center to optimize the task distribution strategy, predict potential failures, and trigger the warning mechanism based on real-time status data, improving the stability and manageability of the execution layer in the distributed data exchange system, and providing an underlying guarantee for continuous and reliable operation in complex business scenarios.
[0031] As an alternative embodiment, the data access module includes a data production sub-module, and the data production sub-module stores data into the data partition of the corresponding Kafka topic according to the data production request issued by the task executor.
[0032] It can be understood that in this embodiment, the data production sub-module of the data access module responds to the data production request of the task executor and stores the data into the data partition according to the corresponding Kafka topic, realizing the real-time and orderly convergence of multi-source heterogeneous data from the upstream application terminal to the Kafka cluster; taking the data exchange scenario of enterprise CRM, ERP and other systems as an example, this mechanism can accurately map the data generated by different systems to the corresponding partitions of Kafka according to business themes (such as customer data, order data), construct a standardized data access process, avoid data chaos and redundant storage, improve the real-time and standardization of data collection in the distributed system, provide an efficient and reliable data underlying support for subsequent data processing, pushing and cross-system business collaboration, and enhance the compatibility and scalability of the system for complex business scenarios.
[0033] As an alternative embodiment, the data push module includes a data consumption sub-module, and the data consumption sub-module extracts the data partition of the corresponding Kafka topic according to the data consumption request issued by the task executor.
[0034] It can be understood that in this embodiment, the data consumption sub-module of the data push module is adopted to respond to the data consumption request of the task executor and extract data from the corresponding Kafka topic data partition, realizing the accurate and real-time distribution of the processed data from the Kafka cluster to the downstream application terminal. Taking the data interaction scenario between the enterprise production monitoring system and the ERP system as an example, this mechanism can quickly retrieve the corresponding partition data from Kafka according to the real-time requirements of the ERP system by business theme (such as production progress data, equipment status data) and push it to the target system, avoiding the impact of data delay on the business process, building a low-latency and highly reliable data distribution link, improving the real-time performance of data flow and business collaboration efficiency in the distributed system, providing efficient data support for cross-system real-time business scenarios (such as dynamic inventory management, production scheduling), and enhancing the system's response ability and scalability to complex business requirements.
[0035] As an alternative embodiment, the Kafka cluster includes a number of data storage units and marks the Kafka topics of the corresponding data storage units according to the data production request; determines the corresponding partition configuration rules according to the data privacy of the Kafka topic; The Kafka cluster associates the data partitions of the corresponding data storage units according to the data consumption request and extracts the data of the corresponding data partitions according to the partition extraction rules.
[0036] It can be understood that in this embodiment, the Kafka cluster is adopted to mark the Kafka topics of the data storage units through the data production request, dynamically match the partition configuration rules according to the data privacy of the topic (such as public level, sensitive level, confidential level) (such as round-robin storage, hash partitioning, encrypted storage in a private channel), and associate the data partitions according to the partition extraction rules (such as topic retrieval, hash matching, authentication and decryption) during data consumption, realizing the themed classification management of multi-source data in the distributed storage environment, the differential security protection and efficient access of data with different sensitive levels. Taking the enterprise cross-system data exchange scenario as an example, this mechanism can evenly store and quickly retrieve the public-level business data (such as product information) through round-robin partitioning, use hash partitioning for sensitive-level customer data to ensure centralized management of similar data and prevent leakage, and encrypt and store the confidential-level financial data in an independent physical unit through a private channel, building a dynamic security policy system covering the entire process of data storage and access. It not only improves the data retrieval and transfer efficiency, but also effectively resists security threats such as data theft and tampering through the hierarchical protection mechanism, providing intelligent and refined technical support for the security and compliance management of data assets in complex business scenarios.
[0037] As an alternative embodiment, determining the corresponding partition configuration rules according to the data privacy of the Kafka topic includes the following steps: The data type management module associates data privacy labels with each Kafka topic and presets partition configuration rules for different privacy data. The data privacy labels include public level, sensitive level, and confidential level.
[0038] It can be understood that this embodiment adopts the technical means of using the data type management module to associate public level, sensitive level, and confidential level data privacy labels with Kafka topics and preset partition configuration rules, realizing refined classification management and differential security policy adaptation of multi-source data in a distributed storage environment. Taking the scenario of cross-departmental data exchange in an enterprise as an example, this mechanism can evenly store public level product information data in data partitions through preset polling rules to improve retrieval efficiency, adopt preset hash partition rules for sensitive level customer contact information to achieve centralized control of similar data and reduce the risk of leakage, and start preset independent physical storage encryption rules for confidential level financial data to strengthen security isolation, constructing a full-chain dynamic security policy system from data classification, storage to access. It not only meets the requirements of the business scenario for data transfer efficiency, but also improves the pertinence and compliance of data security protection through the labeled hierarchical protection mechanism, providing a standardized technical framework for refined management of data assets and secure and compliant operation in a complex business environment.
[0039] As an alternative embodiment, the partition configuration rules include: When the data privacy label is public level, the public level data is stored in the data partition by means of partition polling; When the data privacy label is sensitive level, sensitive fields in the sensitive level data are extracted through a Groovy script, a hash operation is performed on the sensitive fields to generate a partition key, and the sensitive level data with the same partition key is stored in the same data partition; When the data privacy label is confidential level, the partition white list in the data bus management module is retrieved and a private channel is configured, and the confidential level data is sent to an independent physical storage unit through the private channel, and the independent physical storage unit is statically encrypted.
[0040] It can be understood that this embodiment adopts the technical means of implementing differential partition configuration rules for different data privacy labels (public level, sensitive level, confidential level), that is, achieving balanced storage and fast access for public level data through partition polling, using a Groovy script to extract sensitive fields and hash to generate a partition key for centralized control of similar data for sensitive level data, and retrieving the partition white list to configure a private channel and encrypt the independent physical storage unit for confidential level data, realizing refined management and security protection of multi-level data in a distributed storage environment.
[0041] A specific example is as follows: When there is a set of publicly available data (such as wealth management product announcement information), which mainly includes non-sensitive information such as the publicly available wealth management product names, yields, and terms released by the bank APP, it needs to be quickly provided for users to browse and retrieve.
[0042] Adopt a partition round-robin mechanism: The Kafka cluster implements partition round-robin through the built-in RoundRobinPartitioner, and distributes the publicly available data to different data partitions in the order of message production. For example, when there are 10 partitions, the first piece of data is stored in partition 0, the second piece is stored in partition 1, and so on in a cyclic distribution. Ensure that the data is evenly distributed among the partitions, avoid overloading a single partition, and improve the response speed when the front-end APP calls the data (for example, when the user refreshes the page, multiple partition data can be read in parallel to shorten the loading time).
[0043] When there is a set of sensitive data (such as customer ID numbers and mobile phone numbers); mainly including personal information such as ID numbers and mobile phone numbers submitted during user registration, it needs to be stored in compliance and be convenient for the business system to query according to the user dimension.
[0044] Extract sensitive fields through Groovy scripts: In the data processing module of the task executor, parse the raw data (such as JSON-formatted user registration information) through Groovy scripts, and use regular expressions or field mapping rules to extract sensitive fields.
[0045] For example, the Groovy script is as follows: Def data = json.parse(rawData) def sensitive Field = data.phoneNumber / / Extract the mobile phone number field.
[0046] Generate a partition key through hash operation: Process the extracted sensitive field (such as mobile phone number 138****1234) using the SHA-256 hash algorithm to generate a fixed-length hash value (such as e3b0c····), which is used as the Kafka partition key. Kafka determines the target partition for data storage based on the hash value of the partition key modulo the total number of partitions (for example, when the total number of partitions is 10, hashValue%10 calculates the partition index).
[0047] It should be noted that data with the same mobile phone number is always stored in the same partition, which is convenient for the customer management system to quickly query all the associated records of this user (such as historical transactions, account changes, etc.); the hash desensitization processing makes the sensitive field irreversibly converted into the original value, and even if the partition data is leaked, attackers cannot reverse the real mobile phone number through the hash value, reducing the risk of data leakage.
[0048] When there is a set of confidential data (such as account balance, transaction password); it mainly includes core sensitive information such as user account balance, payment password, etc., and needs to meet the regulatory requirements of physical isolation and encrypted storage.
[0049] First, configure a private channel: Through the partition whitelist mechanism of the data bus management module, a separate SSL / TLS encrypted transmission channel (such as a dedicated VPN link) is configured for confidential data to ensure that the data is encrypted throughout the transmission process from the task executor to the Kafka cluster. The whitelist only contains the IP addresses of authorized financial systems, and devices outside the whitelist cannot access this channel.
[0050] Secondly, set up an independent physical storage unit and perform static encryption on it: Confidential data is stored in an independent SSD disk array or a dedicated server, physically separated from the storage hardware of public-level and sensitive-level data (such as a separate cabinet or data center area). Before the data is written into the storage unit, the data is encrypted using the AES-256 symmetric encryption algorithm. The encryption key is generated by the key management module of the management center and stored in the hardware security module (HSM). The access rights of the storage unit are controlled by the whitelist. Only authorized users (such as financial department personnel) can obtain the key from the HSM to decrypt the data after passing multi-factor authentication (such as USBKey + password).
[0051] Through the above technical means, the risk of cross-partition data leakage caused by shared storage is avoided through physical isolation. The combination of static encryption and whitelist authentication ensures that even if the storage medium is stolen, unauthorized persons cannot decrypt the data, thus constructing a triple security protection system of "transmission encryption + physical isolation + key control". In particular, this application can effectively support the compliance data exchange and storage requirements in scenarios sensitive to data security such as finance and healthcare, while ensuring the real-time performance and processing efficiency of the business system.
[0052] As an alternative embodiment, the partition extraction rules include: When the data privacy label is public-level, the task executor retrieves the data of the corresponding data partition according to the Kafka topic corresponding to the data consumption request; When the data privacy label is sensitive-level, sensitive fields in the sensitive level are extracted through a Groovy script, a hash operation is performed on the sensitive fields to generate a partition key, and the task executor retrieves the data of the corresponding data partition according to the discrimination key; When the data privacy label is confidential-level, authentication information for confidential data is generated according to the partition whitelist, access rights and keys are determined according to the authentication information, and the task executor retrieves the encrypted data of the independent physical storage unit and decrypts it with the key.
[0053] It can be understood that in this embodiment, a technical means of implementing differential partition extraction rules for different data privacy labels (public level, sensitive level, confidential level) is adopted, that is, for public-level data, the corresponding partition is directly retrieved based on the Kafka topic to achieve fast access; for sensitive-level data, sensitive fields are extracted through Groovy scripts and hashed to generate a partition key to accurately locate the target partition; for confidential-level data, authentication information is generated based on the partition whitelist, and the mechanism of encrypting the data in the independent physical storage unit by combining access rights and keys is used to achieve intelligent retrieval and secure and controllable access to multi-level data in a distributed storage environment.
[0054] A specific example is as follows: It should be particularly noted that the scenario data adopted in this embodiment is the same as the partition configuration rules; The extraction method for public-level data is as follows: When the task executor receives a data consumption request, it first parses the Kafka topic (such as "public_finance_products") in the request. Directly associate the corresponding partition list (such as 10 partitions: partition0 to partition9) in the Kafka cluster according to the topic. Without additional calculation or encryption processing, directly read data from all partitions according to the topic. Adopt a parallel reading strategy: the task executor pulls data from multiple partitions simultaneously to improve the reading efficiency. For example, when the user refreshes the APP page, the public data in partitions such as partition0, partition2, and partition4 can be read simultaneously to shorten the page loading time. When the data is returned, the results of each partition are automatically merged to present a complete list of public information.
[0055] The extraction method for sensitive-level data is as follows: When the task executor receives a consumption request (such as querying the transaction records of a certain user), it extracts sensitive fields (such as the mobile phone number "138****1234") from the request. Use the same Groovy script as when storing to parse the data and extract the target sensitive fields. For example, the Groovy script is as follows: Def request Data = json.parse(requestBody) Def target Phone = requestData.targetPhone / / Extract the mobile phone number in the request The extracted mobile phone numbers are hashed using SHA-256 to generate the same hash value as when stored (such as e3b0c44···). The partition index is determined by the formula hashValue % total number of partitions (for example, when the total number of partitions is 10, hashValue % 10), accurately locating the partition where the data is located (such as partition3). After reading the data from the target partition, the task executor automatically performs desensitization processing on sensitive fields (such as displaying the mobile phone number as "138****1234") to ensure that the returned data meets the privacy protection requirements; if other user information needs to be associated (such as historical transaction records), since data with the same mobile phone number is stored in the same partition, all associated records in this partition can be quickly read in batches.
[0056] The extraction method for confidential-level data is as follows: When the task executor initiates a consumption request, it first verifies whether its own IP address is in the white list of confidential-level data (such as only allowing access from the IP addresses of the finance department) to the data bus management module of the management center. If the IP passes the white list verification, it further requires the user to provide a USBKey and password for identity verification to ensure legal operation permissions. The encryption key for confidential-level data is generated by the key management module of the management center and stored in the Hardware Security Module (HSM), and it is prohibited to store it in plain text on a general server. After authentication, the task executor applies for a temporary key from the HSM, and this key is only valid during the validity period of this session. Encrypted data is read from an independent physical storage unit (such as a dedicated SSD disk array) and decrypted using the obtained AES-256 key to ensure that the data is in plain text when processed in memory, and the key and plain text data are immediately cleared after processing. Data transmission is completed through an SSL / TLS encrypted channel (such as a dedicated VPN link) to prevent it from being stolen or tampered with midway; the independent physical storage unit is completely isolated from the storage hardware for public-level and sensitive-level data (such as a separate cabinet) to avoid the risk of cross-partition data leakage.
[0057] Through the above technical implementation, the system can automatically match the optimal extraction strategy according to the data privacy level, improving the exchange efficiency while ensuring data security, and is applicable to complex business scenarios sensitive to data such as finance and healthcare.
[0058] The specific implementation manners described above are the preferred implementation manners of the Kafka-based distributed real-time data exchange system of the present invention, and do not limit the specific scope of the present invention. The scope of the present invention includes but is not limited to this specific implementation manner. All equivalent changes made according to the shape and structure of the present invention are within the protection scope of the present invention.
Claims
1. A distributed real-time data exchange system based on Kafka, characterized in that It includes: a management center, an execution area, and a Kafka cluster; The management center configures a data exchange channel and generates a data exchange task; The execution area contains multiple task executors, which respond to data exchange tasks according to the data exchange channel; The Kafka cluster stores data in partitions according to the concealment of the data, and responds to the data production request and data consumption request sent by the task executor.
2. The Kafka-based distributed real-time data exchange system according to claim 1, characterized in that The management center includes a data bus management module, a data type management module, an execution node management module, an application management module, a monitoring and alarm management module, and a data source management module; The bus management module is used to configure the basic information of the Kafka cluster and monitor the status of the Kafka cluster; The data type management module is used to determine the definition information of the exchanged or to-be-exchanged information, and the definition information at least includes the data flow direction, the data name, and the Kafka topic; The application management module is used to configure a data exchange channel and generate a data exchange task, and the data exchange channel includes a data collection channel and a data push channel; The execution node management module is used to monitor the task executor and distribute the execution data exchange task according to the running status of the task executor; The monitoring and alarm management module is used to obtain the data exchange process data and respond to the alarm signal of the task executor; The data source management module includes an interaction protocol for configuring different data sources.
3. The Kafka-based distributed real-time data exchange system according to claim 1, characterized in that The task executor includes a heartbeat monitoring module, a data access module, a data push module, a data processing module, and a monitoring and alarm module; The heartbeat monitoring module sends heartbeat information to the management center according to a preset period; The data access module matches the corresponding upstream application terminal and the data partition of the Kafka cluster according to the data exchange channel and the data exchange task; The data push module matches the corresponding downstream application terminal and the data partition of the Kafka cluster according to the data exchange channel and the data exchange task; The data processing module builds a data processing link relying on the Groovy dynamic scripting language; The monitoring and alarm module monitors the abnormal status of the data exchange process in real time and feeds it back to the management center.
4. The Kafka-based distributed real-time data exchange system according to claim 1, characterized in that The heartbeat information at least includes CPU usage, memory usage, and running status information.
5. The Kafka-based distributed real-time data exchange system according to claim 3, characterized in that The data access module includes a data production sub-module, and the data production sub-module stores the data into the data partition of the corresponding Kafka topic according to the data production request sent by the task executor.
6. The Kafka-based distributed real-time data exchange system according to claim 3, characterized in that The data push module includes a data consumption sub-module, and the data consumption sub-module extracts the data partition of the corresponding Kafka topic according to the data consumption request sent by the task executor.
7. The Kafka-based distributed real-time data exchange system according to claim 1, characterized in that The Kafka cluster includes a number of data storage units and marks the Kafka topics of the corresponding data storage units according to the data production request; determines the corresponding partition configuration rules according to the data privacy of the Kafka topic; The Kafka cluster associates the data partition of the corresponding data storage unit according to the data consumption request, and extracts the data of the corresponding data partition according to the partition extraction rule.
8. The Kafka-based distributed real-time data exchange system according to claim 7, characterized in that Determining the corresponding partition configuration rules according to the data privacy corresponding to the Kafka topic includes the following steps: Associate a data privacy label for each Kafka topic through the data type management module and preset partition configuration rules for different privacy data. The data privacy labels include public level, sensitive level, and confidential level.
9. The Kafka-based distributed real-time data exchange system according to claim 7, characterized in that The partition configuration rules include: When the data privacy label is at the public level, the public-level data is stored in the data partition in a partition polling manner; When the data privacy label is at the sensitive level, extract the sensitive fields in the sensitive-level data through a Groovy script, perform a hash operation on the sensitive fields to generate a partition key, and store the sensitive-level data with the same partition key in the same data partition; When the data privacy label is at the confidential level, retrieve the partition white list in the data bus management module and configure a private channel, send the confidential-level data to an independent physical storage unit through the private channel, and statically encrypt the independent physical storage unit.
10. The Kafka-based distributed real-time data exchange system according to claim 7, characterized in that The partition extraction rules include: When the data privacy label is at the public level, the task executor retrieves the data of the corresponding data partition according to the Kafka topic corresponding to the data consumption request; When the data privacy label is at the sensitive level, extract the sensitive fields in the sensitive-level data through a Groovy script, perform a hash operation on the sensitive fields to generate a partition key, and the task executor retrieves the data of the corresponding data partition according to the discrimination key; When the data privacy label is at the confidential level, generate authentication information for the confidential-level data according to the partition white list, determine the access permission and key according to the authentication information, and the task executor retrieves the encrypted data of the independent physical storage unit and decrypts it with the key.