Kafka multi-cluster switching method and device, medium and program product
By real-time monitoring of the Kafka cluster's task indicators and topic associations, anomalies can be quickly identified and switched to the backup cluster, solving the high availability problem of the Kafka cluster in weak environments and achieving the continuity of data sending tasks and system stability.
Patent Information
- Application Number
- CN202510852341.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-12
AI Technical Summary
The existing Kafka cluster cannot switch to the backup cluster in a timely manner under weak environments, resulting in message backlogs, failing to meet high availability requirements, and affecting the operating efficiency of upstream senders.
By collecting the association between data types and topics in the Kafka resource pool in real time, identifying abnormal topics, and switching data sending tasks to the standby cluster for execution, combined with detection information and thread pool reallocation, the continuity and high availability of data sending tasks are ensured.
Quickly identify anomalies in the primary cluster and switch to the backup cluster in a timely manner, improving the high availability of the Kafka cluster in weak environments and ensuring the continuity of data sending tasks and system stability.
Smart Images

Figure CN120639581A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of dynamic management of Kafka resources, and in particular to a Kafka multi-cluster switching method, device, medium, and program product. Background Art
[0002] In today's digital age, distributed message queue systems have emerged to achieve efficient data transmission and processing to ensure the stable operation of various platform systems. Among them, Kafka clusters can effectively realize asynchronous data transmission, decoupling data producers and consumers and improving data processing capabilities.
[0003] Currently, Kafka cluster disaster recovery solutions primarily focus on failover within a single cluster. For example, by monitoring the activity of the leader within the Kafka cluster, such as heartbeat timeouts and ISR (In-Sync Replicas) shrinkage, failover is triggered and a new leader is automatically elected from followers to take over related operations, ensuring the availability of the Kafka cluster.
[0004] However, when the main Kafka cluster is in a weak environment, such as poor network conditions or excessive server load, there is a lag in the anomaly detection of internal monitoring indicators (such as heartbeat timeout and ISR contraction). As a result, in the early stage of performance degradation of the main cluster, the message flow cannot be quickly switched to the backup cluster. A large number of messages are accumulated in the main cluster and cannot be sent in time, which in turn causes the platform to respond to upstream senders lagging, ultimately slowing down the operating efficiency of upstream senders, failing to meet the high availability requirements of the Kafka cluster, and making it difficult to ensure the stable and efficient operation of the Kafka cluster. Summary of the Invention
[0005] The main purpose of this application is to provide a Kafka multi-cluster switching method, device, medium and program product, aiming to improve the high availability of systems deploying Kafka clusters in weak environments.
[0006] To achieve the above objectives, this application proposes a Kafka multi-cluster switching method, which includes:
[0007] Collect task indicators of data sending tasks corresponding to each data type in real time, where the data type is associated with each topic of the primary cluster in the Kafka resource pool, and the Kafka resource pool has at least one primary cluster and its corresponding backup cluster;
[0008] The topic corresponding to the data type whose task indicator is within the preset abnormal range is determined as the target topic;
[0009] In the case that there is at least one target topic in the primary cluster, the data sending task in the primary cluster is switched to the standby cluster for execution.
[0010] In one embodiment, after the step of switching the data sending task in the active cluster to the standby cluster for execution, the method further includes:
[0011] Sending a probe message to the target topic;
[0012] In the case that the detection information is sent successfully and the sending time is less than a preset time threshold, the data sending task is switched to the main cluster for execution.
[0013] In one embodiment, after the step of switching the data sending task in the active cluster to the standby cluster for execution, the method further includes:
[0014] In the case that the standby cluster has at least one target topic, a new cluster is registered in the Kafka resource pool, and the data sending task is switched to the new cluster for execution.
[0015] In one embodiment, the method further comprises:
[0016] In the event that the Kafka resource pool is unable to provide services, the message data received from the upstream sender is sent to the downstream receiver in real time, and the message data is stored in a third-party database;
[0017] In the event of a real-time transmission failure, the message data is extracted from the third-party database and resent to the downstream recipient in batches according to a preset resending strategy.
[0018] In one embodiment, after the step of resending the message data to the downstream recipient in batches, the method further includes:
[0019] According to the replenishment request of the receiver, the replenishment data of the message data is extracted from the third-party database, and the replenishment data is sent to the downstream receiver.
[0020] In one embodiment, the data sending task includes at least one subtask corresponding to each service code. After the step of switching the data sending task in the active cluster to the standby cluster for execution, the step further includes:
[0021] Collect the time taken to execute the subtasks corresponding to each business code;
[0022] In the case where the number of business codes whose task execution time exceeds the preset time threshold is greater than the preset number threshold, determining the priority of each business code according to the task execution time corresponding to each business code;
[0023] According to each priority, the target business code is determined in sequence, and the thread pool is rebound for the subtask corresponding to the target business code.
[0024] In one embodiment, the step of binding a thread pool for the subtask corresponding to the target business code includes:
[0025] If there is a corresponding dedicated thread pool for the target business code, bind the dedicated thread pool to the subtask corresponding to the target business code;
[0026] If there is no dedicated thread pool corresponding to the target business code, retrieve the target thread pool with the highest priority and remaining quota from the preset thread pool library, and determine whether the remaining quota of the target thread pool is greater than the number of transactions processed per second of the target business code;
[0027] When the remaining quota of the target thread pool is greater than the number of transactions processed per second, binding the subtask corresponding to the target business code to the target thread pool;
[0028] When the remaining quota of the target thread pool is less than or equal to the number of transactions processed per second, the target thread pool with the highest priority and remaining quota is taken out from the remaining thread pools in the thread pool library, and the step of determining whether the remaining quota of the target thread pool is greater than the number of transactions processed per second of the target business code is performed.
[0029] In addition, to achieve the above objectives, the present application also proposes a Kafka multi-cluster switching device, which includes:
[0030] A collection module is used to collect task indicators of data sending tasks corresponding to various data types in real time, wherein the data types are associated with various topics of the primary cluster in the Kafka resource pool, and the Kafka resource pool has at least one primary cluster and its corresponding backup cluster;
[0031] A determination module is used to determine the topic corresponding to the data type whose task indicator is within a preset abnormal range as the target topic;
[0032] The switching module is configured to switch the data sending task in the primary cluster to the standby cluster for execution when there is at least one target topic in the primary cluster.
[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the Kafka multi-cluster switching method described above.
[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the Kafka multi-cluster switching method described above are implemented.
[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the Kafka multi-cluster switching method as described above.
[0036] One or more technical solutions proposed in this application have at least the following technical effects: first, task indicators of data sending tasks corresponding to each data type are collected in real time to understand the running status of data sending tasks of different data types in the main cluster. Since the data type is associated with the topics of the main cluster in the Kafka resource pool, the running status of each topic in the main cluster can be monitored to quickly identify abnormalities in the main cluster; then, the topic corresponding to the data type whose task indicator is within the preset abnormal range is determined as the target topic, and the target topic with an abnormal number of sending task failures is quickly identified through quantitative indicators, so that it can be quickly determined whether the cluster has an abnormality, reducing the impact of the weak environment on the abnormal response judgment of the main cluster; then, when there is at least one target topic in the main cluster, the data sending task in the main cluster is switched to the corresponding standby cluster for execution, ensuring the normal operation of the Kafka cluster as a whole and enhancing the availability of the Kafka cluster. This application collects task indicators related to Kafka cluster and data sending tasks in real time, and judges the health status of the current main cluster from external indicators of Kafka cluster. Even in weak environments, it can quickly identify abnormal responses of the main cluster and perform master-slave switching, thereby improving the high availability of the system deployed with Kafka cluster in weak environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0039] Figure 1 A flowchart of the first embodiment of the Kafka multi-cluster switching method provided in this application;
[0040] Figure 2 A schematic diagram of the framework for Kafka cluster pooling management provided in Example 1 of this application;
[0041] Figure 3 Schematic diagram of the overall framework of Kafka multi-cluster switching provided in Example 1 of this application;
[0042] Figure 4 A schematic diagram of a cache degradation process using a third-party database provided in Example 2 of the present application;
[0043] Figure 5 A schematic diagram of the thread pool reallocation process provided in Example 3 of the present application;
[0044] Figure 6 A schematic diagram of a processing framework for dynamic thread pool expansion provided in Example 3 of the present application;
[0045] Figure 7 A schematic diagram of the system framework for deploying a Kafka cluster provided in Example 3 of this application;
[0046] Figure 8 This is a schematic diagram of the module structure of the Kafka multi-cluster switching device according to an embodiment of the present application;
[0047] Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the Kafka multi-cluster switching method in the embodiment of the present application.
[0048] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0050] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0051] Because when the main Kafka cluster is in a weak environment such as poor network conditions and excessive server load, there is a lag in the anomaly detection of internal monitoring indicators (such as heartbeat timeout and ISR contraction). As a result, in the early stage of performance degradation of the main cluster, the message flow cannot be quickly switched to the backup cluster, and the requirements for system real-time and high availability cannot be met.
[0052] The present application provides a solution. First, the task indicators of the data sending tasks corresponding to each data type are collected in real time to understand the running status of the data sending tasks of different data types in the main cluster. Since the data types are associated with the topics of the main cluster in the Kafka resource pool, the running status of each topic in the main cluster can be monitored to quickly identify the abnormality of the main cluster. Then, the topic corresponding to the data type whose task indicator is in the preset abnormal range is determined as the target topic. The target topic with abnormal number of sending task failures is quickly identified through quantitative indicators, so that it can be quickly determined whether the cluster has abnormalities and reduce the impact of weak environments on the abnormal response judgment of the main cluster. Then, when there is at least one target topic in the main cluster, the data sending task in the main cluster is switched to the corresponding standby cluster for execution, ensuring the normal operation of the Kafka cluster as a whole and enhancing the availability of the Kafka cluster. The present application collects the task indicators related to the Kafka cluster and the data sending task in real time, and judges the health status of the current main cluster from the external indicators of the Kafka cluster. Even in weak environments, it can quickly identify the abnormal response of the main cluster and perform master-slave switching, thereby improving the high availability of the system deployed with the Kafka cluster in weak environments.
[0053] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, personal computer, mobile phone, messaging platform, etc., or an electronic device that can realize the above functions.
[0054] Optionally, the Kafka cluster can be deployed on a software system, thereby obtaining a Kafka communication system, and the Kafka communication system can be deployed on the above-mentioned computer service device. The following describes this embodiment and the following embodiments by taking the Kafka communication system as the execution body as an example.
[0055] Based on this, the embodiment of the present application provides a Kafka multi-cluster switching method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the Kafka multi-cluster switching method of this application.
[0056] In this embodiment, the Kafka multi-cluster switching method includes steps S10 to S30:
[0057] Step S10, collecting task indicators of data sending tasks corresponding to each data type in real time;
[0058] The data type is associated with each topic of the primary cluster in the Kafka resource pool. The Kafka resource pool contains at least one primary cluster and a corresponding backup cluster.
[0059] In one feasible embodiment, the Kafka communication system can quickly identify abnormal conditions in the primary cluster by collecting task metrics for data transmission tasks of different data types and leveraging the association between data types and topics in the primary cluster. Furthermore, the Kafka resource pool is configured with a primary cluster and a corresponding backup cluster to facilitate failover in the event of an anomaly in the primary cluster, ensuring the availability of the Kafka communication system.
[0060] Optionally, the task indicator of the data sending task refers to a set of parameters used to measure the execution status of the data sending task, which may be one or more data indicators such as the sending timeout amount, the number of failures, and the message backlog amount.
[0061] Optionally, during the data transmission process, connection points such as before transmission, after transmission and in case of exception can be added, and indicator collection enhancement can be woven into these connection points to mainly collect the above-mentioned task indicator data.
[0062] Optionally, the data type refers to a data classification identifier used to distinguish data streams generated by different business or functional modules, and can be identified by metadata tags or encoding. For example, in a message processing platform, the data type corresponding to the data stream generated by the authorization system is authorization data, and the data type corresponding to the data stream generated by the installment system is installment data.
[0063] Optionally, the association relationship between the data type and each topic in the main cluster can be a one-to-one correspondence, or one data type can correspond to multiple topics in the main cluster.
[0064] Optionally, a Kafka resource pool is a resource aggregation unit consisting of multiple Kafka clusters (including primary and backup clusters). It implements unified management and monitoring of Kafka resources through a unified interface, including registering and deregistering Kafka resources and collecting and analyzing relevant metrics. The primary cluster is the main cluster in the Kafka resource pool for normal data processing and storage, while the backup cluster is the cluster used for disaster recovery and backup. The primary and backup clusters maintain data synchronization, and if the primary cluster fails, the backup cluster can take over the primary cluster's responsibilities, ensuring high system availability. A topic is a logical unit in a Kafka cluster for storing and distributing data.
[0065] For example, please refer to Figure 2 , Figure 2 This paper provides a framework diagram for pooled management of Kafka clusters. Multiple Kafka clusters are connected through a unified interface to obtain a Kafka resource pool. Administrators can use this interface to uniformly manage the Kafka clusters in the resource pool, including registering or destroying modules such as producers, consumers, and topics in the Kafka cluster. Producers refer to the upstream senders (access applications) of message data in the Kafka cluster, and consumers refer to the downstream receivers (sending applications) of message data in the Kafka cluster. Different types of message data generated by producers are written to different topics. The association between data types and topics can be managed through this unified interface, and the message data that consumers can read (or the topics that consumers can access) can also be managed through this interface. Furthermore, when administrators manage Kafka resources through this interface, their configuration information changes are automatically synchronized to the Kafka communication system without the need for a restart to make the configuration information take effect. In addition, when using the cluster in the Kafka resource pool to perform data sending tasks, you can collect information on the data sending status of each producer, the data receiving status of each consumer, the task execution status of each topic, and other indicators in the Kafka resource pool, so as to promptly detect abnormal situations in the Kafka resource pool.
[0066] It is understandable that by collecting task indicators of various data types in real time and combining the association between data types and topics in the main cluster, the abnormal status of the main cluster can be quickly identified in a weak environment, thereby reducing the abnormal response time, so as to enable timely master-slave switching and improve the high availability of the Kafka communication system in a weak environment.
[0067] Step S20: Determine the topic corresponding to the data type whose task indicator is within a preset abnormal range as the target topic.
[0068] Optionally, different exception ranges may be set for different data types to meet data characteristics corresponding to different data types.
[0069] Optionally, since the abnormal causes of data sending tasks include cluster failure, network abnormality, resource contention, etc., after determining that there is a data type whose task indicators are within a preset abnormal range, the abnormal causes of the data sending tasks can be further attributed and analyzed in combination with the environmental perception algorithm to determine whether the abnormal cause of the current data sending task is a cluster failure, and the topic corresponding to the data type whose task indicators are within the abnormal range and the abnormal cause is a cluster failure is determined as the target topic.
[0070] For example, the environment perception algorithm can be used to monitor network delay, packet loss rate and other indicators in real time to determine the impact of network problems on message sending tasks, and then determine whether the current abnormality is caused by a network failure; the environment perception algorithm can also be used to monitor Kafka resource indicators such as the number of under-replicated partitions and the number of minimum ISR partitions in the Kafka cluster, and when the number of under-replicated partitions and / or the number of minimum ISR partitions is greater than zero, it is determined that the abnormality cause of the current data sending task is a cluster failure.
[0071] It is understandable that by setting the abnormal range, on the one hand, the impact of accidental network fluctuations on cluster failure judgment can be avoided. On the other hand, by marking the target topic, it is convenient to quickly determine whether the main cluster is in an abnormal state, and at the same time facilitate subsequent detection of the health status of the main cluster.
[0072] Step S30: When there is at least one target topic in the primary cluster, the data sending task in the primary cluster is switched to the standby cluster for execution.
[0073] Optionally, since the primary and backup clusters are managed through the unified interface of the Kafka resource pool, hot updates can be performed through the Kafka resource manager when switching between the primary and backup clusters without modifying environment variables and restarting to take effect, thereby reducing the steps and time consumption of primary and backup switching to ensure the stability of system operation.
[0074] Optionally, the specific implementation of switching the data sending task in the main cluster to the backup cluster for execution can be to switch all data sending tasks in the main cluster to the backup cluster for execution, or to switch the data sending tasks in the target topic to the corresponding topic in the backup cluster for execution. This embodiment does not impose any specific restrictions on this.
[0075] Optionally, when there is at least one target topic in the primary cluster, the system publishes a master-slave switch event; then, after the Kafka resource manager monitors the master-slave switch event, it automatically performs a master-slave switch, downgrading the primary cluster to a standby cluster and upgrading the standby cluster to the primary cluster to execute the data sending task.
[0076] Optionally, if the number of failures for each data type does not change within a preset time, it is reset to zero to prevent historical failure data from affecting the health status of the current active cluster and causing misjudgment.
[0077] For example, a data processing system uses a Kafka cluster to process business data such as orders, payments, and refunds. The primary cluster is deployed in the main data center, and the backup cluster is deployed in the disaster recovery center. The primary and backup clusters are in the same resource management pool (Kafka resource pool) and are managed through a unified interface. During a network fluctuation, the number of failures of order-type business data detected reached 25 times / minute, far exceeding its preset threshold of 8 times / minute; then the Kafka communication system marked the topic corresponding to the order type as the target topic and published a primary-backup switching event; then, after the backup cluster was in good condition, the data routing configuration of the Kafka resource pool was modified, and the data sending task currently being executed by the primary cluster was switched to the backup cluster for execution, and the data sending task originally sent to the primary cluster was redirected to the backup cluster.
[0078] It is understandable that by switching to the standby cluster, the continuity of the data sending task is ensured, thereby improving the high availability of the Kafka cluster in abnormal situations.
[0079] In a feasible implementation manner, after step S30, the method further includes:
[0080] Step S40, sending the detection information to the target topic;
[0081] Step S50: When the detection information is sent successfully and the sending time is less than a preset time threshold, the data sending task is switched to the main cluster for execution.
[0082] In a feasible embodiment, in order to detect whether the topic (target topic) of the original primary cluster corresponding to the data type that has undergone primary-backup switching has returned to normal, detection information can be sent to the target topic according to timing, real-time or other sending strategies, and the health status of the original primary cluster can be judged based on the detection results; when the detection information is successfully sent to the target topic and the time consumed is lower than the preset threshold, it indicates that the current target topic of the original primary cluster is available, the detection is determined to be passed, and the primary-backup switching event is published; after the Kafka resource manager monitors the primary-backup switching event, it automatically restores the original primary cluster to the primary cluster to perform the data sending task.
[0083] Optionally, the probe message can be a lightweight data packet with a specific tag and a fixed format, used to test the availability of topic blocks in Kafka.
[0084] Optionally, the detection information can be sent to each partition in the target topic. If the detection information is successfully sent to each partition in the target topic and the sending time is less than a preset time threshold, the detection is determined to be successful and the master-slave switching is performed.
[0085] In this embodiment, the recovery status of the primary cluster is determined by dynamically sending detection information, thereby achieving automatic switchback, ensuring that the data processing capability of the primary cluster is restored in a timely manner after it returns to normal, and improving the high availability of the system.
[0086] In a feasible implementation manner, after step S30, the method further includes:
[0087] In step A21, if the standby cluster has at least one target topic, a new cluster is registered in the Kafka resource pool, and the data sending task is switched to the new cluster for execution.
[0088] Optionally, after switching the data sending task to the standby cluster for execution, while sending detection information to the target topic in the primary cluster, it is also necessary to monitor the number of failures of the data sending tasks for each topic in the standby cluster in real time. When an abnormality is found in the standby cluster (that is, there is at least one topic marked as the target topic in the standby cluster), a new Kafka cluster is automatically registered in the Kafka resource pool and its data configuration is initialized, that is, the current data status of the primary cluster or the standby cluster is synchronized to the new cluster to ensure data consistency; after waiting for the new cluster to be registered and recognized by the applications in the system, the data sending task is switched from the current cluster to the new cluster for execution, realizing multi-level disaster recovery capabilities and ensuring that critical businesses can still operate normally under extreme circumstances.
[0089] In this implementation, when both the primary and backup clusters fail at the same time, a new cluster is dynamically registered and the data sending task is switched to the new cluster for execution, thus breaking through the upper limit of the disaster recovery capability of the traditional primary-backup architecture and improving the overall availability of the Kafka cluster.
[0090] Optionally, after collecting multiple task indicators corresponding to each data type, it is possible to further determine whether the weighted sum of the task indicators corresponding to each data type is within a preset abnormal range based on a preset Kafka health decision, wherein the Kafka health decision includes the weight corresponding to each task indicator.
[0091] For example, please refer to Figure 3 , Figure 3This paper provides a schematic diagram of the overall framework of Kafka multi-cluster switching. Assuming that Kafka master cluster A and Kafka backup cluster B are pre-registered in the Kafka resource pool, first obtain the relevant parameter configuration, such as the threshold parameters for determining the abnormal range, the data source parameters for determining the task indicator data collection object, and the Kafka resource parameters for determining the type of task indicator to be collected, so as to subsequently clarify the indicators to be collected and how to judge the abnormal situation of the Kafka cluster; then perform task management through the sending thread. The task management is specifically as follows: through the topic of the main cluster A, according to the priority of each subtask in each type of data sending task, execute each subtask in sequence, and collect Kafka indicators at the same time, including the amount of successful / failed sending, timeout amount, etc. Task metrics, Kafka resource metrics, and Kafka health metrics are used to assess the performance of the current data sending task, providing a macro-level assessment of the health of the primary cluster. Kafka resource metrics and Kafka health metrics are used to monitor the operational status of each partition within the primary cluster's topic, enabling timely election of new leaders for fault recovery within the topic. Furthermore, based on pre-defined Kafka health decisions, the weighted sum of task metrics for each data type is determined to be within the abnormal range to determine whether the metrics of the current primary cluster A are abnormal. If the metrics of the primary cluster A are abnormal, data sending tasks are switched to the backup Kafka cluster B. Kafka metrics from the primary Kafka cluster A are also collected regularly. Once the primary cluster A is confirmed to be recovered, data sending tasks are switched to the primary cluster A. Furthermore, if both the primary Kafka cluster A and the backup Kafka cluster B fail, a new Kafka cluster C can be dynamically registered through parameter configuration, and data sending tasks can be switched to the new cluster C.
[0092] This embodiment provides a Kafka multi-cluster switching method. By collecting and comparing task indicators of data sending tasks of various data types in real time, the health status of the current master cluster is judged from external indicators of the Kafka cluster. Therefore, even in a weak environment, cluster anomalies can be quickly identified and timely master-slave switching and other methods can be used to ensure the continuity of the execution of data sending tasks, thereby improving the high availability of systems deployed with Kafka clusters in weak environments.
[0093] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereafter. On this basis, the Kafka multi-cluster switching method also includes:
[0094] In step E10, when the Kafka resource pool is unable to provide services, the message data received from the upstream sender is sent to the downstream receiver in real time, and the message data is stored in a third-party database.
[0095] In one feasible embodiment, when the Kafka resource pool is unable to provide services, the system forwards received message data to the corresponding recipient in real time and also stores the message data in a third-party database. This ensures that even if the Kafka cluster is completely unavailable, the message data can still be delivered in a timely manner and stored persistently.
[0096] Alternatively, the Kafka resource pool cannot provide services when the primary cluster, the backup cluster, and the registration configuration-related functional modules cannot work properly due to network failure, hardware failure, or configuration error, and the entire Kafka cluster is unavailable.
[0097] Optionally, the upstream sender refers to a system or service that generates and sends message data; the downstream receiver refers to a downstream system or service that is responsible for responding to or processing the message data.
[0098] Optionally, a third-party database refers to an external persistent storage system independent of the main message processing process, which is used to store message data when Kafka resources are unavailable, such as Elasticsearch (ES), MySql (My Structured Query Language), Radis (Remote Dictionary Server), etc. ES and MySql can also be combined to use as a third-party database, combining full-text search and relational query capabilities to facilitate subsequent data replenishment.
[0099] Step E20: In the event of real-time transmission failure, the message data is extracted from a third-party database and resent to the downstream recipient in batches according to a preset resending strategy.
[0100] In one feasible embodiment, the operational status of sending message data to the downstream recipient is monitored in real time; when the operational status indicates that the real-time sending fails, the corresponding message data can be extracted from a third-party database, and according to a preset retransmission strategy, the message data can be divided into multiple task batches and gradually sent to the downstream recipient.
[0101] Optionally, the resending strategy is a set of preset rules that define how to resend a message after the real-time sending of the message data fails, and generally includes parameters such as a resending time interval, a resending batch size, and a number of retries.
[0102] Optionally, a machine learning model is trained based on the historical reissue data of the recipient's system, and the trained model is used to predict a reissue strategy suitable for the current recipient; then, during the batch reissue process, the recipient's response time and error rate are monitored in real time to dynamically adjust the reissue strategy.
[0103] Exemplarily, historical reissue data is collected, including the size of message data, response time, reissue time interval, reissue batch size, and number of reissues; and the historical reissue data is used as training data to train and optimize the preset machine learning model. During the training process, the reissue success rate (the ratio of successfully reissued message data to the total reissued message data) is used as the objective function, so that the model can learn how to set parameters such as the reissue time interval, reissue batch size, and number of reissues to maximize the reissue success rate when receiving the corresponding input of message data size and / or response time. After the training is completed, a strategy prediction model is obtained; and then, the size of the current message data is used as the input of the strategy prediction model, and its corresponding model output is a reissue strategy including the reissue time interval, reissue batch size, and number of reissues; and then, data is reissued according to the reissue strategy, and during the data reissue process, the response time of the recipient is monitored in real time, and the strategy prediction model and the response time can be further used to dynamically adjust the reissue strategy. For example, when the response time increases, the reissue batch size can be reduced and the reissue time interval can be increased.
[0104] Optionally, the message data from the upstream sender can be received in real time through the preset access cluster, the message data can be stored in a third-party database, and the message data can be forwarded to the downstream receiver in real time through the preset sending cluster; further, in the event that the real-time sending through the sending cluster fails, the message data can be extracted from the third-party database through the preset complement cluster and retransmission strategy, and retransmitted to the downstream receiver in batches.
[0105] For example, please refer to Figure 4 , Figure 4A schematic diagram of a cache degradation process using a third-party database is provided. The access cluster and the sending cluster are used to dynamically receive, process, and forward data, while the complement cluster monitors data transmission failures or loss events and retransmits messages in batches according to a pre-set retransmission strategy to ensure message integrity. Under normal circumstances, the access cluster (producer) receives incoming data from the sender and writes it to the Kafka cluster. The sending cluster (consumer) then reads the incoming data from the Kafka cluster and sends it to the receiver. However, if the Kafka resource pool fails (for example, if both the primary and backup Kafka clusters fail to write), the sending cluster retransmits real-time data while triggering an exception degradation mechanism. Data is written to a third-party database to prevent data loss. For example, an ES cluster is used as a cache cluster in the third-party database to cache frequently accessed data, while a MySQL server is used as a database cluster to store all incoming data. If real-time data retransmission from the sending cluster fails after N attempts, the complement cluster retrieves the data from the third-party database and retransmits it in batches to the receiver.
[0106] It is understandable that by dividing the message data into multiple batches and gradually resending them to the recipient, problems such as network congestion or excessive system pressure caused by sending a large number of messages at one time can be avoided, thereby improving the reliability of message retransmission and the overall system stability.
[0107] In a feasible implementation manner, after step E20, the method further includes:
[0108] Step E30: extracting the complementary data of the message data from the third-party database according to the complementary request of the receiver, and sending the complementary data to the downstream receiver.
[0109] In a feasible embodiment, when the business result determined by the recipient based on the message data is inconsistent with the expected result, the recipient can send a replenishment request to the system, wherein the replenishment request generally includes positioning information such as data identification, time range, and data sequence number interval, which is used to determine the data range that needs to be replenished; then, the system can search and extract the corresponding historical message data from the third-party database based on the replenishment request, and send it to the recipient as replenishment data, thereby realizing on-demand data repair and consistency assurance.
[0110] For example, the time range [T1, T2] can be determined based on the receiver's replenishment request; then, the message data within the time range is queried through ES and its serial number is determined; further, based on the serial number, the corresponding message data in MySql is extracted, determined as replenishment data, and sent to the receiver.
[0111] In this embodiment, a multi-level backup mechanism for all cluster failures is used to provide real-time retransmission and batch retransmission, thereby ensuring the integrity and consistency of message data. For example, messages that fail to be sent in real time are stored in es+mysql, and task batch retransmission is used to ensure that each message is not lost. A multi-dimensional historical message screening and retransmission strategy is also provided, and customized retransmission is performed based on retransmission requests. This supports scenarios where downstream receivers can retransmit data multiple times after receiving the data, providing a backup for the downstream, that is, ensuring the receiver's self-repair capability in the event of data inconsistency, thereby reducing the business risks caused by long-term data inconsistency.
[0112] Based on the first and / or second embodiments of the present application, in the third embodiment of the present application, the same or similar contents as those of the first and second embodiments can be referred to above and will not be described in detail. On this basis, the data sending task includes at least one subtask corresponding to each service code, and after step S30, further includes:
[0113] Step B10: collecting the execution time of the subtasks corresponding to each business code;
[0114] In one feasible embodiment, although the primary and backup clusters share the same topic structure and data layout, due to differences in infrastructure conditions such as allocated CPU resources, system load, and network environment, the thread pools originally bound to each subtask in the primary cluster may not maintain the same performance in the backup cluster environment. Therefore, after a primary-backup switchover event occurs, execution time data for each subtask corresponding to each business code can be collected to assess whether the existing thread pool configuration is still applicable and to rebind the thread pool if necessary.
[0115] Optionally, the business code is used to identify and distinguish different processing flows in a certain type of data sending task. Each type of business code may include multiple subtasks. For example, for a data sending task whose data type is installment data, it may include subtasks such as data reading, data verification, data sending, and receipt confirmation, and the business code for data reading and data verification may be determined as data preparation.
[0116] Optionally, the task execution time refers to the time it takes for a subtask to be executed from the start to the completion. In addition to the task execution time, the priority can also be evaluated by combining multiple indicators such as the resource consumption rate and task load corresponding to each business code.
[0117] Step B20: When the number of subtasks whose task execution time exceeds the preset time threshold is greater than the preset number threshold, the priority of each business code is determined according to the execution time of the tasks corresponding to each business code;
[0118] For example, the execution time of the subtasks corresponding to each business code is monitored in real time. When the number of timed-out subtasks reaches a preset threshold, the priority of the business code is automatically recalculated and assigned based on the actual execution time of the subtasks corresponding to each business code, so that system resources are allocated first to performance-sensitive key businesses, thereby ensuring the overall stability of the system under high load conditions.
[0119] Step B30: determine the target business code in sequence according to each priority, and rebind the thread pool for the subtask corresponding to the target business code.
[0120] Optionally, the target business code refers to a business code selected by the system for thread pool allocation after being sorted according to priority; the thread pool is a form of multi-threaded processing used to centralize subtasks of data sending tasks of various data types for unified scheduling and management.
[0121] For example, the first business code with the highest priority can be determined as the target business code in order from high to low priority, and the thread pool can be bound to the subtask corresponding to the target business code; then the first business code can be eliminated, and the second business code with the highest priority can be re-determined as the target business code, and the thread pool can be bound to it. The above steps are repeated until the subtasks corresponding to all business codes are bound.
[0122] Optionally, this thread pool dynamic binding mechanism is not only applicable to the situation of master-slave switching. During the normal operation of the primary cluster, backup cluster or new cluster performing data sending tasks, the Kafka communication system can also collect the task execution time of the subtask corresponding to each business code in real time, so as to evaluate the health status of the thread pool assigned to the Kafka cluster currently executing the task, and trigger thread pool reallocation when the task processing speed of each thread pool decreases, so as to improve the high availability and stability of the Kafka cluster.
[0123] In a feasible implementation manner, step B30 includes:
[0124] Step B31: If a corresponding dedicated thread pool exists for the target business code, bind the dedicated thread pool to the subtask corresponding to the target business code;
[0125] Step B32: If there is no dedicated thread pool corresponding to the business code, retrieve the target thread pool with the highest priority and remaining quota from the preset thread pool library, and determine whether the remaining quota of the target thread pool is greater than the number of transactions processed per second of the business code;
[0126] Optionally, a preset thread pool library refers to a set of pre-created and configured shared thread pool resources, typically with different performance characteristics and capacity parameters, for use by businesses that do not have dedicated thread pools.
[0127] Optionally, Transactions Per Second (TPS) represents the processing load indicator for all subtasks contained in the business code, indicating the number of transactions that need to be processed per unit time. Remaining quota, on the other hand, represents the remaining capacity of the thread pool available for new tasks, typically expressed in transactions per second, reflecting the level of resource availability.
[0128] Step B33: When the remaining quota of the target thread pool is greater than the number of transactions processed per second, the subtask corresponding to the target business code is bound to the target thread pool;
[0129] Step B34, when the remaining quota of the target thread pool is less than or equal to the number of transactions processed per second, take out the target thread pool with the highest priority and remaining quota from the remaining thread pools in the thread pool library, and execute the step of determining whether the remaining quota of the target thread pool is greater than the number of transactions processed per second of the target business code.
[0130] For example, the basic management and task execution functions of the thread pool can be implemented through the Executor interface and its related sub-interfaces and classes. Figure 5 , Figure 5A flowchart for thread pool reallocation is provided, which mainly includes four major processes: B100, initializing the thread pool; B200, collecting thread running data for each subtask; B300, dynamically monitoring whether thread reallocation is triggered; and B400, reallocating threads according to business code priority. Process B100 includes: first executing B101, calling the load method to load the preset thread pool configuration into a ThreadPoolConfig object, which contains all the configuration information of the thread pool. The load method is used to load the thread pool configuration information from the database; then executing B102, calling the register method, initializing the thread pool according to ThreadPoolConfig, and registering the initialized thread pool with the DtpRegistry. The register method is used to register a database executor instance (such as a thread pool) with a management component or container. The DtpRegistry is a registry for managing dynamic thread pools, so that the thread pool can be uniformly managed through the DtpRegistry. Then the process reaches B200: call the beroreExecute method to initialize the start time of the task (B201), and execute B202, call the run method to execute the data sending task; when the data sending task is completed, execute B203, call the afterExecute method to count the task execution time of each subtask, and classify and save according to the business code corresponding to each subtask. Then the process reaches B300: first execute B301 to determine whether the number of business codes that have timed out of the task execution time exceeds the preset number threshold; and if the number of timed-out business codes exceeds the threshold, execute B302, call the reordor method to sort the priority of the business codes from low to high according to the time consumed (task execution time); then execute B303, call the publish method to publish the thread pool reallocation event, where the publish method is used to publish the monitored operating status or indicator data.Finally, the process reaches B400, and the thread pool is allocated to each business code in order of priority: first execute B401 to determine whether the current business code is configured with an exclusive thread pool; if so, execute B402, call the register method to bind the current business code to its corresponding exclusive thread pool; if not, execute B403, call the getNextExecuter method to retrieve the remaining unallocated thread pool with the highest priority, where the getNextExecutor method is used to obtain the next available thread pool instance from the registered thread pool according to a certain strategy (such as priority, load, etc.); and further execute B404 to determine whether the remaining quota of the thread pool is greater than the TPS of the current business code; if the remaining quota is less than or equal to the business code TPS, exclude the thread pool and repeat the above B403; if the remaining quota is greater than the business code TPS, execute B405, call the register method to bind the current business code to the thread pool, and execute B406 to update the remaining quota of the thread pool.
[0131] Optionally, in B403 above, when the remaining quotas of other higher-priority thread pools are all less than or equal to the TPS of the current business code, the subtask corresponding to the target business code can be directly bound to the thread pool with the lowest priority.
[0132] Optionally, in the above B403, after the remaining quotas of all thread pools are less than or equal to the TPS of the current business code, the thread pool to be expanded can be determined based on the load of each thread pool and the expansion benefit ratio, where the expansion benefit ratio refers to the improvement in processing capacity that can be brought about by unit resource input (such as adding one thread); then, based on the TPS requirement of the current business code (target business code), the remaining quota and reserved running quota of the thread pool to be expanded, the quota to be expanded of the thread pool to be expanded is determined, where the reserved running quota is the amount reserved to ensure the normal operation of the thread pool to be expanded; finally, the thread pool to be expanded is expanded according to the quota to be expanded, that is, the quota to be expanded is increased on the basis of the original quota of the thread pool to be expanded.
[0133] For example, please refer to Figure 6 , Figure 6A schematic diagram of a processing framework for dynamic expansion of a thread pool is provided. Generally, by default, there are: a dedicated thread pool for ensuring that the transmission of high-risk data sources is not affected by other data sources, and a fast thread pool, a normal thread pool, and a slow thread pool adapted for data transmission tasks without a dedicated thread pool. The initial capacities of the dedicated thread pool, the fast thread pool, the normal thread pool, and the slow thread pool are p1, p2, p3, and p4, respectively. Initially, the fast, normal, and slow thread pools evenly distribute data transmission tasks that are not bound to a dedicated thread pool in a polling manner. The specific polling method is also based on the priority of the business code corresponding to each subtask. The thread pool is allocated to each subtask in the task queue to be processed. The thread pool allocation at time t1 can be referred to. As each subtask increases, the thread pool allocation situation at time t1 increases. As the number of services running and subtasks assigned to each thread pool increases, the processing speed of the fast thread pool and the slow thread pool slows down. At this time, the number of timed-out business codes exceeds the preset threshold, triggering thread reallocation. During the reallocation process, the business TPS corresponding to subtasks C2 and D2 is greater than the remaining quotas of all non-exclusive thread pools, triggering thread pool expansion. According to the above expansion method, the quota to be expanded p2' of the fast thread pool and the quota to be expanded p3' of the ordinary thread pool are determined in turn to ensure that each subtask can be executed normally. After the expansion is completed, the capacity of the fast thread pool is P2+P2', and the capacity of the ordinary thread pool is p3+p3'. Referring to the reallocation and dynamic expansion at time t2, the reallocation and dynamic expansion of the thread pool are completed.
[0134] In this embodiment, by reallocating the thread pool when abnormal tasks increase, the adaptability to load changes is improved and the stable operation capability in complex and changing environments is enhanced; at the same time, with regard to thread pool reallocation, by establishing a dedicated thread pool and matching thread pools of different levels according to the priority of the business code, differentiated thread guarantee strategies can be provided for businesses of different importance, thereby ensuring the normal operation of key businesses and the high availability of the system.
[0135] For example, to help understand the implementation process of the Kafka multi-cluster switching method obtained by combining the above-mentioned embodiment 1 and embodiment 2, please refer to Figure 7 , Figure 7A schematic diagram of the system framework for deploying a Kafka cluster is provided. First, it includes a mapping configuration module, in which cluster configuration, data type configuration, and producer / consumer configuration can be performed, and registration can be performed with each other, so that the Kafka resource pool can be registered through this module, and data producers (access applications) and consumers (sending applications) as well as the cluster topics (topics) and partitions corresponding to each data type can be clarified. For example, the access application can write installment data to topic-partition 1 of the Kafka primary cluster A in the Kafka resource pool, and the installment notification module in the sending application can also read the corresponding installment data from the topic-partition 1, generate installment notifications, and send them to the corresponding recipients. At the same time, the Kafka backup cluster B in the Kafka resource pool maintains data synchronization with the primary cluster A. In the event of a failure of the primary cluster A, the access application or the sending application can continue to write or read data from the corresponding partition of the backup cluster B. In addition, the access application and sending application are also equipped with a Kafka indicator collection module and a thread pool indicator collection module, which are mainly used to collect corresponding data to determine whether the Kafka cluster or the corresponding thread pool currently executing the data sending task has any abnormalities, and further determine whether cluster switching or thread pool reallocation is required.
[0136] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the Kafka multi-cluster switching method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0137] This application embodiment also provides a Kafka multi-cluster switching device, please refer to Figure 8 , the Kafka multi-cluster switching device includes:
[0138] The collection module 10 is used to collect task indicators of data sending tasks corresponding to each data type in real time, where the data type is associated with each topic of the primary cluster in the Kafka resource pool. The Kafka resource pool has at least one primary cluster and its corresponding backup cluster;
[0139] A determination module 20 is configured to determine a topic corresponding to a data type whose task indicator is within a preset abnormal range as a target topic;
[0140] The switching module 30 is configured to switch the data sending task in the primary cluster to the standby cluster for execution when there is at least one target topic in the primary cluster.
[0141] The Kafka multi-cluster switching device provided in the embodiments of the present application utilizes the Kafka multi-cluster switching method of the aforementioned embodiments to improve the high availability of systems deploying Kafka clusters in weak environments. Compared to the prior art, the beneficial effects of the Kafka multi-cluster switching device provided in the present application are the same as those of the Kafka multi-cluster switching method provided in the aforementioned embodiments. Other technical features of the Kafka multi-cluster switching device are the same as those disclosed in the aforementioned embodiments and are not further described here.
[0142] An embodiment of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the Kafka multi-cluster switching method in the above-mentioned embodiment 1.
[0143] Reference below Figure 9 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0144] like Figure 9As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or a hard disk; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wired to exchange data. Although the figures show electronic devices with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.
[0145] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0146] The electronic device provided in the embodiments of this application utilizes the Kafka multi-cluster switching method described in the above embodiments to improve the high availability of systems deploying Kafka clusters in weak environments. Compared to the prior art, the electronic device provided in this application has the same beneficial effects as the Kafka multi-cluster switching method described in the above embodiments. Other technical features of the electronic device are the same as those disclosed in the above embodiments and are not further described here.
[0147] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0148] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0149] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the Kafka multi-cluster switching method in the above embodiment.
[0150] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0151] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0152] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by an electronic device, the electronic device: collects task indicators of data sending tasks corresponding to each data type in real time, wherein the data type is associated with each topic of the main cluster in the Kafka resource pool, and there is at least one main cluster and a corresponding backup cluster in the Kafka resource pool; determines the topic corresponding to the data type whose task indicator is within a preset abnormal range as the target topic; and when there is at least one target topic in the main cluster, switches the data sending task in the main cluster to the backup cluster for execution.
[0153] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0154] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0155] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0156] The computer-readable storage medium provided in the embodiments of this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned Kafka multi-cluster switching method, thereby improving the high availability of systems deploying Kafka clusters in weak environments. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the Kafka multi-cluster switching method provided in the aforementioned embodiments, and are not further elaborated here.
[0157] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned Kafka multi-cluster switching method.
[0158] The computer program product provided in the embodiments of this application can improve the high availability of systems deploying Kafka clusters in weak environments. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the Kafka multi-cluster switching method provided in the above embodiments, and will not be repeated here.
[0159] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A Kafka multi-cluster switching method, characterized in that: The Kafka multi-cluster switching method includes: Collect task indicators of data sending tasks corresponding to each data type in real time, where the data type is associated with each topic of the primary cluster in the Kafka resource pool, and the Kafka resource pool has at least one primary cluster and its corresponding backup cluster; The topic corresponding to the data type whose task indicator is within the preset abnormal range is determined as the target topic; In the case that there is at least one target topic in the primary cluster, the data sending task in the primary cluster is switched to the standby cluster for execution.
2. The Kafka multi-cluster switching method according to claim 1, characterized in that: After the step of switching the data sending task in the active cluster to the standby cluster for execution, the method further includes: Sending a probe message to the target topic; In the case that the detection information is sent successfully and the sending time is less than a preset time threshold, the data sending task is switched to the main cluster for execution.
3. The Kafka multi-cluster switching method according to claim 1, characterized in that: After the step of switching the data sending task in the active cluster to the standby cluster for execution, the method further includes: In the case that the standby cluster has at least one target topic, a new cluster is registered in the Kafka resource pool, and the data sending task is switched to the new cluster for execution.
4. The Kafka multi-cluster switching method according to claim 1, wherein: The method further comprises: In the event that the Kafka resource pool is unable to provide services, the message data received from the upstream sender is sent to the downstream receiver in real time, and the message data is stored in a third-party database; In the event of a real-time transmission failure, the message data is extracted from the third-party database and resent to the downstream recipient in batches according to a preset resending strategy.
5. The Kafka multi-cluster switching method according to claim 4, characterized in that: After the step of resending the message data to the downstream recipient in batches, the method further includes: According to the replenishment request of the receiver, the replenishment data of the message data is extracted from the third-party database, and the replenishment data is sent to the downstream receiver.
6. The Kafka multi-cluster switching method according to claim 1, characterized in that: The data sending task includes at least one subtask corresponding to each service code. After the step of switching the data sending task in the active cluster to the standby cluster for execution, the step further includes: Collect the time taken to execute the subtasks corresponding to each business code; In the case where the number of business codes whose task execution time exceeds the preset time threshold is greater than the preset number threshold, determining the priority of each business code according to the task execution time corresponding to each business code; According to each priority, the target business code is determined in sequence, and the thread pool is rebound for the subtask corresponding to the target business code.
7. The Kafka multi-cluster switching method according to claim 6, characterized in that: The step of binding a thread pool for the subtask corresponding to the target business code includes: If there is a corresponding dedicated thread pool for the target business code, bind the dedicated thread pool to the subtask corresponding to the target business code; If there is no dedicated thread pool corresponding to the target business code, retrieve the target thread pool with the highest priority and remaining quota from the preset thread pool library, and determine whether the remaining quota of the target thread pool is greater than the number of transactions processed per second of the target business code; When the remaining quota of the target thread pool is greater than the number of transactions processed per second, binding the subtask corresponding to the target business code to the target thread pool; When the remaining quota of the target thread pool is less than or equal to the number of transactions processed per second, the target thread pool with the highest priority and remaining quota is taken out from the remaining thread pools in the thread pool library, and the step of determining whether the remaining quota of the target thread pool is greater than the number of transactions processed per second of the target business code is performed.
8. An electronic device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the Kafka multi-cluster switching method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the Kafka multi-cluster switching method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the Kafka multi-cluster switching method according to any one of claims 1 to 7 are implemented.