Distributed service intelligent master selection method and system based on dual-engine collaboration

By adopting a distributed service intelligent master selection method based on dual-engine collaboration in the distributed service architecture, dynamic priority sorting and master-secure switching, the problems of insufficient flexibility and stability of master node selection in the existing technology are solved, and efficient and accurate master selection and strong fault tolerance are achieved.

CN119484230BActive Publication Date: 2025-05-16北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510039178.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In the existing distributed service architecture, the master node selection method has insufficient flexibility, high performance overhead, and dependence on databases leads to stability and availability problems.

Method used

The intelligent master selection method of distributed services based on dual-engine collaboration is adopted. Dynamic priority sorting and master-solder switching is carried out by receiving node registration requests, collecting performance data, building node scoring functions, generating intelligent identification codes, storing and listening in Nacos and Redis.

Benefits of technology

It improves the efficiency and accuracy of the master selection, enhances the stability and fault tolerance of the system, and optimizes resource utilization and load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484230B_ABST
    Figure CN119484230B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed service intelligent master selection method and system based on dual-engine collaboration, which relates to the field of intelligent master selection technology, including collecting the processing capacity index, network throughput, and resource utilization of service nodes to build a performance data matrix, generating an intelligent identification code based on a node scoring function and a physical clock value; monitoring Nacos and Redis signals through a double-layer event listener, determining the optimal incremental step size according to the distribution trend of the score sequence, generating a target score to build a dynamic priority sequence; periodically selecting a master node and a preheated standby node, and when the master node fails, starting incremental state synchronization based on Redis transactions after confirmation through voting arbitration, and realizing master-standby non-sense switching. The present invention improves the accuracy of master node election and the reliability of switching, and reduces the risk of system interruption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to intelligent leader selection technology, and in particular to a distributed service intelligent leader selection method and system based on dual-engine collaboration. Background Art

[0002] Distributed service architecture has become a popular choice for building high-availability, high-performance applications. In this architecture, multiple service nodes work together to complete tasks. In order to ensure the stability and efficiency of the service cluster, a master node is usually required to coordinate the work of each node, such as allocating tasks and managing data. Therefore, how to efficiently and reliably select the master node is a key issue in the distributed service architecture.

[0003] Existing master node selection methods mainly include static configuration, ZooKeeper-based election, and database-based election. These methods have some defects:

[0004] The static configuration method lacks flexibility and cannot dynamically adjust the master node according to the real-time status of the node. Once the configured master node fails, manual intervention is required, affecting the availability of the service.

[0005] Although the ZooKeeper-based election mechanism can automatically select the leader, it has a large performance overhead. When there are a large number of nodes, problems such as election timeout or brain split are prone to occur, affecting the stability of the cluster.

[0006] The database-based leader election method relies on the availability of the database. If the database fails, the leader election will fail, which will also affect the normal operation of the service. In addition, the read and write operations of the database will also cause certain performance losses. Summary of the invention

[0007] The embodiments of the present invention provide a distributed service intelligent master selection method and system based on dual-engine collaboration, which can solve the problems in the prior art.

[0008] According to a first aspect of the embodiments of the present invention,

[0009] Provides a distributed service intelligent master election method based on dual-engine collaboration, including:

[0010] After receiving the registration request of the distributed service node, the processing capability index, network throughput and resource utilization of the node are collected to form a performance data matrix, and a node scoring function is constructed based on the performance data matrix; the calculation result of the node scoring function is weightedly combined with the server physical clock value to generate the intelligent identification code of the node; a service instance is created in the Nacos registration center and the intelligent identification code is written into the ordered set of the Redis storage engine, and a two-layer event listener is activated at the same time, which is used to capture the instance state change signal of Nacos and the data structure change signal of Redis;

[0011] When the double-layer event listener detects a signal that a new node has joined, it extracts the score sequence of existing nodes in the ordered set, analyzes the distribution trend of the score sequence, and calculates the optimal score increment step based on the distribution trend; combines the optimal score increment step with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the running time of the node; prioritizes the nodes based on the target score, and writes the ranking result into the ordered set to form a dynamic priority sequence;

[0012] The highest priority node is periodically extracted from the dynamic priority sequence as the master node, and the second priority node is selected as the preheated standby node; when the double-layer event listener detects an abnormal load or fault signal of the master node, the fault state is confirmed through a voting arbitration mechanism; after confirmation, the preheated standby node is promoted to a new master node, and data migration between the new and old master nodes is completed through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos.

[0013] The calculation result of the node scoring function is weightedly combined with the server physical clock value to generate the node's intelligent identification code; a service instance is created in the Nacos registration center and the intelligent identification code is written into an ordered set of the Redis storage engine, and a double-layer event listener is activated at the same time. The double-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis, including:

[0014] Input the performance data matrix into the node scoring function for calculation, and output the node scoring result; collect the server physical clock value, perform a weighted combination operation on the node scoring result and the server physical clock value, and generate a node intelligent identification code, wherein the node intelligent identification code includes a performance scoring part and a timing identification part;

[0015] Create a distributed service instance in the Nacos registration center, write the smart identification code of the node into the extended attribute field of the distributed service instance; synchronously write the smart identification code of the node into an ordered set of the Redis storage engine, and the ordered set uses the performance score part of the smart identification code of the node as the sorting basis; the smart identification codes of the nodes in the ordered set are arranged in descending order according to the performance score;

[0016] Activate a two-layer event listener, which includes a Nacos instance monitoring layer and a Redis data monitoring layer; the Nacos instance monitoring layer captures the state change signal of the service instance through an instance subscription mechanism; the Redis data monitoring layer captures the data structure change signal of the ordered set through a key space notification mechanism.

[0017] When the double-layer event listener detects a new node joining signal, extracts the score sequence of existing nodes in the ordered set, analyzes the distribution trend of the score sequence, and calculates the optimal score increment step according to the distribution trend, including:

[0018] The two-layer event listener integrates the captured state change signal and data change signal into a unified event message stream; when the event message stream contains a new node joining signal, it reads the identification information and corresponding scores of all existing nodes from the ordered set of the Redis storage engine, extracts the corresponding scores to construct an ordered score sequence; calculates the difference between adjacent scores in the ordered score sequence to form a step sample set; performs a normality test on the step sample set to obtain the probability distribution characteristics of the step sample set;

[0019] Selecting an optimal step length calculation strategy based on the probability distribution characteristics: when the probability distribution characteristics of the step length sample set meet the normal distribution requirements, using the maximum likelihood estimation method to calculate the reference step length value; when the probability distribution characteristics of the step length sample set do not meet the normal distribution requirements, using the kernel density estimation method to calculate the reference step length value; inputting the calculated reference step length value into the adaptive adjustment module;

[0020] The adaptive adjustment module first obtains the current number of active nodes in the cluster, selects a capacity coefficient according to the number of active nodes, and multiplies the reference step value by the corresponding capacity coefficient to obtain a capacity adjustment step;

[0021] Collecting the performance index of the newly added node and calculating the average performance level of the existing nodes in the cluster; comparing the performance index of the newly added node with the average performance level to obtain a performance ratio; calculating a performance coefficient according to the performance ratio; multiplying the capacity adjustment step by the performance coefficient to obtain a performance correction step;

[0022] Extracting the generation timestamp of each score in the ordered score sequence, calculating the time interval between each score and the current moment; substituting the time interval into an exponential decay function to calculate a time weight, wherein the decay rate of the exponential decay function is dynamically adjusted according to the system load; performing weighted calculation on the performance correction step and the time weight to obtain a timing weighted step;

[0023] Analyze historical score adjustment records, extract frequency characteristics and amplitude characteristics of score adjustment, and determine the lower limit threshold of step length based on the frequency characteristics; predict the growth trend of the number of nodes, and determine the upper limit threshold of step length according to the growth trend; limit the time series weighted step length between the lower limit threshold and the upper limit threshold to obtain the final optimal score increment step length.

[0024] The optimal score increment step is combined with the intelligent identification code of the new node to generate a target score, wherein the target score shows a decaying trend as the node runs for a certain period of time; the nodes are prioritized based on the target score, and the ranking results are written into the ordered set to form a dynamic priority sequence, which includes:

[0025] The optimal score increment step is used to characterize the priority interval between adjacent nodes; the intelligent identification code of the new node is parsed to obtain a performance score part and a node identification part; the performance score part is normalized to obtain a performance weight coefficient;

[0026] The optimal score increment step is multiplied by the performance weight coefficient to obtain a benchmark score; the benchmark score is combined with the node identification part through a nonlinear mapping function to generate an initial target score of the node; the nonlinear mapping function ensures that the target scores of different nodes are globally unique and monotonically increasing;

[0027] Obtaining a count value of the running time of the node from the node status monitoring module; calculating a time decay factor based on the count value of the running time, wherein the time decay factor adopts a piecewise exponential function: when the running time is less than one hour, the time decay rate is the largest, and when the running time exceeds twenty-four hours, the time decay rate is the smallest;

[0028] Multiplying the initial target score of the node by the time decay factor to obtain a real-time target score adjusted by time decay; the real-time target score shows a decreasing trend as the running time of the node increases;

[0029] Periodically collect the performance indicators of all online nodes in the system and calculate the system load standard deviation; dynamically adjust the priority update cycle according to the load standard deviation: when the load standard deviation is greater than the preset load threshold, shorten the update cycle to half of the reference cycle; when the load standard deviation is less than the preset load threshold, restore the update cycle to the reference cycle;

[0030] At the beginning of each priority update cycle, the current version identifier of the ordered set of the Redis storage engine is obtained; the existing node priority sequence in the ordered set is read; the newly calculated real-time target score is combined with the node identifier to form a priority update record; the priority update record is inserted into the node priority sequence according to the score size to generate an updated dynamic priority sequence.

[0031] Periodically extracting the highest priority node from the dynamic priority sequence as the primary node, and selecting the secondary priority node as the preheated standby node includes:

[0032] Selecting a node with the highest priority value from the dynamic priority sequence as a master node; selecting a node with the second highest priority value as a preheated standby node; constructing a double-buffered data synchronization channel with a write-ahead log function between the master node and the preheated standby node; the double-buffered data synchronization channel includes a first buffer for transmitting the write-ahead log and a second buffer for transmitting incremental data;

[0033] Monitor the data volume of the first buffer and the second buffer; when the data volume of the first buffer reaches a first preset threshold, push the write-ahead log to the preheated standby node; when the data volume of the second buffer reaches a second preset threshold, push the incremental data to the preheated standby node; the preheated standby node maintains data synchronization based on the received write-ahead log and incremental data;

[0034] Construct a lightweight memory snapshot mechanism based on a time window; at the beginning of each snapshot cycle, record the current memory state of the master node to obtain the current memory snapshot; compare the current memory snapshot with the memory snapshot of the previous cycle to extract the difference data of the memory state; transmit the difference data to the preheated standby node; the preheated standby node restores the real-time memory state of the master node based on the difference data.

[0035] When the two-layer event listener detects an abnormal load or fault signal of the master node, the fault status is confirmed through the voting arbitration mechanism including:

[0036] Deploy a two-layer event listener, the local monitoring layer of the two-layer event listener collects the system load index, response delay index, and error rate index of the node; performs exponential moving average processing on the system load index to obtain a smoothed load value; calculates the response delay index in a sliding time window to obtain delay statistical characteristics; calculates the time series change characteristics of the error rate index to obtain an error trend value;

[0037] The smoothed load value, the delay statistical feature, and the error trend value are input into a Z-score anomaly detection model, and the Z-score anomaly detection model calculates the mean and standard deviation of each indicator based on historical data; the deviation between the current indicator value and the corresponding mean is divided by the standard deviation to obtain an anomaly score; and the anomaly score is accumulated within the sliding time window to obtain a cumulative anomaly value;

[0038] When the accumulated abnormal value exceeds the trigger threshold, the distributed voting arbitration process based on the Raft protocol is started; during the voting arbitration process, the trust weight is calculated based on the historical fault judgment accuracy of each node; the trust weight increases with the number of correct judgments and decreases with the number of wrong judgments; the voting result of each node is multiplied by its trust weight to obtain a weighted voting value, and when the weighted voting value exceeds two-thirds of the trigger threshold, the fault state of the master node is determined.

[0039] After confirmation, the preheated standby node is promoted to the new master node, and the data migration between the new and old master nodes is completed through the incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos, including:

[0040] Monitor the status indicators of the preheated standby node through a heartbeat detection mechanism, the status indicators include CPU usage, memory usage, and network latency; use a two-way verification mechanism to confirm the status of the preheated standby node, the two-way verification mechanism includes active query by the monitoring component and active report by the preheated standby node, and generate a status confirmation result of the preheated standby node;

[0041] Based on the status confirmation result, determine whether the preheated standby node meets the promotion condition; when the promotion condition is met, start a switching transaction in Redis to promote the preheated standby node to the new master node; obtain the write lock of the old master node, suspend new write requests, and record the latest operation sequence number of the old master node as the data synchronization starting point;

[0042] Constructing an incremental state synchronization task based on the latest operation sequence number, the incremental state synchronization task comprising: constructing a persistent operation log queue to record write operation information, determining the operation log range to be synchronized, dividing the operation log range into batches according to time segments, and generating an operation batch sequence to be synchronized;

[0043] Perform incremental data migration on the operation batch sequence to be synchronized: package and transmit the operation instructions in each operation batch to the new master node in a pipelined manner, convert the operation batch into a Redis command sequence on the new master node, execute the Redis command sequence atomically through the EXEC command, update the synchronization progress mark, and generate a data migration completion mark;

[0044] Triggering active / standby switching based on the data migration completion flag: submitting the role information and end node information of the new active node to the Nacos registration server, pushing the role information and end node information of the new active node to the subscription client through the Nacos service discovery mechanism, and the subscription client updating the local service routing table;

[0045] Assign a globally unique version number to the new master node, and build a lease-based protection mechanism: the new master node renews the contract with the Nacos registration server at a preset time interval to obtain renewal information; the renewal information includes the version number; if the renewal information is detected to be timed out, the protection mechanism is triggered to terminate the service of the new master node;

[0046] Before the switching transaction is committed, atomic confirmation is performed based on the transaction characteristics of Redis: the data consistency of the new master node is checked, the uniqueness of the version number is verified, and the update status of the local service routing table is confirmed; when the atomic confirmation passes, the switching transaction is committed; when the atomic confirmation fails, the switching transaction is rolled back, and the new master node information is broadcast to the cluster through the service discovery function of Nacos.

[0047] According to a second aspect of the embodiments of the present invention,

[0048] Provides a distributed service intelligent master election system based on dual-engine collaboration, including:

[0049] The first unit is used to collect the processing capability index, network throughput, and resource utilization of the node to form a performance data matrix after receiving a registration request from a distributed service node, and to construct a node scoring function based on the performance data matrix; to perform a weighted combination of the calculation result of the node scoring function and the server physical clock value to generate a smart identification code for the node; to create a service instance in the Nacos registration center and write the smart identification code into an ordered set of the Redis storage engine, and to activate a two-layer event listener at the same time, wherein the two-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis;

[0050] The second unit is used for extracting the score sequence of existing nodes in the ordered set when the double-layer event listener detects a new node joining signal, analyzing the distribution trend of the score sequence, and calculating the optimal score increment step based on the distribution trend; combining the optimal score increment step with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the node running time; sorting the nodes based on the target score, and writing the sorting results into the ordered set to form a dynamic priority sequence;

[0051] The third unit is used to periodically extract the highest priority node from the dynamic priority sequence as the master node, and select the second priority node as the preheated standby node; when the two-layer event listener detects an abnormal load or fault signal of the master node, the fault state is confirmed through a voting arbitration mechanism; after confirmation, the preheated standby node is promoted to a new master node, and the data migration between the new and old master nodes is completed through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos.

[0052] According to a third aspect of the embodiments of the present invention,

[0053] An electronic device is provided, comprising:

[0054] processor;

[0055] a memory for storing processor-executable instructions;

[0056] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0057] According to a fourth aspect of the embodiments of the present invention,

[0058] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0059] The beneficial effects of this application are as follows:

[0060] 1. Improve the efficiency and accuracy of leader selection: Through the node scoring function and intelligent identification code, combined with node performance and time factors, the node capabilities can be evaluated more accurately, avoiding misjudgments caused by relying on a single indicator in traditional methods, thereby improving the efficiency and accuracy of leader selection and selecting a more suitable leader node.

[0061] 2. Enhance system stability and fault tolerance: The dual-layer event monitoring mechanism and the design of preheating standby nodes can timely detect abnormalities or failures of the main node, and quickly complete the main-standby switch, reducing the system unavailable time. The incremental state synchronization mechanism ensures data consistency and further improves the stability and fault tolerance of the system.

[0062] 3. Optimize resource utilization and load balancing: The dynamic priority sequence and the optimal score increment step strategy enable the node priority to be dynamically adjusted according to the operating status, achieving load balancing and avoiding excessive concentration or idleness of resources, thereby optimizing resource utilization and improving overall operating efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1A schematic diagram of a process flow of a distributed service intelligent master selection method based on dual-engine collaboration according to an embodiment of the present invention;

[0064] Figure 2 The figure is a schematic diagram of the structure of a distributed service intelligent master election system based on dual-engine collaboration according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0067] Figure 1 FIG. 1 is a flow chart of a distributed service intelligent master selection method based on dual-engine collaboration according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0068] S11. After receiving the registration request of the distributed service node, the processing capacity index, network throughput, and resource utilization of the node are collected to form a performance data matrix, and a node scoring function is constructed based on the performance data matrix; the calculation result of the node scoring function is weightedly combined with the server physical clock value to generate the node's intelligent identification code; a service instance is created in the Nacos registration center and the intelligent identification code is written into the ordered set of the Redis storage engine, and a double-layer event listener is activated at the same time, which is used to capture the instance state change signal of Nacos and the data structure change signal of Redis;

[0069] S12. When the double-layer event listener detects a signal that a new node has joined, it extracts the score sequence of the existing nodes in the ordered set, analyzes the distribution trend of the score sequence, and calculates the optimal score increment step based on the distribution trend; combines the optimal score increment step with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the running time of the node; prioritizes the nodes based on the target score, and writes the ranking results into the ordered set to form a dynamic priority sequence;

[0070] S13. Periodically extract the highest priority node from the dynamic priority sequence as the master node, and select the second priority node as the preheated standby node; when the two-layer event listener detects an abnormal load or fault signal of the master node, confirm the fault state through a voting arbitration mechanism; after confirmation, promote the preheated standby node to the new master node, and complete the data migration between the new and old master nodes through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos.

[0071] In an optional implementation, the calculation result of the node scoring function is weightedly combined with the server physical clock value to generate a smart identification code for the node; a service instance is created in the Nacos registration center and the smart identification code is written into an ordered set of the Redis storage engine, and a double-layer event listener is activated at the same time. The double-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis, including:

[0072] Input the performance data matrix into the node scoring function for calculation, and output the node scoring result; collect the server physical clock value, perform a weighted combination operation on the node scoring result and the server physical clock value, and generate a node intelligent identification code, wherein the node intelligent identification code includes a performance scoring part and a timing identification part;

[0073] Create a distributed service instance in the Nacos registration center, write the smart identification code of the node into the extended attribute field of the distributed service instance; synchronously write the smart identification code of the node into an ordered set of the Redis storage engine, and the ordered set uses the performance score part of the smart identification code of the node as the sorting basis; the smart identification codes of the nodes in the ordered set are arranged in descending order according to the performance score;

[0074] Activate a two-layer event listener, which includes a Nacos instance monitoring layer and a Redis data monitoring layer; the Nacos instance monitoring layer captures the state change signal of the service instance through an instance subscription mechanism; the Redis data monitoring layer captures the data structure change signal of the ordered set through a key space notification mechanism.

[0075] A dynamic node performance ranking and service discovery method based on Nacos and Redis, the specific implementation method is as follows:

[0076] First, we build a node performance data collection module, which is responsible for collecting the performance data of each node in real time.

[0077] Next, the collected performance data matrix is ​​input into the node scoring function for calculation. The node scoring function can be customized according to actual business needs. For example, the weighted average method can be used to assign different weights according to the importance of different performance indicators.

[0078] At the same time, the physical clock value of the server is obtained. Assume that the physical clock value of the current server is 1678886400000 (millisecond timestamp). Combine the node scoring result with the server physical clock value to generate the intelligent identification code of the node. The intelligent identification code contains the performance scoring part and the timing identification part. The performance scoring part is used for sorting, and the timing identification part is used to distinguish different nodes with the same performance score.

[0079] Then, create a distributed service instance in the Nacos registry and write the node's smart identification code into the extended attribute field of the service instance. For example, create a service instance named "test-service" and write the smart identification codes of nodes A, B, and C into their extended attributes respectively.

[0080] At the same time, the smart identification code of the node is written into the ordered set of Redis, and the performance score part of the smart identification code is used as the sorting basis. For example, create an ordered set named "node-rank" and add the smart identification codes of nodes A, B, and C as members to the set, with scores of 69, 63, and 75 respectively. In this way, the smart identification codes of the nodes in the ordered set will be arranged from high to low according to the performance score.

[0081] Finally, activate the two-layer event listener. The Nacos instance monitoring layer captures the status change signals of the service instance through the subscription mechanism, such as the service instance online and offline events. The Redis data monitoring layer captures the data structure change signals of the ordered set through the key space notification mechanism, such as member addition, deletion, score update and other events. When an event occurs in Nacos or Redis, the listener will trigger the corresponding processing logic, such as updating the service list, recalculating the node score, etc.

[0082] The solution of this application can:

[0083] Implemented dynamic sorting of node performance: By collecting node performance data in real time and inputting it into the scoring function for calculation, dynamic sorting of node performance can be implemented, thereby timely reflecting the actual performance status of the node. Improved the efficiency of service discovery: By writing the intelligent identification code of the node into Nacos and Redis, efficient service discovery can be achieved. The client can quickly select the optimal node based on the performance score of the node, thereby improving the response speed and stability of the service. Enhanced the fault tolerance of the system: Through the two-layer event monitoring mechanism, changes in node status and data structure can be perceived in a timely manner, so as to quickly respond to faults and perform corresponding processing, such as removing the faulty node from the service list, or reselecting the optimal node, thereby enhancing the fault tolerance of the system.

[0084] In an optional implementation, when the double-layer event listener detects a new node joining signal, extracting the score sequence of the existing nodes in the ordered set, analyzing the distribution trend of the score sequence, and calculating the optimal score increment step according to the distribution trend includes:

[0085] The two-layer event listener integrates the captured state change signal and data change signal into a unified event message stream; when the event message stream contains a new node joining signal, it reads the identification information and corresponding scores of all existing nodes from the ordered set of the Redis storage engine, extracts the corresponding scores to construct an ordered score sequence; calculates the difference between adjacent scores in the ordered score sequence to form a step sample set; performs a normality test on the step sample set to obtain the probability distribution characteristics of the step sample set;

[0086] Selecting an optimal step length calculation strategy based on the probability distribution characteristics: when the probability distribution characteristics of the step length sample set meet the normal distribution requirements, using the maximum likelihood estimation method to calculate the reference step length value; when the probability distribution characteristics of the step length sample set do not meet the normal distribution requirements, using the kernel density estimation method to calculate the reference step length value; inputting the calculated reference step length value into the adaptive adjustment module;

[0087] The adaptive adjustment module first obtains the current number of active nodes in the cluster, selects a capacity coefficient according to the number of active nodes, and multiplies the reference step value by the corresponding capacity coefficient to obtain a capacity adjustment step;

[0088] Collecting the performance index of the newly added node and calculating the average performance level of the existing nodes in the cluster; comparing the performance index of the newly added node with the average performance level to obtain a performance ratio; calculating a performance coefficient according to the performance ratio; multiplying the capacity adjustment step by the performance coefficient to obtain a performance correction step;

[0089] Extracting the generation timestamp of each score in the ordered score sequence, calculating the time interval between each score and the current moment; substituting the time interval into an exponential decay function to calculate a time weight, wherein the decay rate of the exponential decay function is dynamically adjusted according to the system load; performing weighted calculation on the performance correction step and the time weight to obtain a timing weighted step;

[0090] Analyze historical score adjustment records, extract frequency characteristics and amplitude characteristics of score adjustment, and determine the lower limit threshold of step length based on the frequency characteristics; predict the growth trend of the number of nodes, and determine the upper limit threshold of step length according to the growth trend; limit the time series weighted step length between the lower limit threshold and the upper limit threshold to obtain the final optimal score increment step length.

[0091] The method of double-layer event listener detecting the addition of new nodes and calculating the optimal score increment step is implemented as follows:

[0092] First, deploy a two-layer event listener to monitor the state changes and data changes of the Redis cluster in real time. The listener can capture various types of events such as node joining, node leaving, data modification, etc., and integrate these different types of event signals into a unified event message stream. For example, the event message stream contains the event type, event occurrence timestamp, related node identifier, data change amount, etc.

[0093] When a new node joins the event message stream, the listener triggers the subsequent score calculation process. At this time, the listener reads the identification information and corresponding scores of all existing nodes from the ordered set of the Redis storage engine. Assume that the key name of the ordered set is "node_scores", which stores the node ID and the corresponding score.

[0094] Next, calculate the difference between adjacent scores in the ordered score sequence to form a step sample set. Taking the above example, the difference between adjacent scores is 15 and 20, so the step sample set is [15, 20].

[0095] Then, the step length sample set is tested for normality to analyze its probability distribution characteristics. The test method can be Shapiro-Wilk test or Kolmogorov-Smirnov test. Through the test, it can be determined whether the step length sample set conforms to the normal distribution. The hypothesis test result shows that the step length sample set does not conform to the normal distribution.

[0096] Since the step length sample set does not conform to the normal distribution, the kernel density estimation method is used to calculate the benchmark step length value. The kernel density estimation method can estimate its probability density function based on the distribution of sample data and extract the benchmark step length value from it. Assume that the benchmark step length value calculated by the kernel density estimation method is 18.

[0097] The calculated benchmark step value 18 is input into the adaptive adjustment module. The module first obtains the current number of active nodes in the cluster. Assume that the current cluster has 3 active nodes. According to the number of active nodes, select the corresponding capacity factor. Assume that the capacity factor corresponding to 3 active nodes is 0.8. Multiply the benchmark step value 18 by the capacity factor 0.8 to get the capacity adjustment step size 14.4.

[0098] Collect the performance indicators of the newly added nodes, such as CPU usage, memory usage, network throughput, etc. At the same time, calculate the average performance level of the existing nodes in the cluster. Assume that the comprehensive score of the performance indicators of the newly added nodes is 80, and the average performance level of the existing nodes in the cluster is 90. Compare the performance indicator of the newly added nodes 80 with the average performance level 90, and the performance ratio is about 0.89. Calculate the performance coefficient based on the performance ratio. Assume that the performance coefficient corresponding to the performance ratio of 0.89 is 0.9. Multiply the capacity adjustment step of 14.4 by the performance coefficient of 0.9 to get the performance correction step of 12.96.

[0099] Extract the generation timestamp of each score in the ordered score sequence. Assume that the score generation timestamps of nodes A, B, and C are 1678886400, 1678972800, and 1679059200 (corresponding to March 15, 16, and 17, 2023, respectively). Calculate the time interval between each score and the current moment. Assuming that the current timestamp is 1679145600 (March 18, 2023), the time intervals between each score and the current moment are 259200 seconds, 172800 seconds, and 86400 seconds, respectively. Substitute the time interval into the exponential decay function to calculate the time weight. The decay rate of the exponential decay function is dynamically adjusted according to the system load. Perform weighted calculation on the performance correction step size 12.96 and the corresponding time weight to obtain the time series weighted step size. Assume that the final time series weighted step size is 12.5.

[0100] Analyze the historical score adjustment records and extract the frequency and amplitude characteristics of the score adjustment. Assume that the analysis results show that the frequency of score adjustment is high and the amplitude is small. Determine the lower limit threshold of the step size based on the frequency characteristics, for example, set it to 5. Predict the growth trend of the number of nodes. Assume that the prediction results show that the number of nodes will grow rapidly. Determine the upper limit threshold of the step size based on the growth trend, for example, set it to 20. Limit the time series weighted step size of 12.5 between the lower limit threshold of 5 and the upper limit threshold of 20. Since 12.5 is between 5 and 20, the final optimal score increment step size is 12.5.

[0101] The solution of this application can:

[0102] Improve the accuracy of node sorting: By analyzing the distribution trend of existing node scores and the performance indicators of newly added nodes, dynamically adjust the score increment step to make the node sorting more in line with the actual situation and improve the accuracy of sorting. Enhance the stability of the cluster: Through the adaptive adjustment module and the timing weighting mechanism, the score can be dynamically adjusted according to the cluster load and node performance to avoid excessive concentration or dispersion of scores and enhance the stability of the cluster. Optimize resource allocation efficiency: By considering node performance, cluster capacity and time factors, resources can be allocated more effectively, resource utilization can be improved, and resource allocation efficiency can be optimized.

[0103] In an optional implementation, the optimal score increment step is combined with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the node running time; the nodes are prioritized based on the target score, and the ranking results are written into the ordered set to form a dynamic priority sequence, including:

[0104] The optimal score increment step is used to characterize the priority interval between adjacent nodes; the intelligent identification code of the new node is parsed to obtain a performance score part and a node identification part; the performance score part is normalized to obtain a performance weight coefficient;

[0105] The optimal score increment step is multiplied by the performance weight coefficient to obtain a benchmark score; the benchmark score is combined with the node identification part through a nonlinear mapping function to generate an initial target score of the node; the nonlinear mapping function ensures that the target scores of different nodes are globally unique and monotonically increasing;

[0106] Obtaining a count value of the running time of the node from the node status monitoring module; calculating a time decay factor based on the count value of the running time, wherein the time decay factor adopts a piecewise exponential function: when the running time is less than one hour, the time decay rate is the largest, and when the running time exceeds twenty-four hours, the time decay rate is the smallest;

[0107] Multiplying the initial target score of the node by the time decay factor to obtain a real-time target score adjusted by time decay; the real-time target score shows a decreasing trend as the running time of the node increases;

[0108] Periodically collect the performance indicators of all online nodes in the system and calculate the system load standard deviation; dynamically adjust the priority update cycle according to the load standard deviation: when the load standard deviation is greater than the preset load threshold, shorten the update cycle to half of the reference cycle; when the load standard deviation is less than the preset load threshold, restore the update cycle to the reference cycle;

[0109] At the beginning of each priority update cycle, the current version identifier of the ordered set of the Redis storage engine is obtained; the existing node priority sequence in the ordered set is read; the newly calculated real-time target score is combined with the node identifier to form a priority update record; the priority update record is inserted into the node priority sequence according to the score size to generate an updated dynamic priority sequence.

[0110] A node dynamic priority sorting method based on intelligent identification code and running time is used to optimize resource scheduling and improve system efficiency. The core idea of ​​this method is to dynamically adjust the priority of the node according to its performance and running time, and store the sorting results in a Redis ordered set for fast access and update.

[0111] First, the system presets an optimal score increment step, which is used to characterize the priority interval between adjacent nodes, for example, set to 10. At the same time, the system maintains a reference period, for example, set to 60 seconds, as a reference time interval for priority update.

[0112] When a new node joins the system, its smart identification code is parsed. Assume that a smart identification code is "A100B001", where "A100" represents the performance score part and "B001" represents the node identification part. The performance score part "A100" is normalized, for example, "A100" is mapped to a performance weight coefficient of 0.8. Multiply the optimal score increment step of 10 by the performance weight coefficient of 0.8 to obtain a baseline score of 8.

[0113] Next, the benchmark score 8 is combined with the node identification part "B001" through a nonlinear mapping function to generate the node's initial target score. The nonlinear mapping function needs to ensure that the target scores of different nodes are globally unique and monotonically increasing. For example, "B001" can be converted to an integer, such as 1, and then the benchmark score is added to the integer to obtain an initial target score of 9. Assuming that the identification code of another node is "A050B002", after the same calculation, its initial target score is 6, which ensures that the target scores of different nodes are different.

[0114] The target score is attenuated according to the node running time. The count value of the node's running time is obtained from the node status monitoring module. Assuming that the node has been running for 2 hours, the time decay factor is calculated according to the piecewise exponential function. For example, the decay factor is 0.9 within 1 hour, 0.95 between 1 hour and 24 hours, and 0.99 for more than 24 hours. Since the node has been running for 2 hours, the decay factor is 0.95. Multiply the node's initial target score of 9 by the time decay factor of 0.95 to obtain a real-time target score of 8.55 adjusted by time decay.

[0115] The system periodically collects the performance indicators of all online nodes and calculates the standard deviation of the system load. Assume that the preset load threshold is 0.5. When the calculated load standard deviation is greater than 0.5, the update cycle is shortened to half of the benchmark cycle, that is, 30 seconds; when the load standard deviation is less than 0.5, the update cycle is restored to the benchmark cycle of 60 seconds.

[0116] At the beginning of each priority update cycle, obtain the current version identifier of the ordered set of the Redis storage engine. Read the existing node priority sequence in the ordered set. Combine the newly calculated real-time target score 8.55 with the node identifier "B001" to form a priority update record (B001,8.55). Insert the record into the node priority sequence according to the score size to generate an updated dynamic priority sequence and store it in the Redis ordered set.

[0117] The solution of this application can:

[0118] Improve resource utilization: By dynamically adjusting node priorities, nodes with high performance and short running time are scheduled first, making full use of system resources and avoiding resource waste, thereby improving overall resource utilization. Enhance system stability: The target score of nodes with long running time will decay, and their priority will be reduced to avoid excessive resource occupation by nodes running for a long time, reduce the risk of system crash, and enhance system stability. Simplify management and maintenance: Use Redis ordered sets to store priority sequences for easy and fast access and update, simplify system management and maintenance, and improve operation and maintenance efficiency.

[0119] In an optional implementation, periodically extracting the highest priority node from the dynamic priority sequence as the primary node, and selecting the secondary priority node as the preheated standby node includes:

[0120] Selecting a node with the highest priority value from the dynamic priority sequence as a master node; selecting a node with the second highest priority value as a preheated standby node; constructing a double-buffered data synchronization channel with a write-ahead log function between the master node and the preheated standby node; the double-buffered data synchronization channel includes a first buffer for transmitting the write-ahead log and a second buffer for transmitting incremental data;

[0121] Monitor the data volume of the first buffer and the second buffer; when the data volume of the first buffer reaches a first preset threshold, push the write-ahead log to the preheated standby node; when the data volume of the second buffer reaches a second preset threshold, push the incremental data to the preheated standby node; the preheated standby node maintains data synchronization based on the received write-ahead log and incremental data;

[0122] Construct a lightweight memory snapshot mechanism based on a time window; at the beginning of each snapshot cycle, record the current memory state of the master node to obtain the current memory snapshot; compare the current memory snapshot with the memory snapshot of the previous cycle to extract the difference data of the memory state; transmit the difference data to the preheated standby node; the preheated standby node restores the real-time memory state of the master node based on the difference data.

[0123] A data synchronization method for master and standby nodes based on dynamic priority and double buffer mechanism is designed to improve system availability and data consistency. The core of this method is to dynamically select master and standby nodes and use double buffer mechanism and memory snapshot mechanism to achieve efficient data synchronization.

[0124] First, select the node with the highest priority from the dynamic priority sequence as the primary node, and the node with the second highest priority as the preheated standby node. Suppose there are three nodes A, B, and C, whose priority values ​​are 10, 8, and 5 respectively. According to the priority order, select A as the primary node and B as the preheated standby node.

[0125] A double-buffered data synchronization channel with a write-ahead log function is constructed between the primary node A and the preheated standby node B. The channel contains two buffers: the first buffer is used to transmit the write-ahead log, and the second buffer is used to transmit incremental data.

[0126] Set the preset thresholds of the first buffer and the second buffer. For example, set the first buffer threshold to 1MB and the second buffer threshold to 10MB.

[0127] Before modifying data, the application on the master node A first records the modification operation in the write-before log and writes the log to the first buffer. At the same time, the modified data increment is written to the second buffer.

[0128] Monitor the data volume of the first buffer and the second buffer. When the data volume of the first buffer reaches 1MB, push the write-before log to the preheated standby node B. When the data volume of the second buffer reaches 10MB, push the incremental data to the preheated standby node B.

[0129] After receiving the write-ahead log and incremental data, the preheated standby node B maintains data synchronization based on the received information to ensure that its own data is consistent with that of the primary node A. For example, if the primary node A modifies the age data of the user "Zhang San", a write-ahead log will be generated to record the age value before and after the modification, and the log will be pushed to the preheated standby node B. At the same time, the increment of the age data (for example, from 20 to 25 years old, the increment is 5) is pushed to the preheated standby node B. Based on the write-ahead log and incremental data, the preheated standby node B updates the age of the user "Zhang San" to 25 years old.

[0130] Build a lightweight memory snapshot mechanism based on time windows. Set the snapshot period, for example, take a snapshot every hour.

[0131] At the beginning of each snapshot cycle, the current memory state of the master node A is recorded to obtain the current memory snapshot. Assume that at the beginning of the current snapshot cycle, 1000 pieces of user data are stored in the memory of the master node A.

[0132] Compare the current memory snapshot with the memory snapshot of the previous cycle to extract the difference data of the memory status. Assuming that the memory snapshot of the previous cycle contains 990 user data, the difference data is the 10 newly added user data.

[0133] The difference data is transmitted to the preheated standby node B. The preheated standby node B restores the real-time memory state of the primary node A based on the difference data, and adds the newly added 10 user data to its own memory, so that its memory also contains 1000 user data.

[0134] Through the above steps, data synchronization between the active and standby nodes is achieved, ensuring high availability and data consistency of the system.

[0135] The solution of this application can:

[0136] Improve system availability: By preheating the standby node mechanism, when the primary node fails, the standby node can quickly take over the service, reducing system downtime. Ensure data consistency: The double buffer mechanism and memory snapshot mechanism ensure data synchronization between the primary and standby nodes, ensuring data consistency. Reduce synchronization costs: The lightweight memory snapshot mechanism reduces the bandwidth consumption and time cost of data synchronization by transmitting differential data.

[0137] In an optional implementation, when the dual-layer event listener detects a load anomaly or a fault signal of the master node, confirming the fault state through a voting arbitration mechanism includes:

[0138] Deploy a two-layer event listener, the local monitoring layer of the two-layer event listener collects the system load index, response delay index, and error rate index of the node; performs exponential moving average processing on the system load index to obtain a smoothed load value; calculates the response delay index in a sliding time window to obtain delay statistical characteristics; calculates the time series change characteristics of the error rate index to obtain an error trend value;

[0139] The smoothed load value, the delay statistical feature, and the error trend value are input into a Z-score anomaly detection model, and the Z-score anomaly detection model calculates the mean and standard deviation of each indicator based on historical data; the deviation between the current indicator value and the corresponding mean is divided by the standard deviation to obtain an anomaly score; and the anomaly score is accumulated within the sliding time window to obtain a cumulative anomaly value;

[0140] When the accumulated abnormal value exceeds the trigger threshold, the distributed voting arbitration process based on the Raft protocol is started; during the voting arbitration process, the trust weight is calculated based on the historical fault judgment accuracy of each node; the trust weight increases with the number of correct judgments and decreases with the number of wrong judgments; the voting result of each node is multiplied by its trust weight to obtain a weighted voting value, and when the weighted voting value exceeds two-thirds of the trigger threshold, the fault state of the master node is determined.

[0141] A master node failure detection method based on a double-layer event listener and the Raft protocol, the specific implementation steps are as follows:

[0142] First, deploy a two-layer event listener. The listener consists of a local monitoring layer and a distributed arbitration layer. The local monitoring layer is responsible for collecting various performance indicators of the master node in real time, including system load indicators, response delay indicators, and error rate indicators.

[0143] Then, the collected indicators are preprocessed. The system load indicators are processed by exponential moving average to obtain a smoothed load value, which can effectively filter out instantaneous fluctuations. For example, the current load is 80%, the load at the previous moment was 70%, and the smoothing coefficient is set to 0.8, then the smoothed load value is:

[0144] 0.8*80%+(1-0.8)*70%=78%.

[0145] In the sliding time window (for example, the past 1 minute), statistical response delay indicators such as maximum value, minimum value, average value, 95% percentile value, etc. are collected to obtain delay statistical characteristics. The time series change characteristics of the error rate indicator are calculated, for example, the error rate in the past 1 minute, 5 minutes, and 15 minutes, as well as the error rate change trend, to obtain the error trend value. Assuming that the error rate in the past 1 minute is 1%, the error rate in 5 minutes is 0.5%, and the error rate in 15 minutes is 0.2%, the error trend value is decreasing.

[0146] Next, the preprocessed indicators are input into the Z-score anomaly detection model. The model calculates the mean and standard deviation of each indicator based on historical data. For example, historical data shows that the mean of the smoothed load value is 50% and the standard deviation is 10%. Divide the deviation of the current smoothed load value from the mean by the standard deviation to get the anomaly score. For example, if the current smoothed load value is 78%, the anomaly score is (78%-50%) / 10%=2.8. Accumulate the anomaly scores within the sliding time window to get the cumulative anomaly value. For example, in the past 1 minute, the anomaly scores were 1.2, 1.5, and 2.8 respectively, and the cumulative anomaly value is 1.2+1.5+2.8=5.5.

[0147] When the cumulative outlier value exceeds the preset trigger threshold, for example, the threshold is set to 5, and the current cumulative outlier value is 5.5, exceeding the threshold, the distributed voting arbitration process based on the Raft protocol is initiated. During the voting arbitration process, the trust weight is calculated based on the historical fault judgment accuracy of each node. The trust weight of the node increases with the number of correct judgments and decreases with the number of incorrect judgments. For example, in the past 10 fault judgments of node A, 8 were correct and 2 were wrong, so its trust weight is 0.8.

[0148] Multiply the voting results of each node by its trust weight to get the weighted voting value. For example, if node A votes that the master node is faulty and its trust weight is 0.8, then the weighted voting value is 0.8*1=0.8. When the weighted voting value exceeds the two-thirds trigger threshold, for example, three nodes participate in the vote, two nodes believe that the master node is faulty, and the sum of their weighted voting values ​​exceeds the two-thirds threshold, the fault state of the master node is finally determined.

[0149] The solution of this application can:

[0150] Improve the accuracy of fault detection: Through the two-layer event monitoring mechanism and Z-score anomaly detection model, the abnormal state of the master node can be more accurately identified and the misjudgment rate can be reduced. Enhance the reliability of fault detection: Adopting a distributed voting arbitration mechanism based on the Raft protocol and introducing node trust weights can effectively avoid the impact of single point failures and improve the reliability of fault detection. Improve the efficiency of fault handling: Through real-time monitoring and fast arbitration, master node failures can be discovered and handled in a timely manner, shortening the fault recovery time and improving the availability of the system.

[0151] In an optional implementation, after confirmation, the preheated standby node is promoted to the new master node, and data migration between the new and old master nodes is completed through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos, including:

[0152] Monitor the status indicators of the preheated standby node through a heartbeat detection mechanism, the status indicators include CPU usage, memory usage, and network latency; use a two-way verification mechanism to confirm the status of the preheated standby node, the two-way verification mechanism includes active query by the monitoring component and active report by the preheated standby node, and generate a status confirmation result of the preheated standby node;

[0153] Based on the status confirmation result, determine whether the preheated standby node meets the promotion condition; when the promotion condition is met, start a switching transaction in Redis to promote the preheated standby node to the new master node; obtain the write lock of the old master node, suspend new write requests, and record the latest operation sequence number of the old master node as the data synchronization starting point;

[0154] Constructing an incremental state synchronization task based on the latest operation sequence number, the incremental state synchronization task comprising: constructing a persistent operation log queue to record write operation information, determining the operation log range to be synchronized, dividing the operation log range into batches according to time segments, and generating an operation batch sequence to be synchronized;

[0155] Perform incremental data migration on the operation batch sequence to be synchronized: package and transmit the operation instructions in each operation batch to the new master node in a pipelined manner, convert the operation batch into a Redis command sequence on the new master node, execute the Redis command sequence atomically through the EXEC command, update the synchronization progress mark, and generate a data migration completion mark;

[0156] Triggering active / standby switching based on the data migration completion flag: submitting the role information and end node information of the new active node to the Nacos registration server, pushing the role information and end node information of the new active node to the subscription client through the Nacos service discovery mechanism, and the subscription client updating the local service routing table;

[0157] Assign a globally unique version number to the new master node, and build a lease-based protection mechanism: the new master node renews the contract with the Nacos registration server at a preset time interval to obtain renewal information; the renewal information includes the version number; if the renewal information is detected to be timed out, the protection mechanism is triggered to terminate the service of the new master node;

[0158] Before the switching transaction is committed, atomic confirmation is performed based on the transaction characteristics of Redis: the data consistency of the new master node is checked, the uniqueness of the version number is verified, and the update status of the local service routing table is confirmed; when the atomic confirmation passes, the switching transaction is committed; when the atomic confirmation fails, the switching transaction is rolled back, and the new master node information is broadcast to the cluster through the service discovery function of Nacos.

[0159] A distributed system master-slave switching method based on Redis and Nacos, the specific implementation method is as follows:

[0160] First, monitor and confirm the status of the preheated standby node. Deploy a monitoring component, which uses a two-way verification mechanism that combines active query and passive reception of active reports from preheated standby nodes to collect key status indicators of preheated standby nodes in real time, including CPU usage, memory occupancy, network latency, etc. For example, the monitoring component sends a heartbeat request to the preheated standby node every 5 seconds, and receives status data actively reported by the preheated standby node every 10 seconds. Cross-validate the two parts of data to generate a status confirmation result. For example, the CPU usage reported by the preheated standby node is 70%, and the CPU usage queried by the monitoring component is 72%. If the deviation between the two is within a reasonable range, the CPU usage indicator is confirmed to be normal. Through this two-way verification mechanism, the health status of the preheated standby node can be judged more accurately.

[0161] Next, determine whether the preheated standby node meets the promotion conditions based on the status confirmation results. Preset promotion conditions include CPU usage less than 80%, memory usage less than 70%, and network latency less than 100ms. If all indicators meet the preset conditions, the preheated standby node is considered to meet the promotion conditions and the active / standby switching process can be started.

[0162] When the preheated standby node meets the promotion conditions, a switch transaction is started in Redis. This transaction is used to ensure the atomicity of the master-slave switch process, either all succeed or all fail, to prevent data inconsistency.

[0163] Obtain the write lock of the old master node and suspend new write requests to prevent data conflicts during data synchronization. Record the latest operation sequence number of the old master node, such as 1000, as the starting point for subsequent incremental data synchronization.

[0164] Build an incremental state synchronization task based on the latest operation sequence number. First, build a persistent operation log queue to record all write operation information, such as SET key value, INCR counter, etc. Then, determine the range of operation logs that need to be synchronized, that is, from operation sequence number 1000 to the current latest operation sequence number. Divide this operation log range into batches according to time segments, for example, every 10 seconds of operation log as a batch, and generate the operation batch sequence to be synchronized.

[0165] Perform incremental data migration on the operation batch sequence to be synchronized. Package the operation instructions in each operation batch and transmit them to the new primary node in a pipelined manner. For example, package the 10 operation instructions with operation numbers 1000 to 1010 into a batch request and send it to the new primary node. On the new primary node, convert the operation batch into a Redis command sequence, and execute the command sequence atomically through the EXEC command to ensure data consistency. Update the synchronization progress mark, for example, update the latest operation number currently synchronized to 1010. After all batches are synchronized, generate a data migration completion mark.

[0166] After data migration is completed, the master-slave switch operation is triggered. Submit the role information of the new master node and the end node information of the old master node to the Nacos registration server. For example, register the IP address and port number of the new master node as the new master node service address, and mark the service address of the old master node as offline. The Nacos service discovery mechanism will push this information to all subscribing clients, and the subscribing clients will update the local service routing table to route the request to the new master node.

[0167] Assign a globally unique version number, such as UUID, to the new master node, and build a lease-based protection mechanism. The new master node renews the contract with the Nacos registration server at a preset time interval, such as 30 seconds, and obtains the renewal information, which contains the assigned version number. If the Nacos registration server does not receive a renewal request from the new master node within a certain period of time, such as 60 seconds, it is considered that the new master node has failed, triggering the protection mechanism to terminate the service of the new master node to prevent brain split.

[0168] Before committing the Redis switch transaction, perform atomic confirmation. Check the data consistency of the new master node, such as verifying the integrity of the data through the checksum mechanism; verify the uniqueness of the version number to prevent multiple nodes from becoming the master node at the same time; confirm that the local service routing table has been updated to the latest status. If all checks pass, commit the switch transaction and complete the master-slave switch. If any check fails, roll back the switch transaction, and broadcast the master-slave switch failure information to the cluster through the service discovery function of Nacos so that other nodes can take corresponding measures.

[0169] The solution of this application can:

[0170] Improve system availability: By preheating the standby node and the incremental state synchronization mechanism, fast master-slave switching can be achieved, minimizing service interruption time and thus improving system availability. Ensure data consistency: Based on Redis's transaction characteristics and incremental state synchronization mechanism, the atomicity and consistency of data during the master-slave switching process are ensured to avoid data loss and inconsistency. Simplify operation and maintenance: Using Nacos's service discovery function, the broadcast of new master node information and the update of client routing tables can be automatically completed, simplifying operation and maintenance operations and improving operation and maintenance efficiency.

[0171] Figure 2 FIG. 1 is a schematic diagram of the structure of a distributed service intelligent master selection system based on dual-engine collaboration according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0172] The first unit is used to collect the processing capability index, network throughput, and resource utilization of the node to form a performance data matrix after receiving a registration request from a distributed service node, and to construct a node scoring function based on the performance data matrix; to perform a weighted combination of the calculation result of the node scoring function and the server physical clock value to generate a smart identification code for the node; to create a service instance in the Nacos registration center and write the smart identification code into an ordered set of the Redis storage engine, and to activate a two-layer event listener at the same time, wherein the two-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis;

[0173] The second unit is used for extracting the score sequence of existing nodes in the ordered set when the double-layer event listener detects a new node joining signal, analyzing the distribution trend of the score sequence, and calculating the optimal score increment step based on the distribution trend; combining the optimal score increment step with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the node running time; sorting the nodes based on the target score, and writing the sorting results into the ordered set to form a dynamic priority sequence;

[0174] The third unit is used to periodically extract the highest priority node from the dynamic priority sequence as the master node, and select the second priority node as the preheated standby node; when the two-layer event listener detects an abnormal load or fault signal of the master node, the fault state is confirmed through a voting arbitration mechanism; after confirmation, the preheated standby node is promoted to a new master node, and the data migration between the new and old master nodes is completed through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos.

[0175] According to a third aspect of the embodiments of the present invention,

[0176] An electronic device is provided, comprising:

[0177] processor;

[0178] a memory for storing processor-executable instructions;

[0179] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0180] According to a fourth aspect of the embodiments of the present invention,

[0181] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0182] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium on which are loaded computer-readable program instructions for executing various aspects of the present invention. Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solution from the scope of the technical solution of the embodiments of the present invention.

Claims

1. A distributed service intelligent master selection method based on dual-engine collaboration, characterized in that: include: After receiving a registration request from a distributed service node, the processing capability index, network throughput, and resource utilization of the node are collected to form a performance data matrix, and a node scoring function is constructed based on the performance data matrix; The calculation result of the node scoring function is weightedly combined with the server physical clock value to generate the node's intelligent identification code; a service instance is created in the Nacos registration center and the intelligent identification code is written into an ordered set of the Redis storage engine, and a double-layer event listener is activated at the same time, and the double-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis; When the double-layer event listener detects a signal that a new node has joined, it extracts the score sequence of existing nodes in the ordered set, analyzes the distribution trend of the score sequence, and calculates the optimal score increment step based on the distribution trend; combines the optimal score increment step with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the running time of the node; prioritizes the nodes based on the target score, and writes the ranking result into the ordered set to form a dynamic priority sequence; The highest priority node is periodically extracted from the dynamic priority sequence as the master node, and the second priority node is selected as the preheated standby node; when the double-layer event listener detects an abnormal load or fault signal of the master node, the fault state is confirmed through a voting arbitration mechanism; after confirmation, the preheated standby node is promoted to a new master node, and data migration between the new and old master nodes is completed through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos.

2. The method according to claim 1, characterized in that The calculation result of the node scoring function is weightedly combined with the server physical clock value to generate the node's intelligent identification code; a service instance is created in the Nacos registration center and the intelligent identification code is written into an ordered set of the Redis storage engine, and a double-layer event listener is activated at the same time. The double-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis, including: Input the performance data matrix into the node scoring function for calculation, and output the node scoring result; collect the server physical clock value, perform a weighted combination operation on the node scoring result and the server physical clock value, and generate a node intelligent identification code, wherein the node intelligent identification code includes a performance scoring part and a timing identification part; Create a distributed service instance in the Nacos registration center, write the smart identification code of the node into the extended attribute field of the distributed service instance; synchronously write the smart identification code of the node into an ordered set of the Redis storage engine, and the ordered set uses the performance score part of the smart identification code of the node as the sorting basis; the smart identification codes of the nodes in the ordered set are arranged in descending order according to the performance score; Activate a two-layer event listener, which includes a Nacos instance monitoring layer and a Redis data monitoring layer; the Nacos instance monitoring layer captures the state change signal of the service instance through an instance subscription mechanism; the Redis data monitoring layer captures the data structure change signal of the ordered set through a key space notification mechanism.

3. The method according to claim 1, characterized in that When the double-layer event listener detects a new node joining signal, extracts the score sequence of existing nodes in the ordered set, analyzes the distribution trend of the score sequence, and calculates the optimal score increment step according to the distribution trend, including: The two-layer event listener integrates the captured state change signal and data change signal into a unified event message stream; when the event message stream contains a new node joining signal, it reads the identification information and corresponding scores of all existing nodes from the ordered set of the Redis storage engine, extracts the corresponding scores to construct an ordered score sequence; calculates the difference between adjacent scores in the ordered score sequence to form a step sample set; performs a normality test on the step sample set to obtain the probability distribution characteristics of the step sample set; Selecting an optimal step length calculation strategy based on the probability distribution characteristics: when the probability distribution characteristics of the step length sample set meet the normal distribution requirements, using the maximum likelihood estimation method to calculate the reference step length value; when the probability distribution characteristics of the step length sample set do not meet the normal distribution requirements, using the kernel density estimation method to calculate the reference step length value; inputting the calculated reference step length value into the adaptive adjustment module; The adaptive adjustment module first obtains the current number of active nodes in the cluster, selects a capacity coefficient according to the number of active nodes, and multiplies the reference step value by the corresponding capacity coefficient to obtain a capacity adjustment step; Collecting the performance index of the newly added node and calculating the average performance level of the existing nodes in the cluster; comparing the performance index of the newly added node with the average performance level to obtain a performance ratio; calculating a performance coefficient according to the performance ratio; multiplying the capacity adjustment step by the performance coefficient to obtain a performance correction step; Extracting the generation timestamp of each score in the ordered score sequence, calculating the time interval between each score and the current moment; substituting the time interval into an exponential decay function to calculate a time weight, wherein the decay rate of the exponential decay function is dynamically adjusted according to the system load; performing weighted calculation on the performance correction step and the time weight to obtain a timing weighted step; Analyze historical score adjustment records, extract frequency characteristics and amplitude characteristics of score adjustment, and determine the lower limit threshold of step length based on the frequency characteristics; predict the growth trend of the number of nodes, and determine the upper limit threshold of step length according to the growth trend; limit the time series weighted step length between the lower limit threshold and the upper limit threshold to obtain the final optimal score increment step length.

4. The method according to claim 1, characterized in that: The optimal score increment step is combined with the intelligent identification code of the new node to generate a target score, wherein the target score shows a decaying trend as the node runs for a certain period of time; the nodes are prioritized based on the target score, and the ranking results are written into the ordered set to form a dynamic priority sequence, which includes: The optimal score increment step is used to characterize the priority interval between adjacent nodes; the intelligent identification code of the new node is parsed to obtain a performance score part and a node identification part; the performance score part is normalized to obtain a performance weight coefficient; The optimal score increment step is multiplied by the performance weight coefficient to obtain a benchmark score; the benchmark score is combined with the node identification part through a nonlinear mapping function to generate an initial target score of the node; the nonlinear mapping function ensures that the target scores of different nodes are globally unique and monotonically increasing; Obtaining a count value of the running time of the node from the node status monitoring module; calculating a time decay factor based on the count value of the running time, wherein the time decay factor adopts a piecewise exponential function: when the running time is less than one hour, the time decay rate is the largest, and when the running time exceeds twenty-four hours, the time decay rate is the smallest; Multiplying the initial target score of the node by the time decay factor to obtain a real-time target score adjusted by time decay; the real-time target score shows a decreasing trend as the running time of the node increases; Periodically collect the performance indicators of all online nodes in the system and calculate the system load standard deviation; dynamically adjust the priority update cycle according to the load standard deviation: when the load standard deviation is greater than the preset load threshold, shorten the update cycle to half of the reference cycle; when the load standard deviation is less than the preset load threshold, restore the update cycle to the reference cycle; At the beginning of each priority update cycle, the current version identifier of the ordered set of the Redis storage engine is obtained; the existing node priority sequence in the ordered set is read; the newly calculated real-time target score is combined with the node identifier to form a priority update record; the priority update record is inserted into the node priority sequence according to the score size to generate an updated dynamic priority sequence.

5. The method according to claim 1, characterized in that Periodically extracting the highest priority node from the dynamic priority sequence as the primary node, and selecting the secondary priority node as the preheated standby node includes: Selecting a node with the highest priority value from the dynamic priority sequence as a master node; selecting a node with the second highest priority value as a preheated standby node; constructing a double-buffered data synchronization channel with a write-ahead log function between the master node and the preheated standby node; the double-buffered data synchronization channel includes a first buffer for transmitting the write-ahead log and a second buffer for transmitting incremental data; Monitor the data volume of the first buffer and the second buffer; when the data volume of the first buffer reaches a first preset threshold, push the write-ahead log to the preheated standby node; when the data volume of the second buffer reaches a second preset threshold, push the incremental data to the preheated standby node; the preheated standby node maintains data synchronization based on the received write-ahead log and incremental data; Construct a lightweight memory snapshot mechanism based on a time window; at the beginning of each snapshot cycle, record the current memory state of the master node to obtain the current memory snapshot; compare the current memory snapshot with the memory snapshot of the previous cycle to extract the difference data of the memory state; transmit the difference data to the preheated standby node; the preheated standby node restores the real-time memory state of the master node based on the difference data.

6. The method according to claim 5, characterized in that When the two-layer event listener detects an abnormal load or fault signal of the master node, the fault status is confirmed through the voting arbitration mechanism including: Deploy a two-layer event listener, the local monitoring layer of the two-layer event listener collects the system load index, response delay index, and error rate index of the node; performs exponential moving average processing on the system load index to obtain a smoothed load value; calculates the response delay index in a sliding time window to obtain delay statistical characteristics; calculates the time series change characteristics of the error rate index to obtain an error trend value; The smoothed load value, the delay statistical feature, and the error trend value are input into a Z-score anomaly detection model, and the Z-score anomaly detection model calculates the mean and standard deviation of each indicator based on historical data; the deviation between the current indicator value and the corresponding mean is divided by the standard deviation to obtain an anomaly score; and the anomaly score is accumulated within the sliding time window to obtain a cumulative anomaly value; When the accumulated abnormal value exceeds the trigger threshold, the distributed voting arbitration process based on the Raft protocol is started; during the voting arbitration process, the trust weight is calculated based on the historical fault judgment accuracy of each node; the trust weight increases with the number of correct judgments and decreases with the number of wrong judgments; the voting result of each node is multiplied by its trust weight to obtain a weighted voting value, and when the weighted voting value exceeds two-thirds of the trigger threshold, the fault state of the master node is determined.

7. The method according to claim 1, characterized in that After confirmation, the preheated standby node is promoted to the new master node, and the data migration between the new and old master nodes is completed through the incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos, including: Monitor the status indicators of the preheated standby node through a heartbeat detection mechanism, the status indicators include CPU usage, memory usage, and network latency; use a two-way verification mechanism to confirm the status of the preheated standby node, the two-way verification mechanism includes active query by the monitoring component and active report by the preheated standby node, and generate a status confirmation result of the preheated standby node; Based on the status confirmation result, determine whether the preheated standby node meets the promotion condition; when the promotion condition is met, start a switching transaction in Redis to promote the preheated standby node to the new master node; obtain the write lock of the old master node, suspend new write requests, and record the latest operation sequence number of the old master node as the data synchronization starting point; Constructing an incremental state synchronization task based on the latest operation sequence number, the incremental state synchronization task comprising: constructing a persistent operation log queue to record write operation information, determining the operation log range to be synchronized, dividing the operation log range into batches according to time segments, and generating an operation batch sequence to be synchronized; Perform incremental data migration on the operation batch sequence to be synchronized: package and transmit the operation instructions in each operation batch to the new master node in a pipelined manner, convert the operation batch into a Redis command sequence on the new master node, execute the Redis command sequence atomically through the EXEC command, update the synchronization progress mark, and generate a data migration completion mark; Triggering active / standby switching based on the data migration completion flag: submitting the role information and end node information of the new active node to the Nacos registration server, pushing the role information and end node information of the new active node to the subscription client through the Nacos service discovery mechanism, and the subscription client updating the local service routing table; Assign a globally unique version number to the new master node, and build a lease-based protection mechanism: the new master node renews the contract with the Nacos registration server at a preset time interval to obtain renewal information; the renewal information includes the version number; if the renewal information is detected to be timed out, the protection mechanism is triggered to terminate the service of the new master node; Before the switching transaction is committed, atomic confirmation is performed based on the transaction characteristics of Redis: the data consistency of the new master node is checked, the uniqueness of the version number is verified, and the update status of the local service routing table is confirmed; when the atomic confirmation passes, the switching transaction is committed; when the atomic confirmation fails, the switching transaction is rolled back, and the new master node information is broadcast to the cluster through the service discovery function of Nacos.

8. A distributed service intelligent master selection system based on dual-engine collaboration, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used for collecting the processing capability index, network throughput and resource utilization of the node to form a performance data matrix after receiving a registration request of the distributed service node, and constructing a node scoring function based on the performance data matrix; The calculation result of the node scoring function is weightedly combined with the server physical clock value to generate the node's intelligent identification code; a service instance is created in the Nacos registration center and the intelligent identification code is written into an ordered set of the Redis storage engine, and a double-layer event listener is activated at the same time, and the double-layer event listener is used to capture the instance state change signal of Nacos and the data structure change signal of Redis; The second unit is used for extracting the score sequence of existing nodes in the ordered set when the double-layer event listener detects a new node joining signal, analyzing the distribution trend of the score sequence, and calculating the optimal score increment step based on the distribution trend; combining the optimal score increment step with the intelligent identification code of the new node to generate a target score, and the target score tends to decay with the node running time; sorting the nodes based on the target score, and writing the sorting results into the ordered set to form a dynamic priority sequence; The third unit is used to periodically extract the highest priority node from the dynamic priority sequence as the master node, and select the second priority node as the preheated standby node; when the two-layer event listener detects an abnormal load or fault signal of the master node, the fault state is confirmed through a voting arbitration mechanism; after confirmation, the preheated standby node is promoted to a new master node, and the data migration between the new and old master nodes is completed through an incremental state synchronization mechanism; the incremental state synchronization mechanism ensures the atomicity of the master-slave switching process based on the transaction characteristics of Redis, and broadcasts the new master node information to the cluster through the service discovery function of Nacos.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Micro-service architecture based on open source component

    CN112968960A

  • State monitoring method and system for middleware node

    CN118349419A