Remote replication link detection method and system
By obtaining the communication load status of the task node in the main storage system and the backup storage system, performing path planning and heartbeat detection path allocation, and building a link state view, the problem of poor replication link detection performance in the prior art is solved, and efficient link detection and load balancing are achieved.
Patent Information
- Application Number
- CN202510595777.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, the detection of the replication link between the main storage system and the backup storage system mostly adopts the traditional heartbeat detection mechanism, resulting in the exponential increase of the number of heartbeat detection requests when the number of nodes is large, which brings huge burden to the system and cannot effectively solve the performance problem of replication link detection.
By obtaining the unique identification information of all task nodes in the main storage system and the backup storage system and their current communication load status, performing path planning processing, calculating the current number of communication requests for each backup storage node, analyzing the communication load status and allocating the heartbeat detection path, generating a heartbeat detection path table, updating the communication status field in the path table, building a link state view, and judging the link health status based on the view.
Reduces frequent communication requests, improves query efficiency, reduces the system's response time under high load conditions, avoids redundancy and unbalanced load problems, improves link detection efficiency, and ensures load balancing of data synchronization tasks.
Smart Images

Figure CN120110959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a remote replication link detection method and system thereof. Background Art
[0002] With the continuous improvement of the degree of informatization, data backup and disaster recovery technologies play an increasingly important role in ensuring data security. As a key disaster recovery method, remote replication technology realizes data backup and recovery by regularly synchronizing data in the primary storage system to the backup storage system. In a distributed storage system, the replication link between the primary storage system and the backup storage system is a channel for data transmission, and the health status of the replication link directly affects the success of each backup task. However, in the prior art, the detection of the replication link between the primary storage system and the backup storage system mostly adopts the traditional heartbeat detection mechanism. This method often has a huge performance bottleneck, especially when the number of nodes is large, the number of heartbeat detection requests increases exponentially, which brings a huge burden to the system. The application effect in large-scale systems is poor, and the performance problem of replication link detection cannot be effectively solved.
[0003] Based on the above shortcomings of the prior art, a remote replication link detection method and system are urgently needed. Summary of the invention
[0004] The purpose of the present invention is to provide a remote replication link detection method and system to improve the above problems. In order to achieve the above purpose, the technical solution adopted by the present invention is as follows: In a first aspect, the present application provides a remote replication link detection method, comprising: Acquire first information, where the first information includes unique identification information of all task nodes in the primary storage system and the backup storage system and their current communication load status; Perform path planning based on the first information, calculate the current number of communication requests of each standby storage node in turn, and allocate a heartbeat detection path by analyzing the communication load state to obtain a heartbeat detection path table between the main storage node and the standby storage node; According to the heartbeat detection path table, triggering a heartbeat request based on the path table record and receiving feedback, updating the communication status field in the path table, and generating an updated heartbeat detection path table; According to the updated heartbeat detection path table, the communication status and timestamp information of each path are extracted to construct a link status view; According to the link status view, a link health status evaluation result is generated by analyzing the path communication status and judging the link health status based on preset rules.
[0005] In a second aspect, the present application also provides a remote replication link detection system, including: An acquisition module, configured to acquire first information, wherein the first information includes unique identification information of all task nodes in the primary storage system and the backup storage system and their current communication load status; A planning module performs path planning processing based on the first information, calculates the current number of communication requests of each standby storage node in turn, and allocates a heartbeat detection path by analyzing the communication load state, thereby obtaining a heartbeat detection path table between the main storage node and the standby storage node; An updating module is used to trigger a heartbeat request based on the path table record and receive feedback according to the heartbeat detection path table, update the communication status field in the path table, and generate an updated heartbeat detection path table; A construction module, used to extract the communication status and timestamp information of each path according to the updated heartbeat detection path table, and construct a link status view; The evaluation module is used to generate a link health status evaluation result according to the link status view by analyzing the path communication status and judging the link health status based on preset rules.
[0006] The beneficial effects of the present invention are: The present invention performs link status query based on the route status view cached locally by the node. Compared with the traditional method which requires real-time interaction, the present invention reduces frequent communication requests, improves query efficiency, and reduces the response time of the system under high load conditions. The present invention avoids the redundancy and unbalanced load problems caused by all nodes sending heartbeat requests to the remote end by planning a specific heartbeat detection path for each node and ensuring that the number of paths for each node is balanced. This not only improves the efficiency of link detection, but also ensures the load balance of data synchronization tasks, and optimizes the utilization of overall network resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0008] Figure 1 A schematic flow chart of a link detection method for remote replication described in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a link detection system for remote replication described in an embodiment of the present invention; Figure 3 A diagram of the virtual components of a storage system.
[0009] Markings in the figure: 901, acquisition module; 902, planning module; 903, update module; 904, construction module; 905, evaluation module. DETAILED DESCRIPTION
[0010] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0011] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0012] Embodiment 1: This embodiment provides a remote replication link detection method.
[0013] See also Figure 1 , the figure shows that the method includes steps S100 to S500.
[0014] Step S100: Acquire first information, where the first information includes unique identification information of all task nodes in the primary storage system and the backup storage system and their current communication load status; It is understandable that in actual applications, the primary storage system and the backup storage system usually adopt a distributed topology, and each node is responsible for processing different data synchronization tasks. Therefore, accurately obtaining the identification information and real-time communication load status of each node is crucial to optimizing system performance and achieving load balancing. To achieve this goal, this step requires real-time monitoring of each node in the system to collect their current load status, such as key performance indicators such as network latency, bandwidth usage, and request processing capabilities. The load status data not only reflects the real-time health status of the node, but also helps the path control component to reasonably plan the heartbeat path in subsequent steps to avoid performance bottlenecks caused by overloading of a single node.
[0015] Step S200: performing path planning processing based on the first information, calculating the current number of communication requests of each standby storage node in turn, and allocating a heartbeat detection path by analyzing the communication load state, so as to obtain a heartbeat detection path table between the primary storage node and the standby storage node; Further, step S200 includes step S210 to step S240.
[0016] Step S210: collecting and processing the load information of the standby storage node according to the first information, and obtaining a load status table of each standby storage node by counting the current number of communication requests of each standby storage node; Step S220: Initialize the heartbeat route reference count according to the load status table, and obtain an initialized heartbeat route reference count table by setting the initial reference count value of each backup storage node to 0; Step S230: allocating a heartbeat route to the first node of the primary storage system according to the heartbeat route reference count table, traversing all nodes of the backup storage system, selecting two nodes with the smallest reference counts as heartbeat targets, and updating the reference counts to obtain the heartbeat route allocation status of the first node of the primary storage system; Step S240: allocate heartbeat routes to the remaining nodes in the main storage system according to the heartbeat route allocation of the first node of the main storage system. Loop through each main storage node, select the two backup storage nodes with the smallest current reference counts for heartbeat route allocation, update the reference count, and obtain the heartbeat detection path table between the main storage system nodes and the backup storage nodes.
[0017] In this process, the path control component plays a key role in the main storage system. It is located on the same node as the link arbitration component and is responsible for selecting a suitable backup storage system node for each main storage system node to perform heartbeat detection path allocation. Specifically, the path control component dynamically adjusts the path allocation based on the communication load status of each backup storage node and the current number of communication requests. In the system, the load status of each backup storage node changes in real time, so the path control component needs to rely on these real-time data for calculations to ensure load balancing and allocate at least two suitable backup storage nodes for each task node of the main storage system for heartbeat detection based on the load conditions. This allocation not only prevents a single node from bearing too many requests, but also ensures a balanced load on each path, thereby optimizing the performance of the entire system.
[0018] Step S300: According to the heartbeat detection path table, trigger the heartbeat request based on the path table record and receive feedback, update the communication status field in the path table, and generate an updated heartbeat detection path table; Further, step S300 includes step S310 to step S340.
[0019] Step S310: triggering a heartbeat request according to the heartbeat detection path table, and obtaining a heartbeat request sending record by sending a heartbeat signal to each backup storage node recorded in the path table; Step S320: Send a record according to the heartbeat request, perform reception processing of the heartbeat feedback, and obtain a heartbeat response record by collecting heartbeat response data from each backup storage node; Step S330: Analyze and process the communication status according to the heartbeat response record, update the communication status of each heartbeat route by evaluating the response time and success rate, and obtain the latest communication status of each route; Step S340: update the heartbeat detection path table according to the latest communication status, and generate an updated heartbeat detection path table by integrating all updated communication status information.
[0020] It should be noted that, in this process, the heartbeat sending component and the heartbeat receiving component are respectively deployed in the main storage system and the standby storage system. The heartbeat sending component is located in each task node of the main storage system, responsible for triggering and sending heartbeat requests to the remote nodes of the standby storage system; the heartbeat receiving component is located in each task node of the standby storage system, receiving the heartbeat request from the main storage node and returning the response. Specifically, the heartbeat sending component sends a heartbeat request to the designated remote standby storage node according to the heartbeat detection path table as planned. Each request contains information such as timestamp and communication status, and the feedback after sending will be received by the heartbeat receiving component and transmitted back to the main storage system. According to the feedback result, the main storage system updates the communication status field in the path table through the heartbeat sending component, marking the latest communication status of each path, such as "online" or "offline". This step triggers and receives the heartbeat request in real time, so that the status of each path in the path table can be fed back and updated in real time. This dynamic update mechanism can timely reflect any network failure or load change in the system, and ensures accurate detection of the link health status. By using a one-way heartbeat and feedback mechanism, the overall performance and response speed of the system are optimized, while unnecessary communication requests are greatly reduced, reducing the burden on the system. The updated path table provides accurate and real-time basic data for subsequent link status evaluation and health detection, thereby ensuring the reliability and efficiency of data synchronization tasks.
[0021] Step S400: extract the communication status and timestamp information of each path according to the updated heartbeat detection path table, and construct a link status view; Further, step S400 includes step S410 to step S430.
[0022] Step S410: extract the communication status and timestamp according to the updated heartbeat detection path table, and obtain a data list including the local node ID, remote node ID, route status and last communication timestamp by analyzing the communication success, response time and last communication timestamp of each entry in the path table; Step S420: construct a data structure according to the data list, and obtain a 4-tuple list by integrating the local node ID, remote node ID, route status and last update time of each path into a 4-tuple format; Step S430: Generate a link state view according to the 4-tuple list, and generate a link state view by integrating all 4-tuples into a unified view framework.
[0023] It can be understood that in this step, the heartbeat view component is responsible for extracting the communication status (such as "online" or "offline") and timestamp (recording the last update time) of each path from the path table, and integrating these data into a structured data format, such as a 4-tuple form <local node ID, remote node ID, route status, last update time>. The updated heartbeat detection path table contains the communication status and timestamp data of each path. The heartbeat view component extracts the information of each path based on these data and generates a 4-tuple according to the status of each path. Each 4-tuple represents the complete status of a path, including the identification of the local node, the identification of the remote node, the status of the current path, and the update time of the last heartbeat request. In this way, the heartbeat view component not only establishes a detailed status view for each path, but also enables the link health of the entire system to be represented in a structured manner, which is convenient for subsequent analysis and decision-making.
[0024] Step S500: According to the link status view, a link health status evaluation result is generated by analyzing the path communication status and judging the link health status based on preset rules.
[0025] Further, step S500 includes step S510 to step S530.
[0026] Step S510: Analyze the path communication status according to the link status view, and obtain the current communication status statistics of each path by counting the communication status of each path in the statistical view; Step S520: Perform health status determination processing on communication status statistics based on preset health status rules to obtain a preliminary health status assessment of the entire link; Step S530: Perform an assessment process based on the preliminary health status assessment, and generate a link health status assessment result by integrating the status assessment of each path and the overall health status of the system.
[0027] In this step, the link arbitration component plays a role in the primary storage system, collecting link status views from all nodes and evaluating the communication status of each path according to preset rules. The communication status information of each path (such as "online" or "offline") and timestamp data will be used as the basis for evaluating link health. Specifically, the link arbitration component analyzes the status of each path in the link status view and generates a link health status evaluation result based on the preset health status judgment rules. The preset rules include: Health status: If the status of all paths is "online", the link health status is "healthy"; Sub-healthy status: If a path is "offline", but the number of offline paths does not exceed half, the link health status is "sub-healthy"; Fault status: If more than half of the paths are "offline", the link health status is "fault"; Offline status: If all paths are "offline", the link health status is "offline".
[0028] By applying these rules, the link arbitration component can determine the health of the entire replication link based on the communication status of each path and generate the final evaluation result. This evaluation result provides a basis for subsequent system optimization, fault response, and resource allocation. This step provides a real-time and accurate evaluation of the health of the entire link through a rule-based health status evaluation mechanism. Through automated health status judgment, potential problems in the link can be quickly identified, network failures can be responded to in a timely manner, and the reliability and fault tolerance of the system can be improved.
[0029] Embodiment 2: The difference between this embodiment and the above-mentioned embodiment 1 is that step S330 also includes steps S331 to S334.
[0030] Step S331, performing data aggregation processing according to the updated heartbeat detection path table, and obtaining a grouped heartbeat response data set by evaluating and classifying the heartbeat response patterns using cluster analysis; Specifically, cluster analysis algorithms, such as the K-means algorithm or the DBSCAN algorithm, are used in this step to evaluate and classify the heartbeat response patterns. Through cluster analysis, the heartbeat response data can be divided into several different groups, each representing a specific response pattern. For example, some nodes may exhibit a stable response pattern (low latency, high success rate), while other nodes may show an unstable pattern (high latency, low success rate). Cluster analysis helps the system better understand the health of the network by automatically identifying these patterns.
[0031] The cluster analysis process includes the following steps: Step S3311, data preparation: First, extract the communication status data of each path from the updated heartbeat detection path table and convert it into a data format suitable for cluster analysis. The data of each path may include indicators such as communication success rate and response time.
[0032] Step S3312, clustering model application: clustering algorithm is applied to group the data and automatically identify different heartbeat response patterns. Common clustering algorithms such as K-means divide the heartbeat response data into multiple categories according to the set number of clusters (K value), while DBSCAN automatically clusters according to the density of data points.
[0033] Step S3313, grouping result output: finally a grouped heartbeat response data set is obtained, each data set contains heartbeat paths with similar response patterns for subsequent analysis.
[0034] This step uses cluster analysis to perform in-depth pattern recognition and classification of heartbeat response data, allowing the system to understand the diversity of network status at a higher level. Through automated classification, the system can discover potential network problems and bottleneck areas, thereby more accurately locating and optimizing faults. In addition, this data-driven analysis method can dynamically adapt to changes in the network environment and has high scalability and application value in large-scale distributed systems.
[0035] Step S332: Perform success rate prediction processing according to the heartbeat response data set, predict the future success rate by applying a support vector machine, and obtain a prediction success rate model adjusted based on historical data and real-time feedback; In this process, the link arbitration component and the heartbeat view component work closely together to generate a success rate prediction model that can be adjusted in real time by inputting historical data extracted from the heartbeat response data set into the support vector machine model for training. In this step, the support vector machine is used to build a regression model to predict future changes in success rate by learning the success rate pattern of historical heartbeat response data. This step uses the support vector machine to predict the success rate of the heartbeat response data, effectively combining historical data with real-time feedback, thereby realizing dynamic prediction of the future link health status. Compared with traditional static evaluation methods, the support vector machine model can capture more complex patterns and provide more accurate success rate predictions. This prediction model not only improves the accuracy of link health status assessment, but also improves the system's early warning capability for possible future failures, and enhances the system's adaptability and fault tolerance.
[0036] Step S333: Optimize and analyze the response time according to the prediction success rate model, and obtain an optimized response time strategy by searching for the best response time configuration to minimize delay and maximize communication efficiency; It is understandable that in this step, the link arbitration component will use the output of the prediction success rate model as input, combined with optimization algorithms such as genetic algorithms, simulated annealing or particle swarm optimization, to optimize the response time of the system. This optimization can not only effectively reduce the response time and improve the real-time response capability of the system, but also reduce unnecessary delays and optimize the overall performance of the system while ensuring the stability and success rate of data synchronization.
[0037] Step S334: Dynamically update the communication status according to the response time strategy, analyze the success rate and response time of the communication data in real time and compare them with the preset performance threshold, adjust the status judgment standard based on the comparison result, and obtain the latest communication status of each route.
[0038] It is understandable that the link arbitration component evaluates the status of the current network link based on the communication data collected in real time, including the success rate and response time of the heartbeat request, and compares it with the pre-set performance threshold. These performance thresholds may include key indicators such as the upper limit of communication delay, response time and success rate, which are used to determine whether the path is in an "online", "sub-healthy" or "offline" state. Compared with the static state determination in traditional methods, the dynamic adjustment of thresholds can more accurately reflect the health status of the current network, identify problems in the link in a timely manner and make adjustments. This method not only enhances the real-time and flexibility of the system, but also improves the fault tolerance of the system, ensuring the efficient implementation of data synchronization tasks in complex network environments.
[0039] Embodiment 3: like Figure 3 As shown, this embodiment specifically relates to the configuration of task nodes in the primary storage system and the planning and updating of the heartbeat detection path in the backup storage system. This embodiment describes in detail how to optimize the link detection process and improve the performance and stability of the system by deploying arbitration components, control components, view components, and sending / receiving components in the primary storage system and the backup storage system. Among them, node1, node2, node3, node_n, and node_m represent node names respectively.
[0040] Suppose the primary storage system has m task nodes and the backup storage system has n task nodes.
[0041] In the primary storage system, node1 and node2 deploy arbitration components and control components, where arbitration component 1 is responsible for work and arbitration component 2 is in standby state. Each task node deploys a view component and a send component. In the backup storage system, node1 and node2 deploy arbitration components and control components, where arbitration component 1 is responsible for work and arbitration component 2 is in standby state. Each task node deploys a view component and a receive component.
[0042] The control component in the main storage system selects two heartbeat routes for node_1, node_2, and node_3. In the main storage system, arbitration component 1 regularly broadcasts query requests to view components 1 to m to query the remote node status view saved in the view component. Each sending component regularly triggers heartbeat detection, sends heartbeat requests to the remote end according to the previously planned heartbeat route, and feeds back the results to the local view component to build the remote node status view.
[0043] In the backup storage system, arbitration component 1 broadcasts query requests to view components 1 to m at regular intervals to query the remote node status view saved in the view component. When the receiving component receives the request sent by the primary storage system node, it notifies the view component to build the view and update the timestamp. When arbitration component 1 broadcasts the query view, if it detects that the timestamp of a line in the remote node status view has not been updated for more than 10 seconds, the line is marked as "offline".
[0044] Embodiment 4: like Figure 2 As shown, this embodiment provides a link detection system for remote replication, the system comprising: An acquisition module 901 is used to acquire first information, where the first information includes unique identification information of all task nodes in the primary storage system and the backup storage system and their current communication load status; The planning module 902 performs path planning processing based on the first information, calculates the current number of communication requests of each standby storage node in turn, and allocates the heartbeat detection path by analyzing the communication load state, thereby obtaining a heartbeat detection path table between the primary storage node and the standby storage node; An updating module 903 is used to trigger a heartbeat request based on the path table record and receive feedback, update the communication status field in the path table, and generate an updated heartbeat detection path table according to the heartbeat detection path table; A construction module 904 is used to extract the communication status and timestamp information of each path according to the updated heartbeat detection path table, and construct a link status view; The evaluation module 905 is used to generate a link health status evaluation result according to the link status view by analyzing the path communication status and judging the link health status based on preset rules.
[0045] In a specific embodiment of the present invention, the planning module 902 includes: A first planning unit, configured to collect and process the load information of the standby storage node according to the first information, and obtain a load status table of each standby storage node by counting the current number of communication requests of each standby storage node; The second planning unit is used to initialize the heartbeat route reference count according to the load status table, and obtain an initialized heartbeat route reference count table by setting an initial reference count value of each backup storage node to 0; The third planning unit is used to allocate a heartbeat route to the first node of the primary storage system according to the heartbeat route reference count table, traverse all nodes of the backup storage system, select two nodes with the smallest reference count as heartbeat targets, and update the reference count to obtain the heartbeat route allocation status of the first node of the primary storage system; The fourth planning unit is used to allocate heartbeat routes to the remaining nodes in the main storage system according to the heartbeat route allocation of the first node of the main storage system, by looping each main storage node and selecting the two backup storage nodes with the smallest current reference counts for heartbeat route allocation, updating the reference count, and obtaining the heartbeat detection path table between the main storage system node and the backup storage node.
[0046] In a specific implementation of the present invention, the updating module 903 includes: A first updating unit is used to trigger a heartbeat request according to the heartbeat detection path table, and obtain a heartbeat request sending record by sending a heartbeat signal to each backup storage node recorded in the path table; A second updating unit is used to send a record according to a heartbeat request, receive and process the heartbeat feedback, and obtain a heartbeat response record by collecting heartbeat response data from each backup storage node; The third updating unit is used to analyze and process the communication status according to the heartbeat response record, update the communication status of each heartbeat route by evaluating the response time and success rate, and obtain the latest communication status of each route; The fourth updating unit is used to update the heartbeat detection path table according to the latest communication status, and generate an updated heartbeat detection path table by integrating all updated communication status information.
[0047] In a specific embodiment of the present invention, the construction module 904 includes: The first construction unit is used to extract the communication status and timestamp according to the updated heartbeat detection path table, and obtain a data list including the local node ID, the remote node ID, the route status and the last communication timestamp by analyzing the communication success, response time and the last communication timestamp of each entry in the path table; The second construction unit is used for constructing a data structure according to the data list, by integrating the local node ID, the remote node ID, the route state and the last update time of each path into a 4-tuple format to obtain a 4-tuple list; The third construction unit is used to perform a link state view generation process according to the 4-tuple list, and generate a link state view by integrating all 4-tuples into a unified view framework.
[0048] In a specific embodiment of the present invention, the evaluation module 905 includes: A first evaluation unit is used to analyze the path communication state according to the link state view, and obtain the current communication state statistics of each path by counting the communication state of each path in the statistical view; A second evaluation unit performs health status determination processing on communication status statistics based on a preset health status rule to obtain a preliminary health status evaluation of the entire link; The third evaluation unit is used to perform evaluation processing according to the preliminary health status evaluation, and generate a link health status evaluation result by integrating the status evaluation of each path and the overall health status of the system.
[0049] In a specific implementation of the present invention, the third updating unit includes: A first aggregation unit is used to perform data aggregation processing according to the updated heartbeat detection path table, and obtain a grouped heartbeat response data set by evaluating and classifying the heartbeat response mode through cluster analysis; A first prediction unit is used to perform success rate prediction processing according to the heartbeat response data set, and predict the future success rate by applying a support vector machine to obtain a prediction success rate model based on historical data and real-time feedback adjustment; A first optimization unit is used to perform optimization analysis and processing of the response time according to the prediction success rate model, and obtain an optimized response time strategy by searching for the best response time configuration to minimize delay and maximize communication efficiency; The first adjustment unit is used to perform dynamic communication status update processing according to the response time strategy, analyze the success rate and response time of the communication data in real time and compare them with the preset performance threshold, adjust the status judgment standard based on the comparison result, and obtain the latest communication status of each route.
[0050] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A remote replication link detection method, characterized in that: include: Acquire first information, where the first information includes unique identification information of all task nodes in the primary storage system and the backup storage system and their current communication load status; Perform path planning based on the first information, calculate the current number of communication requests of each standby storage node in turn, and allocate a heartbeat detection path by analyzing the communication load state to obtain a heartbeat detection path table between the main storage node and the standby storage node; According to the heartbeat detection path table, triggering a heartbeat request based on the path table record and receiving feedback, updating the communication status field in the path table, and generating an updated heartbeat detection path table; According to the updated heartbeat detection path table, the communication status and timestamp information of each path are extracted to construct a link status view; According to the link status view, a link health status evaluation result is generated by analyzing the path communication status and judging the link health status based on preset rules.
2. A remote replication link detection method according to claim 1, characterized in that: Performing path planning based on the first information, calculating the current number of communication requests of each standby storage node in turn, and allocating a heartbeat detection path by analyzing the communication load state, to obtain a heartbeat detection path table between the primary storage node and the standby storage node, including: Collecting and processing the load information of the standby storage node according to the first information, and obtaining a load status table of each standby storage node by counting the current number of communication requests of each standby storage node; According to the load status table, initialization processing of the heartbeat route reference count is performed, and an initialized heartbeat route reference count table is obtained by setting an initial reference count value of each backup storage node to 0; According to the heartbeat route reference count table, a heartbeat route is allocated to the first node of the primary storage system, and by traversing all nodes of the backup storage system, two nodes with the smallest reference counts are selected as heartbeat targets, and the reference counts are updated to obtain the heartbeat route allocation status of the first node of the primary storage system; According to the heartbeat route allocation of the first node of the main storage system, heartbeat routes are allocated to the remaining nodes in the main storage system. By looping each main storage node and selecting the two backup storage nodes with the smallest current reference counts for heartbeat route allocation, the reference count is updated to obtain a heartbeat detection path table between the main storage system nodes and the backup storage nodes.
3. A remote replication link detection method according to claim 1, characterized in that: According to the heartbeat detection path table, triggering a heartbeat request based on the path table record and receiving feedback, updating the communication status field in the path table, and generating an updated heartbeat detection path table, including: According to the heartbeat detection path table, trigger processing of the heartbeat request is performed, and a heartbeat request sending record is obtained by sending a heartbeat signal to each backup storage node recorded in the path table; According to the heartbeat request sending record, receiving and processing the heartbeat feedback is performed, and the heartbeat response record is obtained by collecting the heartbeat response data from each backup storage node; Analyze and process the communication status according to the heartbeat response record, update the communication status of each heartbeat route by evaluating the response time and success rate, and obtain the latest communication status of each route; The heartbeat detection path table is updated according to the latest communication status, and an updated heartbeat detection path table is generated by integrating all updated communication status information.
4. A remote replication link detection method according to claim 1, characterized in that: According to the updated heartbeat detection path table, the communication status and timestamp information of each path are extracted to construct a link status view, including: According to the updated heartbeat detection path table, the communication status and timestamp extraction process is performed, and the communication success, response time and last communication timestamp of each entry in the path table are analyzed to obtain a data list including the local node ID, the remote node ID, the route status and the last communication timestamp; Performing a data structure construction process according to the data list, by integrating the local node ID, the remote node ID, the route status and the last update time of each path into a 4-tuple format, to obtain a 4-tuple list; A link state view is generated based on the 4-tuple list, and a link state view is generated by integrating all 4-tuples into a unified view framework.
5. A remote replication link detection method according to claim 1, characterized in that: According to the link status view, by analyzing the path communication status and judging the link health status based on preset rules, a link health status evaluation result is generated, including: Analyze the path communication status according to the link status view, and obtain the current communication status statistics of each path by counting the communication status of each path in the statistical view; Performing health status determination processing on the communication status statistics based on preset health status rules to obtain a preliminary health status assessment of the entire link; An assessment process is performed based on the preliminary health status assessment, and a link health status assessment result is generated by integrating the status assessment of each path and the overall health status of the system.
6. A remote replication link detection system, characterized in that: include: An acquisition module, configured to acquire first information, wherein the first information includes unique identification information of all task nodes in the primary storage system and the backup storage system and their current communication load status; A planning module performs path planning processing based on the first information, calculates the current number of communication requests of each standby storage node in turn, and allocates a heartbeat detection path by analyzing the communication load state, thereby obtaining a heartbeat detection path table between the main storage node and the standby storage node; An updating module is used to trigger a heartbeat request based on the path table record and receive feedback according to the heartbeat detection path table, update the communication status field in the path table, and generate an updated heartbeat detection path table; A construction module, used to extract the communication status and timestamp information of each path according to the updated heartbeat detection path table, and construct a link status view; The evaluation module is used to generate a link health status evaluation result according to the link status view by analyzing the path communication status and judging the link health status based on preset rules.
7. A remote replication link detection system according to claim 6, characterized in that: The planning module includes: a first planning unit, configured to collect and process the load information of the standby storage node according to the first information, and obtain a load status table of each standby storage node by counting the current number of communication requests of each standby storage node; A second planning unit is used to perform initialization processing of the heartbeat route reference count according to the load status table, and obtain an initialized heartbeat route reference count table by setting an initial reference count value of each backup storage node to 0; A third planning unit is used to allocate a heartbeat route to the first node of the primary storage system according to the heartbeat route reference count table, traverse all nodes of the backup storage system, select two nodes with the smallest reference count as heartbeat targets, and update the reference count to obtain the heartbeat route allocation status of the first node of the primary storage system; The fourth planning unit is used to allocate heartbeat routes to the remaining nodes in the main storage system according to the heartbeat route allocation of the first node of the main storage system, by looping each main storage node and selecting the two backup storage nodes with the smallest current reference counts for heartbeat route allocation, updating the reference count, and obtaining the heartbeat detection path table between the main storage system node and the backup storage node.
8. A remote replication link detection system according to claim 6, characterized in that: The update module includes: A first updating unit is used to trigger a heartbeat request according to the heartbeat detection path table, and obtain a heartbeat request sending record by sending a heartbeat signal to each backup storage node recorded in the path table; A second updating unit is used to send a record according to the heartbeat request, perform reception processing of the heartbeat feedback, and obtain a heartbeat response record by collecting heartbeat response data from each backup storage node; A third updating unit is used to analyze and process the communication status according to the heartbeat response record, update the communication status of each heartbeat route by evaluating the response time and success rate, and obtain the latest communication status of each route; The fourth updating unit is used to update the heartbeat detection path table according to the latest communication status, and generate an updated heartbeat detection path table by integrating all updated communication status information.
9. A remote replication link detection system according to claim 6, characterized in that: The building blocks include: The first construction unit is used to extract the communication status and timestamp according to the updated heartbeat detection path table, and obtain a data list including the local node ID, the remote node ID, the route status and the last communication timestamp by analyzing the communication success, response time and the last communication timestamp of each entry in the path table; a second construction unit, configured to construct a data structure according to the data list, by integrating the local node ID, the remote node ID, the route status and the last update time of each path into a 4-tuple format to obtain a 4-tuple list; The third construction unit is used to generate a link state view according to the 4-tuple list, and generate a link state view by integrating all 4-tuples into a unified view framework.
10. A remote replication link detection system according to claim 6, characterized in that: The evaluation module includes: A first evaluation unit, configured to analyze the path communication state according to the link state view, and obtain current communication state statistics of each path by counting the communication state of each path in the statistical view; A second evaluation unit, based on a preset health status rule, performs health status determination processing on the communication status statistics to obtain a preliminary health status evaluation of the entire link; The third evaluation unit is used to perform evaluation processing according to the preliminary health status evaluation, and generate a link health status evaluation result by integrating the status evaluation of each path and the overall health status of the system.
Citation Information
Patent Citations
Path communication quality detection method and device
CN103404080A
Cable-free cluster health state detection method
CN112486761A
Method and device for detecting health state of host link in storage area network
CN118152244A
Distributed heartbeat detection method and system
CN118784536A
Distributed network node heartbeat message transmission optimization method and device
CN119449676A
Cited By
Remote copy link processing method and device, medium and program product
CN120692215A