Method and apparatus for testing communication reliability, and electronic device and storage medium

CN117729131BActive Publication Date: 2026-10-09INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311606077.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2026-10-09
Estimated Expiration
2043-11-28

AI Technical Summary

Benefits of technology

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program/instructions stored thereon, which, when executed by a processor, implements the communication reliability testing method as described in the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117729131B_ABST
    Figure CN117729131B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a communication reliability test method and device, electronic equipment and storage medium, relates to the storage technology field, and can realize the communication reliability test suitable for the remote replication of a storage system. The method comprises the following steps: dividing data transmitted through each remote replication relationship in a storage system into a plurality of data streams; for each data stream, counting a first duration that a host delay of the data stream exceeds a first set value; in the case that the first duration of at least one data stream is not less than a second set value, outputting a first test result, wherein the first test result is used for representing that the data stream with the first duration not less than the second set value has a remote replication relationship persistence error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, and in particular to a method, apparatus, electronic device, and storage medium for testing communication reliability. Background Technology

[0002] To ensure the continuity, recoverability, and high availability of business data, remote disaster recovery and backup solutions have emerged, and remote copy (RC) technology is one of the key technologies in remote disaster recovery and backup solutions.

[0003] In remote replication application scenarios, there will be one or more remote replication relationships and their associated links in the storage system. Due to the long distance and many communication links, it is difficult to reasonably implement communication reliability testing using the current communication reliability testing system of multi-controller cluster storage systems (i.e., between close-range nodes). Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, electronic device, and storage medium for testing communication reliability, which can realize communication reliability testing suitable for remote replication of storage systems.

[0005] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide a method for testing communication reliability, the method comprising:

[0006] The data transmitted through various remote replication relationships in the storage system is divided into multiple data streams;

[0007] For each of the data streams, the duration during which the host latency of the data stream exceeds a first preset value is recorded.

[0008] If at least one of the data streams has a first duration of not less than a second set value, a first test result is output, wherein the first test result is used to characterize that the data stream with a first duration of not less than the second set value has a remote replication relationship persistence error.

[0009] Secondly, embodiments of this application provide a communication reliability testing apparatus, the apparatus comprising:

[0010] The first partitioning module is used to divide the data transmitted in the storage system through various remote replication relationships into multiple data streams;

[0011] The first statistics module is used to, for each data stream, calculate the first duration during which the host latency of the data stream exceeds a first preset value;

[0012] A first output module is configured to output a first test result when at least one of the data streams has a first duration of not less than a second set value. The first test result is used to characterize that the data stream with a first duration of not less than the second set value has a remote replication relationship persistence error.

[0013] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the communication reliability testing method as described in the first aspect.

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the communication reliability testing method as described in the first aspect.

[0015] Fifthly, embodiments of this application also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the communication reliability testing method as described in the first aspect.

[0016] As can be seen from the above technical solution, this application divides the data transmitted in the remote replication relationship into data streams. Based on the host latency of each data stream, a first set value and a second set value are set to conduct communication reliability tests on the remote replication relationship associated with each data stream by measuring the frequency of high-latency communication anomalies in the data streams. For the data streams that frequently exhibit high-latency communication anomalies, the test results of the remote replication relationship persistence error are promptly output to provide early warning, thereby realizing a communication reliability test applicable to remote replication. Attached Figure Description

[0017] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a two-site, three-center approach in a related technology.

[0019] Figure 2 A flowchart illustrating the implementation of a communication reliability testing method provided in this application embodiment;

[0020] Figure 3 A schematic diagram illustrating the cumulative bad cycle duration provided in this application embodiment;

[0021] Figure 4A schematic diagram illustrating a target duration provided in an embodiment of this application;

[0022] Figure 5 A schematic diagram of a communication reliability testing device provided in an embodiment of this application;

[0023] Figure 6 A schematic diagram of an electronic device provided in an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0025] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0027] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.

[0028] As digitalization progresses across industries, data is increasingly becoming the core of enterprise operations, and users are placing higher demands on the stability of storage systems that hold this data. While enterprises can utilize highly stable storage devices to build storage systems, they still cannot completely prevent irreparable damage to production systems caused by various natural disasters. To ensure the continuity, recoverability, and high availability of business data, remote disaster recovery and backup solutions have emerged, and remote replication technology is one of the key technologies within these solutions.

[0029] The current communication reliability testing system for multi-controller cluster storage systems is mainly applicable to storage systems with short link distances, few communication links, and a certain level of communication quality assurance. It mainly obtains the communication reliability test results by checking whether the communication timeout duration between nodes exceeds its corresponding threshold value.

[0030] In remote replication applications, storage systems may contain one or more partnerships, remote replication relationships, and associated links. Taking a three-data-centers-in-two-sites (3DC) scenario as an example... Figure 1 As shown, the storage system is distributed in two locations. One location deploys the production data center corresponding to the production volume and the same-city disaster recovery data center corresponding to the first disaster recovery volume. The other location deploys the off-site disaster recovery data center corresponding to the second disaster recovery volume. Partnership and remote replication relationships can be established between each pair of production data centers and disaster recovery data centers.

[0031] However, due to the long distances and numerous communication links associated with these partnerships and remote replication relationships, various problems such as signal attenuation and noise interference exist. Therefore, it is difficult to directly apply the communication reliability testing system of the multi-controller cluster storage system to reasonably test the remote replication communication reliability of this storage system.

[0032] To address the problems existing in the aforementioned related technologies, this application proposes a communication reliability testing scheme. For different abnormal situations in remote replication of storage systems, the scheme refines the indicators and threshold values ​​used in reliability testing and provides corresponding response measures for different abnormal situations, thereby realizing a comprehensive communication reliability test applicable to remote replication.

[0033] The following description, in conjunction with the accompanying drawings, details a communication reliability testing method, apparatus, electronic device, and storage medium provided in this application through some embodiments and application scenarios.

[0034] Firstly, see [the following] Figure 2 The diagram shown is an implementation flowchart of a communication reliability testing method provided in this application embodiment. The method may include the following steps:

[0035] Step S101: Divide the data transmitted in the storage system through various remote replication relationships into multiple data streams.

[0036] In practice, the data transmitted on the logical links (i.e., one or more remote replication links) where one or more remote replication relationships reside can be divided into a single data stream according to actual needs.

[0037] Step S102: For each data stream, calculate the first duration during which the host latency of the data stream exceeds a first set value.

[0038] In practice, for each data stream, the host latency of that data stream is monitored individually. If the host latency of that data stream exceeds a first preset value, it indicates that the data stream has experienced high-latency input / output (IO). The duration of the high-latency IO is recorded, and the duration of all high-latency IOs recorded for that data stream is calculated to obtain the first duration. For example, the sum of the durations of all high-latency IOs currently recorded for that data stream can be used as the first duration; alternatively, the difference between the sum of the durations of currently recorded high-latency IOs and the sum of the durations of currently recorded non-high-latency IOs (those corresponding to data stream host latency not exceeding the first preset value) can be used as the first duration.

[0039] Step S103: If the first duration of at least one of the data streams is not less than a second set value, output a first test result, wherein the first test result is used to characterize that the data stream with the first duration not less than the second set value has a remote replication relationship persistence error.

[0040] In practical implementation, for any data stream, if the first duration of the data stream is not less than the second set value, it indicates that the remote replication link associated with the data stream has experienced a large number of high-latency I / Os (i.e., there is a persistent high-latency communication anomaly in the remote replication relationship) within a certain period of time. Therefore, it is determined that the data stream has a communication anomaly with a persistent error in the remote replication relationship, and the first test result (such as outputting error 1920 for the data stream) is output to provide an early warning of the communication anomaly, thereby prompting timely manual intervention. Meanwhile, the remote replication links associated with other data streams can still perform disaster recovery services normally.

[0041] As can be seen from the above technical solution, this application divides the data transmitted in the remote replication relationship into data streams. Based on the host latency of each data stream, a first set value and a second set value are set to conduct communication reliability tests on the remote replication relationship associated with each data stream by measuring the frequency of high-latency communication anomalies in the data streams. For the data streams that frequently exhibit high-latency communication anomalies, the test results of the remote replication relationship persistence error are promptly output to provide early warning, thereby realizing a communication reliability test applicable to remote replication.

[0042] Optionally, in one embodiment, the above-mentioned method of calculating the first duration for which the host latency of each data stream exceeds a first preset value includes:

[0043] Based on the set duration of the first period, the second duration for which the host latency of the data stream exceeds the first set value in each of the first periods is statistically analyzed.

[0044] Each of the first cycles with a second duration not less than a third set value is defined as a second cycle, and each of the first cycles with a second duration less than the third set value is defined as a third cycle, wherein the third set value is the duration of a set percentage of a single first cycle.

[0045] The difference between the total duration of the second cycle and the total duration of the third cycle is determined as the first duration during which the host latency of the data stream exceeds the first set value.

[0046] In practice, for each data stream, the host latency of that data stream is monitored periodically. Based on the proportion of the second duration of host latency exceeding the first set value in the total duration of a single monitoring period (i.e., the first period), it is determined whether the monitoring period is a bad period with a large number of high-latency I / Os (i.e., belonging to the second period) or a good period without a large number of high-latency I / Os (i.e., belonging to the third period). Then, the difference between the total duration of each period in the currently determined second period and the total duration of each period in the currently determined third period is calculated. This calculated difference is used as the first duration and compared with the second set value. This can measure the frequency of bad periods occurring within a period of time, thereby accurately testing whether there is a persistent high-latency communication anomaly in the remote replication relationship associated with the data stream.

[0047] For example, refer to Figure 3 The diagram showing the cumulative bad cycle duration (i.e., the first duration) divides the data transmitted by each remote replication relationship into 16 data streams. The duration of the first cycle is set to 10 seconds, and the third setting is set to 1 / 3 of the duration of a single first cycle (i.e., 10 / 3 seconds). Each data stream is periodically monitored individually.

[0048] For any data stream, the first period corresponding to each first period is defined as a bad period if the host latency is greater than the first set value acmaxhostdelay for at least 1 / 3 of the period, and the first period in which the host latency is greater than the first set value acmaxhostdelay for less than 1 / 3 of the period is defined as a good period.

[0049] For each bad cycle that occurs in a data stream, the corresponding counter value `badintervalcount` is incremented by 1; for each good cycle that occurs, the counter value `badintervalcount` is decremented by 1. The cumulative bad cycle duration of a data stream is the corresponding counter value `badintervalcount` multiplied by the duration of the first cycle. When the cumulative bad cycle duration of a data stream is not less than the second setting `aclinktolerance`, a 1920 error (i.e., a remote replication persistence error) is triggered for that data stream. For example, if the second setting `aclinktolerance` is 30, then when the counter value `badintervalcount` of a data stream reaches 3 (with the first cycle duration set to 10 seconds), a 1920 error will be triggered.

[0050] The first setting, acmaxhostdelay, represents the maximum host delay after the remote replication operation's linktolerance counter starts counting, in milliseconds (ms), with a default value of 5. The second setting, aclinktolerance, represents the total duration of differential bandwidth that the remote replication operation can tolerate, in seconds, with a value range of 20-86400 and a default value of 300. This feature can be disabled by setting a value of 0.

[0051] Optionally, in one embodiment, outputting the first test result when the first duration of at least one of the data streams is not less than a second preset value includes:

[0052] If the first duration of at least one of the data streams is not less than a second set value, then, based on the busy level of each remote replication relationship associated with the target data stream, select one remote replication relationship from the associated remote replication relationships to stop, and output the first test result.

[0053] The target data stream is the data stream with a first duration not less than a second set value.

[0054] In practice, when the first duration of a data stream reaches the second set value, the first test result can be output for the data stream to prompt manual intervention, and the remote replication relationship associated with the data stream can be stopped at the same time to reduce the impact of high-latency IO on disaster recovery services.

[0055] When performing a stop operation on a remote replication relationship for a data stream, if the data stream is currently associated with only one remote replication relationship (e.g., the data stream contains only data transmitted by a single remote replication relationship), then the remote replication relationship is stopped directly. If the data stream is associated with multiple remote replication relationships (e.g., the data stream contains data transmitted by multiple remote replication relationships), then based on actual needs and the busy level of each remote replication relationship, one remote replication relationship is selected to be stopped to avoid significantly impacting the disaster recovery capability of the storage system by stopping multiple remote replication relationships at once.

[0056] For example, when a data stream is associated with multiple remote replication relationships, in order to eliminate persistent errors in the remote replication relationships as quickly as possible, one can choose to stop the remote replication relationship with the highest level of activity (e.g., the highest data transmission frequency); or one can choose to stop the remote replication relationship with the lowest level of activity and that is not idle, so as to reduce the high-latency I / O in the data stream while reducing the amount of transmitted data affected by the stopping operation of the remote replication relationship.

[0057] As one possible implementation, the step of selecting one remote replication relationship to stop based on the busy level of each remote replication relationship associated with the target data stream, when the first duration of at least one of the data streams is not less than a second preset value, includes:

[0058] If the first duration of at least one of the data streams is not less than a second preset value, detect whether the time interval between the time when the target data stream last stopped the remote replication relationship and the current time exceeds a fourth preset value.

[0059] If the time interval between the last time the remote replication relationship of the target data stream was stopped and the current time exceeds a fourth preset value, a remote replication relationship is selected from the associated remote replication relationships and stopped according to the busy level of each remote replication relationship associated with the target data stream.

[0060] In practical implementation, when the first duration of a data stream reaches the second preset value, it is detected whether the data stream (i.e., the target data stream) has stopped remote replication relationships within a previous period of time (e.g., 20 seconds). If the data stream has stopped remote replication relationships within a previous period of time (corresponding to the case where the time interval between the last time the target data stream stopped remote replication relationships and the current time does not exceed the fourth preset value), then the remote replication relationships associated with the data stream are maintained. If the data stream has not stopped remote replication relationships within a previous period of time (corresponding to the case where the time interval between the last time the target data stream stopped remote replication relationships and the current time exceeds the fourth preset value), then a remote replication relationship is selected to be stopped based on the busy level of each remote replication relationship associated with the data stream, so that only one remote replication relationship is stopped for a single data stream in a short period of time (e.g., 20 seconds), thereby ensuring the disaster recovery of the storage system.

[0061] Optionally, for a stopped remote replication relationship, the remote replication relationship can be restarted only after the response measures for the remote replication relationship persistence error have been completed (i.e., after the error repair operation has been completed), in order to avoid restarting the remote replication relationship too early and causing repeated reporting of the remote replication relationship persistence error.

[0062] It's important to note that remote replication persistence errors (i.e., error 1920) are not strictly a problem, but rather a state. Taking asynchronous remote replication as an example, the storage system periodically checks whether the host latency and its duration exceed set thresholds (such as the first, second, and third thresholds mentioned above) to perform communication reliability testing in response to error 1920. If the storage system's performance degrades or link congestion occurs, host latency exceeding the set threshold will occur frequently, resulting in error 1920.

[0063] Error 1920 can be caused by problems on the primary system, the secondary system, or the inter-system link. For example, it could be a component failure: a component becomes unavailable or its performance degrades due to maintenance operations, or the component's performance degrades to the point that it cannot maintain the remote replication relationship. It could also be an application problem: such as an application using remote replication sending performance requirement changes.

[0064] After organizing and analyzing the causes of error 1920, this application determined that error 1920 is strongly correlated with storage write speed, actual bandwidth, system parameters, RC partnership parameters, etc., and therefore proposes the following response measures for error 1920:

[0065] For the data stream with a first duration not less than a second set value, perform at least one of the following operations:

[0066] Item A-1: ​​Based on the host latency of the data stream, determine whether the setting of the first set value is reasonable, and if it is determined that the setting of the first set value is unreasonable, adjust the first set value based on the maximum value of the host latency of the data stream;

[0067] Item A-2: Determine whether the write speed of the storage system volume associated with the data stream is abnormal. If it is determined that the write speed of the storage system volume associated with the data stream is abnormal, process the storage system volume with abnormal write speed.

[0068] Item A-3: Determine whether the CPU utilization rate of the storage device associated with the data stream is abnormal; if it is determined that the CPU utilization rate of the storage device associated with the data stream is abnormal, process the storage device with the abnormal CPU utilization rate.

[0069] Item A-4: Determine whether there is an abnormality in the latency of the communication port between the storage cluster nodes associated with the data stream; if it is determined that there is an abnormality in the latency of the communication port between the storage cluster nodes associated with the data stream, process the communication port between the storage cluster nodes with the abnormal latency.

[0070] Item A-5: Determine whether there is an abnormality in the bandwidth between the main system and the auxiliary system associated with the data stream. If it is determined that there is an abnormality in the bandwidth between the main system and the auxiliary system associated with the data stream, process the main system and the auxiliary system where the bandwidth between the systems is abnormal.

[0071] Item A-1 is mainly used to detect whether error 1920 is caused by an unreasonable setting of the first setting value, and to adjust the unreasonable first setting value to fix error 1920.

[0072] For example, if the first setting is 5ms, frequent IO delays of 5 to 5.5ms will quickly trigger a 1920 error. At this time, the actual latency within the monitoring period (i.e., the first period) is observed through the runtime status tool to see if it is stable, and whether the actual latency exceeds the current maximum tolerance threshold for more than 1 / 3 of the period. If the actual latency within the monitoring period is stable, and / or the actual latency does not exceed the current maximum tolerance threshold for more than 1 / 3 of the period, it is determined that the setting of the first setting is unreasonable. Then, the current latency level is accepted and the current IO pressure is maintained, and the first setting is adjusted to a value slightly greater than the maximum value of the current actual latency, thereby realizing the adjustment of the first setting.

[0073] Item A-2 is mainly used to detect whether the 1920 error is caused by an abnormal write speed of the storage system volume (such as the primary system volume or the secondary system volume), and to process the storage system volume with abnormal write speed in order to fix the 1920 error.

[0074] For example, if the storage in a storage system can only accept a write speed of 75 megabytes per second (MBps), a write speed of 100 MBps can quickly fill the node cache, thus triggering a 1920 error. At this time, the write response time can be obtained from the volume read / write statistics of the storage system. If the background write response time is greater than the expected value, it can be determined that there is an anomaly in the write speed of the storage system volume (i.e., the 1920 error is very likely caused by slow writes on the storage side). Subsequently, storage device resources can be added to the storage system, or the storage device can be moved to a faster storage device, or the storage system can be migrated to a new storage system, thereby handling the storage system volume with anomaly in write speed.

[0075] Item A-3 is mainly used to detect whether the 1920 error is caused by abnormal CPU usage of the storage device, and to process the storage device with abnormal CPU usage in order to fix the 1920 error.

[0076] Understandably, since remote replication transmission accepts I / O and writes data to disk, both require CPU resource scheduling and execution. If the CPU of any node (i.e., storage device) is overloaded, it can affect all replication relationships, thus triggering error 1920. In this case, check the CPU utilization in the storage device monitoring statistics. If the CPU utilization increases when performance problems occur, it can be determined that the CPU utilization of the storage device is abnormal (i.e., error 1920 is very likely caused by excessive CPU utilization (i.e., resource exhaustion)). If the number of read / write operations per second (IOPS) on the primary and secondary systems has not increased, it is necessary to find the event causing the increase in CPU utilization (e.g., starting many Fiber Channels (FCs) and consider moving the event to a time of low load). If the IOPS on the primary or secondary system increases, it is necessary to upgrade the hardware to solve the overload, or add additional I / O groups to share the node load. This is how to handle storage devices with abnormal CPU utilization.

[0077] For item A-4, it is mainly used to detect whether the 1920 error is caused by an abnormality in the latency of the communication port between storage cluster nodes, and to process the communication port between storage cluster nodes with abnormal latency in order to fix the 1920 error.

[0078] For example, if an increased volume write latency on the auxiliary system, coupled with a constant background response time and low CPU utilization, triggers a 1920 error, check the port to local node send response time in the storage cluster node statistics. This value should ideally be less than 1ms. If it is greater than 1ms, it will result in a volume write latency of 3ms or more, which can easily lead to a 1920 error. Therefore, if the value is greater than 1ms, it can be determined that there is an anomaly in the latency of the communication port between storage cluster nodes. Subsequently, more bandwidth can be provided between nodes, a faster FC link can be used, or an additional port can be added to the channel between nodes. This is how to handle the communication port between storage cluster nodes with anomalies in latency.

[0079] Item A-5 is mainly used to detect whether the 1920 error is caused by an abnormality in the bandwidth between the main system and the auxiliary system, and to process the main system and auxiliary system with abnormal bandwidth between the systems in order to fix the 1920 error.

[0080] For example, if the available replication bandwidth between systems is 1 gigabit per second (Gbps), writing to a replicated volume at a speed of 150 MBps can easily trigger a 1920 error. In this case, it is necessary to check whether the data compression ratio has changed. For example, check the port to remote node send response time value in the statistics of the master system node. When the link is saturated, this value increases significantly. Therefore, if this value is greater than the set value, it can be determined that there is an anomaly in the bandwidth between the master system and the auxiliary system. Subsequently, the total available bandwidth can be increased, the replication amount can be reduced, or remote replication with volume modification can be considered. This is how to handle the master system and auxiliary system with anomalies in the bandwidth between the systems.

[0081] It should be noted that in practical applications, users can choose one or more response measures from items A-1 to A-5 based on experience to implement, thereby further improving the efficiency of root cause localization and repair of 1920 errors based on experience. Alternatively, users can implement the above response measures in the order of items A-1 to A-5 (which represents the order of the probability of anomaly existence from high to low). If an anomaly is found during the implementation of a certain response measure and is handled, the implementation of subsequent response measures can be stopped, thereby further improving the efficiency of root cause localization and repair of 1920 errors based on the probability of anomaly existence.

[0082] Optionally, in one embodiment, the method further includes:

[0083] For each remote replication link in the storage system, determine whether the target duration of the remote replication link exceeds the set time threshold corresponding to the target duration;

[0084] If the target duration of at least one of the remote replication links exceeds a set time threshold corresponding to the target duration, a second test result is output. The second test result is used to characterize that the remote replication link whose target duration exceeds the set time threshold corresponding to the target duration is unavailable.

[0085] The target duration includes at least one of the following:

[0086] Item B-1: The time taken for the message receiving end of the communication layer module associated with the remote replication link to receive the feedback result for the input / output request from the time it initiates the input / output request;

[0087] Item B-2: The time interval between two consecutive input / output requests received by the message sending end of the communication layer module associated with the remote replication link;

[0088] Item B-3: The time taken for the message sending end of the communication layer module associated with the remote replication link to complete the data block splicing from the time it receives the feedback result that the driver layer has successfully completed the direct memory access operation.

[0089] For item B-1, it can be called the total I / O process time (full name CL_SND_ANTI_DEADLOCK_TIMEOUT, abbreviated as CMD_TIMEOUT). For example... Figure 4As shown, the total IO process time refers to the duration of the entire process (i.e., the whole process) from the message receiving end (also known as the initiator end, or the communication process initiator end) of the Communication Layer (CL) module initiating an IO request, which is transmitted through the link to the other party node, and the feedback result issued by the message sending end (also known as the target end) of the other party node in response to the IO request, which is transmitted through the link to the initiator end of the local CL module.

[0090] If the entire IO process takes too long, such as exceeding its corresponding set time threshold (e.g., 5.2s), it indicates poor communication performance between the initiator and target ends of the associated remote replication link. Considering the upper limit on IO concurrency, to avoid this communication anomaly significantly impacting IO concurrency, the remote replication link can be deemed unavailable, and a second test result (e.g., an alarm 71889 for this remote replication link) can be output to warn of this communication anomaly and prompt timely manual intervention. Other remote replication links can still perform disaster recovery services normally. The feedback result can include status, or a combination of data and status; the set time threshold corresponding to the IO process timeout can be called the IO process timeout threshold.

[0091] For item B-2, it can be called the receive I / O request interval (full name CL_COMPLETION_ANTI_DEADLOCK_TIMEOUT, abbreviated as COMPLETION_TIMEOUT). Figure 4 As shown, the I / O request interval refers to the time interval between two I / O requests received by the target end of the CL module. Similar to detecting communication link blockages, timeouts, or other anomalies from the initiator end via item B-1, this I / O request interval allows the target end to detect such anomalies. Specifically, if the I / O request interval exceeds its corresponding set time threshold (e.g., 6 seconds), it indicates that the associated remote replication link is experiencing blockages, timeouts, or other anomalies. In this case, a second test result is output to provide an early warning of the communication anomaly, prompting timely manual intervention. The set time threshold corresponding to the I / O request interval can be referred to as the upper limit of the I / O request interval.

[0092] For item B-3, it can be called the data successfully sent time (full name CL_COMPLETION_DMA_TIMEOUT, abbreviated as DMA_TIMEOUT). For example... Figure 4As shown, the data successful transmission time refers to the time elapsed from when the target end of the CL module finishes assembling the data block, sets the Direct Memory Access (DMA) time start point, sends data to the driver layer, to receiving feedback from the driver layer indicating successful completion of the DMA operation. This data successful transmission time is mainly used to test whether the port resources of the sending end are sufficient, i.e., whether there is message congestion. When message transmission congestion occurs, the sending of messages will compete with resources, which is difficult to alleviate automatically in a short time, thus affecting subsequent messages to be sent and causing widespread message timeouts. Therefore, this embodiment of the application tests whether the port resources of the sending end of a single remote replication link are sufficient by setting the data successful transmission time and its corresponding set time threshold. If the data successful transmission time exceeds its corresponding set time threshold (e.g., 800ms), it indicates that the port resources of the sending end are insufficient. At this time, it can be determined that the remote replication link is unavailable, and a second test result is output to provide an early warning of this communication anomaly, thereby prompting timely manual intervention. The set time threshold corresponding to the data successful transmission time can be called the data successful transmission time threshold.

[0093] As one possible implementation, the step of outputting a second test result when the target duration of at least one of the remote replication links exceeds a set time threshold corresponding to the target duration includes:

[0094] If the target duration of at least one of the remote replication links exceeds the set time threshold corresponding to the target duration, the second test result is output, and the logical link of the communication layer associated with the target remote replication link is disconnected. After the set time, the disconnected logical link is reconnected.

[0095] The target remote replication link is the remote replication link whose target duration exceeds a set time threshold corresponding to the target duration.

[0096] In practical implementation, when the target duration of a remote replication link exceeds its corresponding set time threshold, it indicates that the link communication quality of the remote replication link (i.e., the target remote replication link) has deteriorated to an extreme level, and the communication reliability of the remote replication link cannot be guaranteed. At this time, the second test result is output, the communication anomaly is reported, and the logical link of the relevant CL is actively disconnected. From the perspective of the disaster recovery system, this means that a communication link of a remote replication relationship has been disconnected. After a set duration (e.g., 15 seconds), the system actively attempts to reconnect. The remote replication function module will continue to transmit data after reconnection. Under normal data pressure, the data synchronization of the primary and secondary systems can be completed quickly (short-term disconnection has little impact on business). Thus, self-repair is achieved by actively disconnecting the logical link.

[0097] It should be noted that although short-term interruptions in remote replication services (i.e., short-term link disconnections triggered by alarm 71889) have little impact on services, link disconnections sometimes cannot effectively repair alarm 71889. Therefore, when maintenance personnel discover that remote replication links have or frequently experience alarm 71889 or link disconnections, they still need to promptly rectify and investigate the issue.

[0098] After organizing and analyzing the causes of alarm 71889, this application determined that alarm 71889 is strongly correlated with port configuration, communication component performance, CPU utilization, storage write speed, and actual bandwidth. Therefore, the following response measures for alarm 71889 are proposed:

[0099] For remote replication links where the target duration exceeds a set time threshold corresponding to the target duration, perform at least one of the following operations:

[0100] Item C-1: Determine whether the port associated with the remote replication link is reused by other communications besides remote replication communication. If it is determined that the port associated with the remote replication link is reused by other communications besides remote replication communication, process the reused port.

[0101] Item C-2: Determine whether the performance of the communication components associated with the remote replication link is abnormal. If it is determined that the performance of the communication components associated with the remote replication link is abnormal, process the communication components with abnormal performance.

[0102] C-3: Determine whether the CPU utilization rate of the storage device associated with the remote replication link is abnormal. If it is determined that the CPU utilization rate of the storage device associated with the remote replication link is abnormal, process the storage device with the abnormal CPU utilization rate.

[0103] Item C-4: Determine whether the write speed of the auxiliary system volume associated with the remote replication link is abnormal. If it is determined that the write speed of the auxiliary system volume associated with the remote replication link is abnormal, process the auxiliary system volume with abnormal write speed.

[0104] Item C-5: Determine whether there is an abnormality in the bandwidth between the master system and the auxiliary system associated with the remote replication link. If it is determined that there is an abnormality in the bandwidth between the master system and the auxiliary system associated with the remote replication link, process the master system and the auxiliary system where the bandwidth between the systems is abnormal.

[0105] For item C-1, it is mainly used to detect whether the 71889 alarm is caused by the port associated with the remote replication link being reused by other communications besides remote replication communication, and to process the reused port to fix the 71889 alarm.

[0106] For example, when host I / O communication and remote replication communication share a port on a storage device, DMA timeouts may occur due to port resource sharing, triggering alarm 71889. In this case, it's necessary to log in to the switch and check the network port configuration associated with the remote replication link, primarily examining the zone information to check if host I / O ports and remote replication cluster communication ports are in the same zone. If so, it indicates port reuse for the remote replication link. The zone configuration can then be adjusted to isolate host I / O communication and remote replication cluster communication. Alternatively, smaller zones or port masks can be set to allow only one type of communication, thus handling the reused ports.

[0107] Item C-2 is mainly used to detect whether the 71889 alarm is caused by abnormal performance of communication components, and to process the communication components with abnormal performance in order to fix the 71889 alarm.

[0108] Understandably, as equipment ages and wears out over time, the performance of some communication components will degrade (e.g., reduced power of the FC port optical module), triggering alarm 71889. In this case, it's necessary to check the performance statistics of the remote replication link port in the storage system, including average IO latency, packet loss rate, and error rate. If abnormal performance indicators are found, the link component problem can be located (i.e., the communication component performance is determined to be abnormal). Subsequently, each communication component is checked one by one to pinpoint the specific device causing the problem. This specific device is then updated and maintained, and the performance statistics are checked again to see if they have returned to normal. This process effectively addresses the communication components exhibiting abnormal performance.

[0109] For item C-3, it is mainly used to detect whether the 71889 alarm is caused by abnormal CPU usage of the storage device, and to process the storage device with abnormal CPU usage to fix the 71889 alarm.

[0110] It's understandable that storage systems can experience CPU resource contention due to various service configurations, low device performance, or running on a virtual platform, leading to a significant increase in CPU utilization and triggering alarm 71889. In this case, it's necessary to use storage memory monitoring tools to check the CPU utilization of each core associated with the remote replication link. If the core CPU utilization exceeds 85% or approaches 100%, it will severely impact message sending and receiving in remote replication, causing increased communication latency and message sending stuttering. Therefore, if the CPU utilization exceeds a set value, it can be determined that the storage device's CPU utilization is abnormal. Subsequently, check if the IOPS on the primary and secondary systems has increased. If the IOPS has not increased, identify the event causing the increased CPU utilization and consider moving the event to a lower load period for processing. If the IOPS has increased, upgrade the hardware to resolve the overload or add additional IO groups to distribute the node load. This process effectively addresses storage devices with abnormal CPU utilization.

[0111] For item C-4, it is mainly used to detect whether the 71889 alarm is caused by abnormal write speed of the secondary system volume, and to process the secondary system volume with abnormal write speed to fix the 71889 alarm.

[0112] For example, if the storage in a storage system can only accept a write speed of 75MBps, a write speed of 100MBps will quickly fill the node cache, triggering a 71889 alarm. In this case, it's necessary to obtain the Backend Write Response Time (HRRT) value from the secondary system volume statistics. If this value is greater than the expected value, it can be determined that the write speed of the secondary system volume is abnormal (i.e., the 71889 alarm is very likely caused by slow secondary system storage). Subsequently, the number of volumes in the storage pool in the secondary system can be increased, moved to faster storage, or migrated to a new storage system, thereby handling the secondary system volume with abnormal write speed.

[0113] For item C-5, it is mainly used to detect whether the 71889 alarm is caused by an abnormality in the bandwidth between the main system and the auxiliary system, and to process the main system and auxiliary system with abnormal bandwidth between the systems in order to fix the 71889 alarm.

[0114] For example, if the available replication bandwidth between systems is 1Gbps, writing to the replicated volume at a speed of 150MBps can easily trigger alarm 71889. In this case, it is necessary to check whether the data compression ratio has changed. For example, check the value of the port to remote node send response time in the statistics of the master system node. When the link is saturated, this value increases significantly. Therefore, if this value is greater than the set value, it can be determined that there is an anomaly in the bandwidth between the master system and the auxiliary system. Subsequently, the total available bandwidth can be increased, or the replication amount can be reduced, or remote replication with change volume can be considered. This is how to handle the master system and auxiliary system with anomalies in the bandwidth between systems.

[0115] It should be noted that in practical applications, users can choose one or more response measures from items C-1 to C-5 based on experience to perform, thereby further improving the efficiency of root cause location and repair of 71889 alarms based on experience. Alternatively, users can execute the above response measures in the order of items C-1 to C-5 (which represents the order of the probability of anomaly existence from high to low). If an anomaly is found during the execution of a certain response measure and the anomaly is handled, the execution of subsequent response measures can be stopped, thereby further improving the efficiency of root cause location and repair of 71889 alarms based on the probability of anomaly existence.

[0116] Optionally, in one embodiment, the method further includes:

[0117] If at least one partnership in the storage system meets the set conditions, a third test result is output, which is used to characterize that the connection of the remote cluster configured by the partnership that meets the set conditions has been lost.

[0118] The setting conditions include at least one of the following:

[0119] Item D-1: The communication layer of the storage system does not contain a partnership link associated with the partnership;

[0120] Item D-2: The remote replication module associated with the partnership in the storage system cannot communicate bidirectionally.

[0121] For item D-1, if the partnership condition is met, it means that the associated CL does not have a partnership link. This is generally caused by a link logout, such as a physical link disconnection (FC line) or a logical link disconnection due to alarm 71889. If all remote replication links (i.e., all remote replication logical links of the CL) logout, alarm 987301 will be triggered.

[0122] For item D-2, if the partnership meets the conditions, it means that the associated CL link is normal, but the RC module cannot communicate bidirectionally. This may be caused by the RC module itself, or it may be caused by error 1920 causing all remote replication relationships to be disconnected.

[0123] In practice, the above-mentioned conditions are used to test whether the partnerships in the storage system are abnormal (i.e., whether the partnerships exist). If a partnership does not exist, it means that the storage system has caused abnormal situations such as link logout or disconnection of all remote replication relationships due to reduced reliability of remote replication communication. This results in the storage system losing the connection to the configured remote cluster (or system). At this time, a third test result is output (such as outputting a 987301 alarm for the partnership) to warn of the communication abnormality and prompt timely manual intervention. Other partnerships can still carry out disaster recovery services normally.

[0124] After organizing and analyzing the causes of alarm 987301, this application determined that alarm 987301 is strongly correlated with remote replication and remote replication links. Therefore, the following response measures are proposed for alarm 987301:

[0125] For partnerships that meet the defined criteria, perform at least one of the following operations:

[0126] E-1: Determine whether the remote replication link associated with the partnership is unavailable. If the remote replication link associated with the partnership is unavailable, process the unavailable remote replication link.

[0127] E-2: Determine whether there is a remote replication relationship persistence error in the data stream associated with the partnership. If it is determined that there is a remote replication relationship persistence error in the data stream associated with the partnership, process the data stream with the remote replication relationship persistence error.

[0128] The E-1 item is mainly used to detect whether the 987301 alarm is caused by the unavailability of the remote replication link (i.e., the 71889 alarm), and to fix the 987301 alarm by handling the anomaly that caused the 71889 alarm.

[0129] For example, for a partnership that is associated with two remote replication links, the adjacent 71889 error records can be checked based on the list of alarm error items to confirm whether the remote replication links have been completely disconnected. If both remote replication links are disconnected due to the occurrence of 71889 alarms, it can be determined that the remote replication links associated with the partnership are unavailable, and then the response measures for 71889 alarms in the above embodiments can be processed.

[0130] For item E-2, it is mainly used to detect whether the 987301 alarm is caused by a persistent error in the remote replication relationship of the associated data stream (i.e., error 1920), and to fix the 987301 alarm by handling the anomaly that caused the error 1920.

[0131] For example, if a data stream associated with a partnership has a large number of high-latency I / O operations that trigger a 1920 error, the associated RC module may be unable to communicate bidirectionally due to the high I / O latency, thus triggering a 987301 alarm. In this case, the adjacent 1920 error records can be checked based on the alarm error item list to confirm whether the data stream has triggered a 1920 error. If the data stream triggers a 1920 error, it is determined that the data stream associated with the partnership has a remote replication relationship persistence error, and then the response measures for the 1920 error in the above embodiment are followed.

[0132] It should be noted that in practical applications, users can choose one or two response measures from items E-1 to E-2 based on experience to perform, thereby further improving the efficiency of root cause location and repair of the 987301 alarm based on experience. Alternatively, users can execute the above response measures in the order of items E-1 to E-2 (which represents the order of the probability of anomaly existence from high to low). If an anomaly is found during the execution of a certain response measure and the anomaly is handled, the execution of subsequent response measures can be stopped, thereby further improving the efficiency of root cause location and repair of the 987301 alarm based on the probability of anomaly existence.

[0133] Based on the above embodiments, this application proposes a comprehensive detection and processing scheme to address reliability issues in remote replication communication of storage systems. This scheme performs threshold verification on a large number of high-latency I / O operations in the remote replication communication link, issuing warnings for remote replication relationships exceeding the threshold to prompt manual intervention, while other remote replication relationships can continue to communicate normally. When a more serious reliability problem occurs, such as high-latency I / O timeouts exceeding a higher threshold, an error alarm is generated, and the relevant link is automatically disconnected and automatically reconnected within a short period. After reconnection, remote replication communication resumes from the breakpoint, avoiding impact on disaster recovery services. When a surge in communication I / O causes link port resources to be exhausted, resulting in communication congestion, a short-term disconnection will also occur. When two or all redundant remote replication links are disconnected, a higher-level error alarm is generated. In the above testing process, most of the processing is automatic, requiring only manual equipment testing or business pressure analysis. When alarms occur frequently, or when remote replication links are reported as unavailable, the manufacturer can intervene.

[0134] In summary, this application tested the reliability issues that may occur during the remote replication communication process from multiple perspectives, and improved the communication quality of the remote replication link through alarm prompts and self-healing methods, thereby improving the reliability and stability of disaster recovery data transmission. Furthermore, this application provides specific response and handling measures for different detected problems, providing effective guidance for on-site operation and maintenance personnel, improving the efficiency of on-site problem handling, and enhancing customer satisfaction.

[0135] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of this application.

[0136] Secondly, embodiments of this application provide a communication reliability testing apparatus, such as... Figure 5 As shown, the device includes:

[0137] The first partitioning module is used to divide the data transmitted in the storage system through various remote replication relationships into multiple data streams;

[0138] The first statistics module is used to, for each data stream, calculate the first duration during which the host latency of the data stream exceeds a first preset value;

[0139] A first output module is configured to output a first test result when at least one of the data streams has a first duration of not less than a second set value. The first test result is used to characterize that the data stream with a first duration of not less than the second set value has a remote replication relationship persistence error.

[0140] Optionally, the first statistics module includes:

[0141] The first statistics submodule is used to calculate the second duration for which the host latency of the data stream exceeds a first set value in each of the first periods, based on the set duration of the first period.

[0142] The second statistics submodule is used to determine each of the first cycles with a second duration not less than a third set value as the second cycle, and to determine each of the first cycles with a second duration less than the third set value as the third cycle, wherein the third set value is the duration of a set percentage of a single first cycle.

[0143] The third statistics submodule is used to determine the difference between the total duration of the second period and the total duration of the third period as the first duration during which the host latency of the data stream exceeds the first set value.

[0144] Optionally, the first output module includes:

[0145] The first output submodule is configured to, when the first duration of at least one of the data streams is not less than a second set value, select one remote replication relationship from the associated remote replication relationships to stop, based on the busy level of each remote replication relationship associated with the target data stream, and output the first test result.

[0146] The target data stream is the data stream with a first duration not less than a second set value.

[0147] Optionally, the first output submodule includes:

[0148] The second output submodule is used to detect whether the time interval between the last time the remote replication relationship of the target data stream stopped and the current time exceeds a fourth preset value, provided that the first duration of at least one of the data streams is not less than a second preset value.

[0149] The third output submodule is used to select one of the associated remote replication relationships to stop when the time interval between the last time the remote replication relationship was stopped and the current time exceeds a fourth preset value, based on the busy level of each remote replication relationship associated with the target data stream.

[0150] Optionally, the device further includes:

[0151] The first processing module is used to determine, for each remote replication link in the storage system, whether the target duration of the remote replication link exceeds a set time threshold corresponding to the target duration.

[0152] The second output module is used to output a second test result when the target duration of at least one of the remote replication links exceeds a set time threshold corresponding to the target duration. The second test result is used to characterize that the remote replication link whose target duration exceeds the set time threshold corresponding to the target duration is unavailable.

[0153] The target duration includes at least one of the following:

[0154] The time taken for the message receiving end of the communication layer module associated with the remote replication link to receive the feedback result for the input / output request from the time it initiates the input / output request.

[0155] The time interval between two consecutive input / output requests received by the message sending end of the communication layer module associated with the remote replication link;

[0156] The time taken from the completion of data block splicing to receiving feedback from the driver layer indicating successful completion of direct memory access operation at the message sending end of the communication layer module associated with the remote replication link.

[0157] Optionally, the second output module includes:

[0158] The fourth output submodule is used to output the second test result and disconnect the logical link of the communication layer associated with the target remote replication link when the target duration of at least one of the remote replication links exceeds the set time threshold corresponding to the target duration, and reconnect the disconnected logical link after the set time.

[0159] The target remote replication link is the remote replication link whose target duration exceeds a set time threshold corresponding to the target duration.

[0160] Optionally, the device further includes:

[0161] The first execution module is configured to perform at least one of the following operations on the remote replication link where the target duration exceeds a set time threshold corresponding to the target duration:

[0162] Determine whether the port associated with the remote replication link is reused by other communications besides remote replication communication. If it is determined that the port associated with the remote replication link is reused by other communications besides remote replication communication, process the reused port.

[0163] Determine whether the performance of the communication components associated with the remote replication link is abnormal. If it is determined that the performance of the communication components associated with the remote replication link is abnormal, process the communication components with abnormal performance.

[0164] Determine whether the CPU utilization rate of the storage device associated with the remote replication link is abnormal. If it is determined that the CPU utilization rate of the storage device associated with the remote replication link is abnormal, process the storage device with abnormal CPU utilization rate.

[0165] Determine whether the write speed of the auxiliary system volume associated with the remote replication link is abnormal. If it is determined that the write speed of the auxiliary system volume associated with the remote replication link is abnormal, process the auxiliary system volume with abnormal write speed.

[0166] Determine whether there is an anomaly in the bandwidth between the master system and the auxiliary system associated with the remote replication link. If it is determined that there is an anomaly in the bandwidth between the master system and the auxiliary system associated with the remote replication link, process the master system and the auxiliary system where the bandwidth between the systems is abnormal.

[0167] Optionally, the device further includes:

[0168] The third output module is used to output a third test result when at least one partnership in the storage system meets the set conditions. The third test result is used to characterize that the connection of the remote cluster configured by the partnership that meets the set conditions has been lost.

[0169] The setting conditions include at least one of the following:

[0170] The communication layer of the storage system does not contain the partnership link associated with the partnership.

[0171] The remote replication module associated with the partnership in the storage system cannot communicate bidirectionally.

[0172] Optionally, the device further includes:

[0173] The second execution module is configured to perform at least one of the following operations for the partnership that meets the set conditions:

[0174] Determine whether the remote replication link associated with the partnership is unavailable. If the remote replication link associated with the partnership is unavailable, process the unavailable remote replication link.

[0175] Determine whether there is a remote replication relationship persistence error in the data stream associated with the partnership. If it is determined that there is a remote replication relationship persistence error in the data stream associated with the partnership, process the data stream with the remote replication relationship persistence error.

[0176] Optionally, the device further includes:

[0177] The third execution module is configured to perform at least one of the following operations on the data stream for which the first duration is not less than the second set value:

[0178] Based on the host latency of the data stream, determine whether the setting of the first set value is reasonable, and if it is determined that the setting of the first set value is unreasonable, adjust the first set value according to the maximum value of the host latency of the data stream;

[0179] Determine whether the write speed of the storage system volume associated with the data stream is abnormal. If it is determined that the write speed of the storage system volume associated with the data stream is abnormal, process the storage system volume with abnormal write speed.

[0180] Determine whether the CPU utilization rate of the storage device associated with the data stream is abnormal. If it is determined that the CPU utilization rate of the storage device associated with the data stream is abnormal, process the storage device with the abnormal CPU utilization rate.

[0181] Determine whether there is an abnormality in the latency of the communication port between the storage cluster nodes associated with the data stream. If it is determined that there is an abnormality in the latency of the communication port between the storage cluster nodes associated with the data stream, process the communication port between the storage cluster nodes with abnormal latency.

[0182] Determine whether there is an anomaly in the bandwidth between the main system and the auxiliary system associated with the data stream. If it is determined that there is an anomaly in the bandwidth between the main system and the auxiliary system associated with the data stream, process the main system and the auxiliary system where the bandwidth between the systems is abnormal.

[0183] As can be seen from the above technical solution, this application divides the data transmitted in the remote replication relationship into data streams. Based on the host latency of each data stream, a first set value and a second set value are set to conduct communication reliability tests on the remote replication relationship associated with each data stream by measuring the frequency of high-latency communication anomalies in the data streams. For the data streams that frequently exhibit high-latency communication anomalies, the test results of the remote replication relationship persistence error are promptly output to provide early warning, thereby realizing a communication reliability test applicable to remote replication.

[0184] It should be noted that the device embodiments are similar to the method embodiments, so the description is relatively simple. For relevant details, please refer to the method embodiments.

[0185] This application also provides an electronic device, see embodiments thereof. Figure 6 , Figure 6 This is a schematic diagram of the electronic device proposed in an embodiment of this application. Figure 6 As shown, the electronic device 100 includes a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus for communication. The memory 110 stores a computer program that can run on the processor 120 to implement the steps in the communication reliability testing method disclosed in the embodiments of this application.

[0186] This application also provides a computer-readable storage medium, see [link to relevant documentation]. Figure 7 , Figure 7 This is a schematic diagram of a computer-readable storage medium proposed in an embodiment of this application. Figure 7 As shown, a computer-readable storage medium 200 stores a computer program / instructions 210, which, when executed by a processor, implements the steps in the communication reliability testing method disclosed in the embodiments of this application.

[0187] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps in the communication reliability testing method disclosed in this application.

[0188] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0189] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0190] This application describes embodiments of methods, systems, devices, storage media, and program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0191] The present application provides a detailed description of a communication reliability testing method, apparatus, electronic device, and storage medium. Specific examples have been used to illustrate the principles and implementation methods of the present application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present application. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present application. Therefore, the content of this specification should not be construed as a limitation of the present application.

Claims

1. A method for testing communication reliability, characterized in that, The method includes: The data transmitted through various remote replication relationships in the storage system is divided into multiple data streams; For each of the data streams, the duration during which the host latency of the data stream exceeds a first preset value is recorded. If the first duration of at least one of the data streams is not less than a second set value, a first test result is output. The first test result is used to characterize that the data stream with the first duration exceeding the second set value has a remote replication relationship persistence error. Wherein, for each of the data streams, the first duration during which the host latency of the data stream exceeds a first preset value includes: Based on the set duration of the first period, the second duration during which the host latency of the data stream exceeds the first set value in each of the first periods is statistically analyzed. Each of the first cycles with a second duration not less than a third set value is defined as a second cycle, and each of the first cycles with a second duration less than the third set value is defined as a third cycle, wherein the third set value is the duration of a set percentage of a single first cycle. The difference between the total duration of the second cycle and the total duration of the third cycle is determined as the first duration during which the host latency of the data stream exceeds the first set value. Wherein, the step of outputting a first test result when the first duration of at least one of the data streams is not less than a second preset value includes: If the first duration of at least one of the data streams is not less than a second set value, then, based on the busy level of each remote replication relationship associated with the target data stream, select one remote replication relationship from the associated remote replication relationships to stop, and output the first test result. The target data stream is the data stream with a first duration not less than a second set value.

2. The method according to claim 1, characterized in that, When the first duration of at least one of the data streams is not less than a second preset value, selecting one remote replication relationship to stop based on the busy level of each remote replication relationship associated with the target data stream includes: If the first duration of at least one of the data streams is not less than a second preset value, detect whether the time interval between the time when the target data stream last stopped the remote replication relationship and the current time exceeds a fourth preset value. If the time interval between the last time the remote replication relationship of the target data stream was stopped and the current time exceeds a fourth preset value, a remote replication relationship is selected from the associated remote replication relationships and stopped according to the busy level of each remote replication relationship associated with the target data stream.

3. The method according to claim 1, characterized in that, The method further includes: For each remote replication link in the storage system, determine whether the target duration of the remote replication link exceeds the set time threshold corresponding to the target duration; If the target duration of at least one of the remote replication links exceeds a set time threshold corresponding to the target duration, a second test result is output. The second test result is used to characterize that the remote replication link whose target duration exceeds the set time threshold corresponding to the target duration is unavailable. The target duration includes at least one of the following: The time taken for the message receiving end of the communication layer module associated with the remote replication link to receive the feedback result for the input / output request from the time it initiates the input / output request. The time interval between two consecutive input / output requests received by the message sending end of the communication layer module associated with the remote replication link; The time taken from the completion of data block splicing to receiving feedback from the driver layer indicating successful completion of direct memory access operation at the message sending end of the communication layer module associated with the remote replication link.

4. The method according to claim 3, characterized in that, If the target duration of at least one of the remote replication links exceeds a set time threshold corresponding to the target duration, a second test result is output, including: If the target duration of at least one of the remote replication links exceeds the set time threshold corresponding to the target duration, the second test result is output, and the logical link of the communication layer associated with the target remote replication link is disconnected. After the set time, the disconnected logical link is reconnected. The target remote replication link is the remote replication link whose target duration exceeds a set time threshold corresponding to the target duration.

5. The method according to claim 3, characterized in that, After outputting the second test result, the method further includes: For remote replication links where the target duration exceeds a set time threshold corresponding to the target duration, perform at least one of the following operations: Determine whether the port associated with the remote replication link is reused by other communications besides remote replication communication. If it is determined that the port associated with the remote replication link is reused by other communications besides remote replication communication, process the reused port. Determine whether the performance of the communication components associated with the remote replication link is abnormal. If it is determined that the performance of the communication components associated with the remote replication link is abnormal, process the communication components with abnormal performance. Determine whether the CPU utilization rate of the storage device associated with the remote replication link is abnormal. If it is determined that the CPU utilization rate of the storage device associated with the remote replication link is abnormal, process the storage device with abnormal CPU utilization rate. Determine whether the write speed of the auxiliary system volume associated with the remote replication link is abnormal. If it is determined that the write speed of the auxiliary system volume associated with the remote replication link is abnormal, process the auxiliary system volume with abnormal write speed. Determine whether there is an anomaly in the bandwidth between the master system and the auxiliary system associated with the remote replication link. If it is determined that there is an anomaly in the bandwidth between the master system and the auxiliary system associated with the remote replication link, process the master system and the auxiliary system where the bandwidth between the systems is abnormal.

6. The method according to claim 1, characterized in that, The method further includes: If at least one partnership in the storage system meets the set conditions, a third test result is output, which is used to characterize that the connection of the remote cluster configured by the partnership that meets the set conditions has been lost. The setting conditions include at least one of the following: The communication layer of the storage system does not contain the partnership link associated with the partnership. The remote replication module associated with the partnership in the storage system cannot communicate bidirectionally.

7. The method according to claim 6, characterized in that, After outputting the third test result, the method further includes: For partnerships that meet the defined criteria, perform at least one of the following operations: Determine whether the remote replication link associated with the partnership is unavailable. If the remote replication link associated with the partnership is unavailable, process the unavailable remote replication link. Determine whether there is a remote replication relationship persistence error in the data stream associated with the partnership. If it is determined that there is a remote replication relationship persistence error in the data stream associated with the partnership, process the data stream with the remote replication relationship persistence error.

8. The method according to any one of claims 1-7, characterized in that, After outputting the first test result, the method further includes: For the data stream with a first duration not less than a second set value, perform at least one of the following operations: Based on the host latency of the data stream, determine whether the setting of the first set value is reasonable, and if it is determined that the setting of the first set value is unreasonable, adjust the first set value according to the maximum value of the host latency of the data stream; Determine whether the write speed of the storage system volume associated with the data stream is abnormal. If it is determined that the write speed of the storage system volume associated with the data stream is abnormal, process the storage system volume with abnormal write speed. Determine whether the CPU utilization rate of the storage device associated with the data stream is abnormal. If it is determined that the CPU utilization rate of the storage device associated with the data stream is abnormal, process the storage device with the abnormal CPU utilization rate. Determine whether there is an abnormality in the latency of the communication port between the storage cluster nodes associated with the data stream. If it is determined that there is an abnormality in the latency of the communication port between the storage cluster nodes associated with the data stream, process the communication port between the storage cluster nodes with abnormal latency. Determine whether there is an anomaly in the bandwidth between the main system and the auxiliary system associated with the data stream. If it is determined that there is an anomaly in the bandwidth between the main system and the auxiliary system associated with the data stream, process the main system and the auxiliary system where the bandwidth between the systems is abnormal.

9. A testing device for communication reliability, characterized in that, The device includes: The first partitioning module is used to divide the data transmitted in the storage system through various remote replication relationships into multiple data streams; The first statistics module is used to, for each data stream, calculate the first duration during which the host latency of the data stream exceeds a first preset value; The first output module is configured to output a first test result when at least one of the data streams has a first duration of not less than a second set value. The first test result is used to characterize that the data stream with a first duration of not less than the second set value has a remote replication relationship persistence error. The first statistical module includes: The first statistics submodule is used to calculate the second duration for which the host latency of the data stream exceeds a first set value in each of the first periods, based on the set duration of the first period. The second statistics submodule is used to determine each of the first cycles with a second duration not less than a third set value as the second cycle, and to determine each of the first cycles with a second duration less than the third set value as the third cycle, wherein the third set value is the duration of a set percentage of a single first cycle. The third statistics submodule is used to determine the difference between the total duration of the second period and the total duration of the third period as the first duration during which the host latency of the data stream exceeds the first set value. The first output module includes: The first output submodule is configured to, when the first duration of at least one of the data streams is not less than a second set value, select one remote replication relationship from the associated remote replication relationships to stop, based on the busy level of each remote replication relationship associated with the target data stream, and output the first test result. The target data stream is the data stream with a first duration not less than a second set value.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the communication reliability testing method as described in any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the communication reliability testing method as described in any one of claims 1 to 8.