Network state detection method, system and device and storage medium

By acquiring network topology information of the storage cluster and dynamically adjusting the number of probe packets, combined with detection models for different network types, the problem of insufficient accuracy in network sub-health detection in existing technologies has been solved, enabling timely and accurate identification and detection of network status.

CN122069210APending Publication Date: 2026-05-19DAWNING INFORMATION IND (BEIJING) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DAWNING INFORMATION IND (BEIJING) CO LTD
Filing Date
2026-03-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing network status detection methods cannot accurately detect network sub-health conditions, leading to system performance degradation and potentially evolving into failures. Furthermore, existing methods can affect system operation under high load or make misjudgments under low load.

Method used

By acquiring network topology information between controllers in the storage cluster, the number of active probe messages sent is dynamically adjusted, and a state detection model for different network types is constructed. The detection is then performed by combining the active probe response results and the business network performance results.

Benefits of technology

It enables accurate detection of network sub-health conditions, reduces false alarms, ensures high availability and high performance of the storage system, and avoids interference with business operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069210A_ABST
    Figure CN122069210A_ABST
Patent Text Reader

Abstract

The invention discloses a network state detection method, system and device and a storage medium, and the method comprises the steps: obtaining network topology information between controllers in a storage cluster, the network topology information comprising physical network cards associated with network connection relationships of different network types; determining the service cycle load capacity of each physical network card, and dynamically determining the active detection message sending number of each physical network card according to each service cycle load capacity; determining a target network address based on the network topology information, and initiating an active detection request to the target network address; obtaining an active detection response result responding to the active detection request, and obtaining a service network performance result in an active detection request period; constructing each state detection model corresponding to different network types; and inputting the active detection response result and the service network performance result into a corresponding state detection model to obtain a network state detection result. According to the method provided by the invention, the detection accuracy of the network state is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network status detection technology, and in particular to a network status detection method, system, device, and storage medium. Background Technology

[0002] Centralized all-flash storage systems typically consist of multiple controllers, intelligent hard disk enclosures, and other components. In some scenarios, dual storage arrays are required to achieve replication and active-active functionality. The network, as the core link connecting the independent components and integrating the system to provide services to the outside world, directly determines the performance and stability of the storage system.

[0003] However, existing network status detection methods determine whether a network is normal or abnormal based on current indicators such as packet loss and latency. These methods cannot achieve more refined status detection, such as detecting sub-healthy network conditions. Sub-healthy networks differ from completely unusable network failures; they manifest as inefficient operating states such as packet loss, abnormal latency, and congestion. This not only makes some components performance bottlenecks, leading to decreased system read / write performance, but may also evolve into a complete network failure after prolonged neglect, causing storage system paralysis.

[0004] Current methods for detecting network sub-health rely heavily on business network performance metrics. For example, they identify sub-health by determining whether the network latency of communication and heartbeat messages between two OSDs (Object Storage Devices) exceeds a preset threshold. However, detection methods based solely on business network performance data are lagging, only becoming apparent after sub-health affects upper-layer applications. Furthermore, insufficient data volume during periods of low workload can easily lead to misjudgments. While sending a large number of proactive probe messages can detect anomalies in advance, it increases the burden on physical network cards, impacting normal system operation under high workloads. Summary of the Invention

[0005] This invention provides a network status detection method to address the technical problem of how to improve existing network status detection methods and achieve the effect of improving the accuracy of network status detection.

[0006] To address the aforementioned technical problems, this invention provides a network state detection method, including... Obtain network topology information between controllers in the storage cluster, wherein the network topology information includes each physical network interface card associated with network connection relationships of different network types; Determine the service cycle load of each physical network interface card (NIC), and dynamically determine the number of active probe packets sent by each physical NIC based on the service cycle load. Based on the network topology information, the target network address is determined, and an active probe request is initiated to the target network address. The number of active probe requests is determined by the number of corresponding active probe packets sent. Obtain the active probe response result in response to the active probe request, and obtain the service network performance results within the active probe request period; Construct state detection models corresponding to different network types; The active detection response result and the service network performance result are input into the corresponding state detection model to obtain the network state detection result of the storage cluster.

[0007] As one preferred embodiment, the network type includes an inter-controller network, a replica active-active network, and a back-end network; The process of obtaining the network topology information of the storage cluster includes: Based on the preset configuration information of the storage cluster, the IP address connection relationship between each node in the cluster is obtained, and the corresponding physical network card identifier is determined through the IP address connection relationship, so as to obtain the first connection relationship of the control network and its associated first physical network card. Obtain the preset logical interface of the storage cluster, determine the corresponding physical network card identifier based on the port group information bound to the preset logical interface, and obtain the second connection relationship and associated second physical network card of the replication dual-active network; Based on the preset configuration information of the storage cluster, the IP address connection relationship between the nodes in the cluster and the back-end storage device is obtained. The corresponding physical network card identifier is determined through the IP address connection relationship, and the third connection relationship of the back-end network and the associated third physical network card are obtained. The network topology information of the storage cluster is generated by integrating the first connection relationship, the first physical network card, the second connection relationship, the second physical network card, the third connection relationship, and the third physical network card.

[0008] As one preferred embodiment, determining the service cycle load of each physical network interface card (NIC) and dynamically determining the number of active probe packets sent for each physical NIC based on the service cycle load includes: A reference value for the single-cycle service load is determined based on historical load data, and the load of the service cycle is compared with the reference value. If the service cycle load is lower than the reference value, the number of active probe messages sent is generated based on the difference between the service cycle load and the reference value. If the load of the service cycle is not lower than the reference value, no additional number of active probe messages will be generated.

[0009] As one preferred embodiment, the step of determining the target network address based on the network topology information and initiating an active probe request to the target network address includes: Using the network to which the physical network card currently belongs as the source address, target network addresses are filtered based on the connection relationships in the network topology information, wherein the target network addresses cover all associated nodes under the network type corresponding to the physical network card; Based on the number of active probe messages sent, probe requests are initiated to each of the target network addresses, and the number of responses and response times of the probe requests are obtained; Based on the number of responses and the response time, the active detection response result is calculated.

[0010] As one preferred embodiment, the construction of various state detection models corresponding to different network types includes: All physical network cards corresponding to the same network type are divided into the same independent detection domain, wherein the network performance evaluation benchmark of the independent detection domain is preset according to the characteristics of the corresponding network type; Construct an initial network state detection model corresponding to each independent detection domain, and obtain historical service network performance results and historical active probe response results in the storage cluster to generate a training dataset; The initial network state detection model is trained based on the training dataset to obtain the network state detection model for each independent detection domain.

[0011] As a preferred embodiment, the step of inputting the active detection response result and the service network performance result into the corresponding state detection model to obtain the network state detection result of the storage cluster includes: The service network performance results and active detection response results of each physical network card within the same independent detection domain are input into the network status detection model corresponding to the detection domain, and the candidate sub-health markers of each physical network card in the current time period are output. Record the time period during which each physical network card is marked as a candidate sub-healthy, and take the time period of the first time it is marked as the candidate sub-healthy as the statistical starting point, and accumulate the number of consecutive periods in which it is marked as the candidate sub-healthy. If the number of consecutive cycles reaches the preset cycle determination threshold, the physical network card is determined to be in a sub-healthy state; if any cycle during the accumulation process is not marked as a candidate sub-healthy state, the consecutive cycle count is reset. The network status detection results of the storage cluster are generated by integrating the judgment results of all the physical network cards.

[0012] As one preferred embodiment, after generating the network status detection results of the storage cluster, the method further includes: Based on the network status detection results, identify the target physical network interface card (NIC) that is in a sub-healthy network state. Generate alarm information containing the identification information, network type, and sub-health characteristic parameters of the target physical network card, and send the alarm information to a preset receiving end; The network redundancy configuration of the storage cluster is detected, and it is determined whether there are redundant links or redundant network cards that can replace the target physical network card. If the redundant link or redundant network card exists and the preset isolation conditions are met, then an isolation operation is performed on the target physical network card.

[0013] Another aspect of the present invention provides a network status detection system, comprising: The acquisition module is used to acquire network topology information between controllers in the storage cluster. The network topology information includes each physical network card associated with the network connection relationship of different network types. An evaluation module is used to determine the service cycle load of each physical network card within a preset time period, and dynamically determine the number of active probe packets sent by each physical network card based on the service cycle load. The detection module is used to determine the target network address based on the network topology information and initiate an active detection request to the target network address. The number of active detection requests is determined by the number of corresponding active detection packets sent. The collection module is used to obtain the active probe response results in response to the active probe request, and to obtain the service network performance results within the active probe request period; The building module is used to construct various state detection models corresponding to different network types. The detection module is used to input the active detection response result and the service network performance result into the corresponding state detection model to obtain the network state detection result of the storage cluster.

[0014] In another aspect, the present invention provides a network status detection device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the network status detection method as described above.

[0015] In another aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein when the device containing the computer-readable storage medium executes the computer program, it implements the network status detection method described above.

[0016] Compared with the prior art, the beneficial effects of the present invention are at least one of the following: 1) This invention accurately acquires the network topology information of the storage cluster, clarifies the correlation between different network types and physical network cards, and provides comprehensive and accurate basic data support for subsequent detection. At the same time, it dynamically adjusts the number of active probe packets sent based on the service cycle load of the physical network card. When the service load is low, it supplements probe packets to ensure the sufficiency of performance data and avoids misjudgment in low-load scenarios. When the service load is high, it reduces the sending of additional probe packets to avoid interference with storage services. This achieves dynamic adaptation between probe behavior and service load, which not only ensures the timeliness of detection but also takes into account the operational stability of the storage system.

[0017] 2) This invention effectively solves the problem of insufficient accuracy caused by the use of the same detection standard for different network types in traditional methods by constructing dedicated state detection models for different network types and integrating active detection response results with business network performance results. The performance characteristics and requirements of different network types are fully considered, making the model detection more targeted. The fusion of dual-dimensional performance data makes the network state assessment more comprehensive, significantly improving the accuracy of network sub-health detection. It can timely and accurately identify sub-health states such as packet loss and latency anomalies, providing a reliable basis for subsequent alarm and isolation operations, and further ensuring the high availability and high performance of the centralized all-flash storage system. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a network state detection method in one embodiment of the present invention; Figure 2 This is a schematic diagram of the network structure of a centralized all-flash storage system in one embodiment of the present invention; Figure 3 This is an overall flowchart of a network state detection method in one embodiment of the present invention; Figure 4 This is a structural block diagram of a network state detection system according to one embodiment of the present invention; Figure 5 This is a schematic diagram of a network status detection device in one embodiment of the present invention; Figure label: Among them, 11 is the acquisition module; 12 is the evaluation module; 13 is the detection module; 14 is the collection module; 15 is the construction module; 16 is the detection module; 21 is the processor; and 22 is the memory. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0020] In the description of this invention, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0021] In the description of this invention, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0022] One embodiment of the present invention provides a network status detection method; for details, please refer to [link to specific documentation]. Figure 1 , Figure 1 The diagram shown is a flowchart of a network state detection method according to one embodiment of the present invention, which includes steps S1-S6: S1: Obtain network topology information between controllers in the storage cluster. The network topology information includes the physical network cards associated with the network connection relationships of different network types.

[0023] In centralized all-flash storage systems, the network serves as a crucial link connecting core components such as controllers and intelligent hard drive enclosures. Accurate acquisition of its topology is fundamental for subsequent network health monitoring. Different network types (inter-controller networks, replica active-active networks, and back-end networks) undertake different data transmission functions, and the corresponding physical network interface cards (NICs) have significant differences in performance requirements and connection logic. Therefore, it is necessary to collect connection relationships and NIC information in a targeted manner to ensure the completeness and accuracy of the topology information.

[0024] See Figure 2 , Figure 2This is a schematic diagram of a centralized all-flash storage network structure provided by an embodiment of the present invention. The diagram clearly shows the hardware component layout and network connection relationship of the storage cluster: it includes two control boxes, control box 1 and control box 2, and each control box deploys two controllers (control box 1 contains controller 1 and controller 2, and control box 2 contains controller 3 and controller 4). The back-end of the cluster is configured with two smart hard disk enclosures.

[0025] Each controller connects to three types of networks via different types of physical network interface cards (NICs). The inter-controller network connects controllers in different control frames within the same cluster (e.g., controller 1 connects to controller 3, and controller 2 connects to controller 4 via NICs 1 and 2). The replication-active network enables data transfer across storage arrays (e.g., controllers 1-4 communicate with a remote array controller via NIC n). The backend network connects controllers to smart hard disk enclosures (e.g., controllers 1 and 2 connect to smart hard disk enclosure 1 via NIC 3, and controllers 3 and 4 connect to smart hard disk enclosure 2 via NIC 3). Figure 2 It is known that physical network cards of different network types have clear distinctions in terms of connection objects and transmission scenarios, which is also the core basis for this invention to collect network connection information in a targeted manner.

[0026] This embodiment takes the above-mentioned centralized all-flash storage cluster containing 4 controllers and 2 smart hard disk enclosures as an example to explain in detail the process of obtaining network topology information.

[0027] In this embodiment, the preset configuration information of the storage cluster is stored in the configuration file of the main controller, including basic data such as the IP address allocation of each node, the binding relationship between ports and network cards, and the access parameters of the backend storage devices. This data can be directly accessed through the system's built-in configuration read interface. The preset logical interface (LIF) is configured by the user through the storage system management platform and is used for cross-array data transmission in active-active replication scenarios. It should be noted that this invention does not limit the storage format of the configuration information (such as XML or JSON) or the specific implementation method of the query interface, as long as the IP address connection relationship and the port group and network card binding relationship can be obtained.

[0028] For obtaining the first connection relationship and the first physical network card in the inter-controller network: Based on the preset configuration information of the storage cluster, the IP addresses of each controller are first extracted (e.g., controller 1's IP is 192.168.1.1, controller 2's is 192.168.1.2, controller 3's is 192.168.1.3, and controller 4's is 192.168.1.4). The direct connection relationship between each controller is obtained through IP routing table resolution (e.g., controller 1 is directly connected to controllers 2, 3, and 4, and controller 2 is directly connected to controllers 3 and 4, etc.), which is the first connection relationship; then, the corresponding physical network card identifier is looked up based on the above IP addresses using the network card management tool.

[0029] For example, the physical network interface card (NIC) names corresponding to IP address 192.168.1.1 are eth0 and eth1, and the physical NIC names corresponding to IP address 192.168.1.2 are eth0 and eth2. These physical NICs associated with the inter-controller network connection are the first physical NICs.

[0030] For replicating the second connection relationship of the active-active network and obtaining the second physical network interface card: First, obtain the three LIFs configured by the user (named LIF1, LIF2, and LIF3 respectively). Among them, LIF1 is bound to the port group PortGroup-A, LIF2 is bound to the port group PortGroup-B, and LIF3 is bound to the port group PortGroup-C. It should be noted that as a virtual interface for communication between the storage system and the host, the port group bound to the LIF consists of multiple physical ports, and each physical port corresponds to a unique physical network interface card.

[0031] By querying the mapping table between port groups and physical network cards, it is determined that PortGroup-A corresponds to eth3 of controller 1 and eth3 of controller 2, PortGroup-B corresponds to eth3 of controller 3 and eth3 of controller 4, and PortGroup-C corresponds to eth4 of controller 1 and eth4 of controller 2. Then, according to the communication logic of the port groups, the connection relationship of the replicated active-active network is obtained (such as controller 1 establishing a connection with the remote array controller through eth3, controller 3 establishing a connection with the remote array controller through eth3, etc.), which is the second connection relationship. The corresponding eth3 and eth4 are the second physical network cards.

[0032] For obtaining the third connection relationship and the third physical network interface card (NIC) of the backend network: Based on the preset configuration information of the storage cluster, extract the IP addresses of each controller and the backend smart disk enclosure (e.g., the IP of smart disk enclosure 1 is 192.168.2.1, and the IP of smart disk enclosure 2 is 192.168.2.2). The connection relationship between each controller and the smart disk enclosure is obtained by parsing the device access log (e.g., controllers 1 and 2 are directly connected to smart disk enclosure 1, and controllers 3 and 4 are directly connected to smart disk enclosure 2), which is the third connection relationship. Then, by reverse lookup through the binding relationship between IP address and NIC, it is determined that eth5 and eth6 of controller 1 are connected to smart disk enclosure 1, and eth5 and eth6 of controller 3 are connected to smart disk enclosure 2. These physical NICs used for backend data transmission are the third physical NICs.

[0033] By integrating the first connection relationship, first physical network card, second connection relationship, second physical network card, third connection relationship, and third physical network card obtained above, network topology information of the storage cluster is generated. This topology information is stored in the form of structured data, including network type field, IP fields of both connected parties, physical network card identifier field, port mapping field, etc., which supports subsequent active detection of target address filtering and detection domain division.

[0034] Preferably, the topology information is synchronized to the local cache of each controller and smart hard disk enclosure in real time. When the network configuration of the storage cluster changes (such as adding a controller or adjusting the LIF binding relationship), the topology information is triggered to be reacquired and updated, ensuring that the topology data used in the detection process is consistent with the actual network status.

[0035] S2: Determine the service cycle load of each physical network card, and dynamically determine the number of active probe packets to be sent for each physical network card based on the service cycle load.

[0036] In network sub-health detection of centralized all-flash storage systems, the number of active probe packets sent directly affects detection accuracy and service stability. Existing technologies using a fixed number of probe packets are prone to misjudgments due to insufficient data volume under low service load, while under high service load, they consume additional network bandwidth, impacting service transmission. Therefore, this step dynamically adapts the number of active probe packets sent by comprehensively evaluating the service cycle load of the physical network interface cards (NICs). This ensures sufficient detection data under low load scenarios while avoiding service interference under high load scenarios. This embodiment is still based on a storage cluster containing four controllers and two intelligent hard disk enclosures for detailed explanation.

[0037] In this embodiment, the storage cluster presets a fixed detection time period for each physical network card, preferably 10 seconds. This period can be flexibly adjusted according to service response requirements, and is not limited herein. The reference value of single-cycle service load (denoted as T) is determined based on historical load data.

[0038] Specifically, single-cycle load data for the physical network interface card (NIC) is collected within the same time window (e.g., 9:00-10:00 on weekdays and 23:00-24:00 at night) over the past 30 days. A reference value T is calculated using the arithmetic mean method. If the physical NIC has been in use for less than 30 days, the average load of other physical NICs in the same network during the same period is used as the initial reference value, and subsequently updated dynamically over time. It should be noted that historical load data is stored in the local database of each controller and can be accessed in real-time via a data query interface. This invention does not restrict the data storage format or the implementation method of the query interface.

[0039] In this embodiment, at fixed time intervals, each controller evaluates the service load of each physical network interface card (NIC) and dynamically adjusts the number of probe packets actively sent by each NIC based on the service load. The formula for calculating the service load of each physical NIC is as follows: in, Indicates the first Periodic physical network interface card The load capacity; and Indicates up to the up to the end of the cycle, physical network card The total number of packets sent and received since the controller was powered on; and Indicates up to the up to the end of the cycle, physical network card The total number of packets sent and received since the controller was powered on; and This is a service load adjustment factor that indicates the number of packets received and sent within the period, preset based on network interface card (NIC) load experience.

[0040] After obtaining the business cycle load, it is compared with a reference value to dynamically determine the number of active probe messages to be sent. Periodic physical network interface card The formula for calculating the number of probe messages that need to be actively sent is: in, This is a pre-set threshold based on experience, representing the total number of packets sent by each physical network interface card (NIC) within each time period. This is relevant when the service load... If the amount is below the threshold, a replacement is required. Each performance probe message avoids inaccurate detection of network sub-health due to insufficient statistical performance data; when the service load... When the load exceeds the threshold, the service load is high, and there is no need to resend performance probe messages, thus preventing the services on the physical network card from being affected.

[0041] Preferably, to avoid a sudden increase or decrease in the number of probe packets due to load fluctuations in extreme cases, this embodiment also sets upper and lower limit thresholds for the number of active probe packets. The upper limit threshold is set to 1.5 times the reference value T, and the lower limit threshold is set to 0, meaning the range of N is 0 ≤ N ≤ 1.5T. For example, when T = 1000, N cannot exceed 1500. If the difference between the value and T exceeds 1500, probe packets will still be sent as if the value were 1500 to prevent excessive probe packets from consuming too much network bandwidth; when When the difference between N and T is negative, it is treated as N=0. In addition, each controller will record the L, T, N and other parameters of each physical network card in real time and synchronize them to the main controller for subsequent performance data analysis and model optimization, ensuring the continuous adaptability of the dynamic adjustment strategy.

[0042] S3: Determine the target network address based on network topology information, and initiate an active probe request to the target network address. The number of active probe requests is determined by the number of corresponding active probe packets sent.

[0043] In network sub-health detection of centralized all-flash storage systems, the accurate initiation of proactive probe requests is crucial for obtaining effective network performance data. Existing technologies often suffer from incomplete target address filtering, leading to biased detection results, and fixed packet allocation methods can cause uneven load distribution among nodes. This step, however, filters target addresses based on complete network topology information and initiates requests using a dynamically determined number of probe packets. This ensures comprehensive probe coverage while maintaining the authenticity of performance data through reasonable packet allocation, providing reliable data support for subsequent sub-health detection. This embodiment is still based on a storage cluster containing four controllers (Controllers 1-4) and two smart hard disk enclosures (Smart Hard Disk Enclosures 1-2). The detection time period remains consistent with Embodiment 2, at 10 seconds.

[0044] In this embodiment, the active probe message is encapsulated using the ICMP protocol. ICMP (Internet Control Message Protocol), as a core protocol for network diagnostics, can efficiently transmit control messages between hosts and routers. Its message structure is simple, transmission overhead is low, and it can quickly obtain network connectivity and latency data, meeting the lightweight requirements of active probes. It should be noted that this invention does not impose a unique limitation on the protocol type of the probe message; other protocols with network diagnostic functions, such as TCP and UDP, are also within the scope of protection of this invention.

[0045] The selection of target network addresses should use the network to which the physical network interface card (NIC) currently belongs as the source address. Based on the connection relationships in the network topology information, it should ensure coverage of all associated nodes under the corresponding network type: For the first physical NIC in the inter-controller network (such as eth0 and eth1 of controller 1), its network is the inter-controller network, and its associated nodes are all other controllers in the cluster. Therefore, the IP addresses of controllers 2-4 (192.168.1.2, 192.168.1.3, and 192.168.1.4) are extracted from the topology information as the target network addresses; for replication dual-active... The second physical network interface card (such as eth3 and eth4 of controller 1) belongs to a replicated active-active network and is associated with the controller of the remote storage array. The IP address of the remote array controller (192.168.3.1, 192.168.3.2) is extracted as the target network address. For the third physical network interface card (such as eth5 and eth6 of controller 1) in the backend network, it belongs to the backend network and is associated with the corresponding smart hard disk enclosure. The IP address of smart hard disk enclosure 1 (192.168.2.1) is extracted as the target network address.

[0046] Preferably, the target network address will be synchronized in real time with the update of network topology information. If a new associated node is added or the node IP changes, the target address will be automatically re-filtered to ensure that the detection coverage is complete.

[0047] After determining the target network address, probe requests are initiated to each target address based on the number of active probe packets (N) dynamically calculated in Embodiment 2. In this embodiment, probe packets are allocated according to the "uniform distribution" principle. If N=655 for a certain physical network card, corresponding to 3 target network addresses, then 218 packets are allocated to each target address, and the remaining packet is randomly allocated to any target address to ensure that the amount of probe data for each associated node is balanced. If N=0, no additional probe requests are initiated, and only service network performance data is required.

[0048] It should be noted that this invention does not limit the message allocation method. In addition to uniform allocation, the number of messages can also be allocated according to the importance weight of the target node. As long as the validity of the probe data can be guaranteed, it is within the protection scope of this invention. When a probe request is initiated, each controller will record the sending timestamp of each probe message (accurate to the microsecond level) and listen for the response message from the target address.

[0049] After the detection period ends, the number of responses to the probe request and the response time are counted: the number of responses is the total number of valid response packets returned by each target address. If no response is received after the probe packet is sent for more than the preset timeout period (preferably set to 500 milliseconds in this embodiment, but can be adjusted according to the characteristics of the network type), it is judged as a timeout and is not included in the number of responses; the response time is the difference between the receiving timestamp and the sending timestamp of each valid response packet, that is, the round-trip time (RTT) of a single packet. As a core indicator of network performance, RTT directly reflects the round-trip time of data transmission and is a key parameter for judging whether the network is in a sub-healthy state.

[0050] The active probing response results are calculated based on the number of responses and response time, with core metrics including packet loss rate and average round-trip time: The formula for calculating packet loss rate (Loss) is as follows: Where R is the number of responses, for example, when N=655 and R=640, Loss=(655-640) / 655×100%≈2.29%; Average round-trip time ( The formula for calculating ) is: in, Let be the round-trip time (RTT) of the i-th valid response message. For example, if the total RTT of 640 response messages returned from three destination addresses is 3200 milliseconds, then... =3200 / 640=5 milliseconds.

[0051] Preferably, before calculating the average round-trip time (RTT), outliers exceeding the normal range (such as RTT data exceeding the timeout period) are removed to avoid extreme data affecting the accuracy of performance evaluation. Each controller will report the calculated packet loss rate, average RTT, and other proactive detection response results, along with concurrent service network performance data (such as service packet loss rate and RTT latency), to the storage cluster master controller for subsequent sub-health model prediction.

[0052] S4: Obtain the active probe response result in response to the active probe request, and obtain the service network performance results within the active probe request period.

[0053] In network sub-health detection, the comprehensiveness and accuracy of performance data directly determine the reliability of the detection results. Existing technologies often rely solely on proactive probe data or business data, resulting in a one-sided detection dimension. This step achieves complementary data from two dimensions by simultaneously acquiring proactive probe response results and business network performance results. This covers both the forward-looking nature of early detection and the authenticity of business operations, providing comprehensive input data for subsequent state detection models. This embodiment is still based on a storage cluster containing 4 controllers and 2 intelligent hard disk enclosures, and the detection time period remains at 10 seconds, consistent with the previous embodiment.

[0054] In this embodiment, the core indicators of the active probing response results are packet loss rate and average round-trip time (RTT). These two indicators are key parameters reflecting network connectivity and transmission efficiency, and can directly characterize whether the network exhibits sub-health characteristics such as packet loss or abnormal latency. It should be noted that this invention does not limit the types of indicators for the active probing response results to a single type. In addition to the core indicators mentioned above, auxiliary indicators such as jitter rate and packet retransmission rate can be added according to actual detection needs, and all of these fall within the scope of protection of this invention.

[0055] The acquisition of the active probe response results is based on the probe request initiated in step S3. The specific implementation process is as follows: Each controller captures the response packets returned by the target network address in real time and records the receiving timestamp of each response packet (with the same precision as the sending timestamp in step S3, both at the microsecond level).

[0056] After the 10-second detection period ends, the number of valid responses (R) is first counted. If no response is received after a preset timeout period (preferably 500 milliseconds in this embodiment, but can be dynamically adjusted according to the network type; for example, the timeout period for a dual-active network can be set to 1000 milliseconds) following the sending of a probe message, it is determined to be an invalid message and is not included in R. Then, the packet loss rate (Loss) is calculated using the following formula: Where N is the number of active probe messages sent as determined in step S2.

[0057] Finally, the average RTT delay was calculated. The sum of the differences between the received and sent timestamps of each valid response message, divided by the number of valid responses, yields the result. .

[0058] The service network performance results are obtained from the performance parameters of each physical network interface card (NIC) during service operation, ensuring complete synchronization of the data and the time period of the active probe data (both are 10 seconds). In this embodiment, the core indicators of the service network performance results include service packet loss rate, service average RTT latency, and service throughput. The service packet loss rate is calculated by the difference between the total number of packets sent and received in the current period, using the following formula: For example, if a physical network card currently sends 5000 packets and receives 4980 packets in the current cycle, then... =(5000-4980) / 5000×100%=0.4%; The average RTT latency of the service is calculated by extracting the timestamp field from the service packets (provided by the service layer protocol, such as the timestamp field of the iSCSI protocol) and calculating the average round-trip latency of all service packets; the service throughput is obtained by dividing the total number of bytes of service packets in the current period by the time period, in Mbps. It should be noted that the implementation of service data statistics can be based on the network card driver statistics function of the existing storage system or the analysis of upper-layer service logs. This invention does not limit this, as long as the above service performance indicators can be accurately collected.

[0059] Preferably, to ensure data accuracy, this embodiment also includes a data verification mechanism: for active probe response results, if the packet loss rate exceeds 80% or the average RTT latency exceeds twice the timeout period, the probe data is deemed abnormal, the active probe results for that period are marked as invalid, and a supplementary probe mechanism for the next period is triggered; for service network performance results, if the total number of service packets counted is lower than a preset minimum (set to 100 in this embodiment), only valid data is retained, and features are supplemented through data augmentation algorithms during subsequent model training to avoid data bias caused by low service volume. Furthermore, each controller stores the acquired active probe response results and service network performance results in a key-value pair format in its local database, with the key being "physical network card identifier - time period," facilitating subsequent targeted queries and calls by the main controller.

[0060] Within 3 seconds of the end of each detection cycle, each controller reports the two types of performance data to the storage cluster master controller via its internal communication channel. Upon receiving the data, the master controller verifies the data format (e.g., completeness of indicator fields and reasonableness of numerical ranges). If the verification passes, the data is stored according to physical network card identifier and network type, forming a unified performance dataset to prepare for subsequent detection domain division and state detection model input. If data verification fails, the master controller sends a retransmission request to the corresponding controller to ensure no data is missed. This invention does not limit the communication protocol used for data reporting; protocols such as TCP and UDP can be used.

[0061] S5: Construct state detection models corresponding to different network types.

[0062] In network sub-health detection, the adaptability of the detection model directly determines the accuracy of the detection results. Existing technologies include all network interface cards (NICs) of all network types in the same model for detection, ignoring the fundamental differences in performance requirements and transmission characteristics between control networks, replicated active-active networks, and backend networks. Furthermore, relying on manually preset thresholds easily leads to insufficient adaptability. This step divides the detection domains according to network type, constructs and trains a dedicated network status detection model for each detection domain, enabling the model to learn the performance characteristics of the corresponding network type, and completely solves the detection bias problems caused by mixed detection of different network types and manually set thresholds. This embodiment is still based on a storage cluster containing 4 controllers and 2 intelligent hard disk enclosures, and uses the 10-second detection cycle of the previous embodiment.

[0063] In this embodiment, the division of independent detection domains is based solely on network type, and the network performance evaluation benchmark for each detection domain is preset according to the characteristics of the corresponding network type, ensuring that the evaluation criteria match the actual application scenario of the network.

[0064] Specifically, all physical network interface cards (NICs) used for inter-controller network communication (such as eth0 and eth1 of controllers 1-4) are classified as the first detection domain. As the core network for data interaction between controllers within the cluster, the inter-controller network has stringent transmission latency requirements (usually in the microsecond range). Therefore, the preset evaluation benchmark is "average RTT latency ≤ 50 microseconds, packet loss rate ≤ 1%". All physical NICs used for replication dual-active network communication (such as eth3 and eth4 of controllers 1-4) are classified as the second detection domain. Replication dual-active networks involve cross-array data transmission, and distance factors result in higher latency (usually in the millisecond range). The preset evaluation benchmark is "average RTT latency ≤ 500 milliseconds, packet loss rate ≤ 3%". All physical NICs used for backend network communication (such as eth5 and eth6 of controllers 1-4) are classified as the third detection domain. The backend network is responsible for data reading and writing between the controller and the smart hard disk enclosure. Its performance requirements are between the first two. The preset evaluation benchmark is "average RTT latency ≤ 200 microseconds, packet loss rate ≤ 2%".

[0065] It should be noted that the specific values ​​of the above evaluation benchmarks can be flexibly adjusted according to the hardware configuration and business requirements of the storage system, and this invention does not limit them.

[0066] The initial network state detection model is constructed using a BP neural network. As a multi-layer feedforward neural network, the BP neural network iteratively optimizes parameters through the backpropagation algorithm, which can accurately learn the mapping relationship between input features and output state and adapt to the nonlinear characteristics of network performance data.

[0067] In this embodiment, the initial BP neural network structure is consistent for each detection domain but the parameters are independent: the number of nodes in the input layer is set to 6, corresponding to 6 core feature parameters (service packet loss rate, service average RTT latency, service throughput, active detection packet loss rate, active detection average RTT latency, and active detection packet response rate), ensuring comprehensive coverage of the key performance dimensions of service and active detection; the hidden layer is set to 2 layers, with 12 nodes in the first layer and 8 nodes in the second layer, increasing the number of nodes in the hidden layer to improve the model's ability to fit complex features; the number of nodes in the output layer is set to 1, and the output value is 0 or 1, where 0 indicates that the network is in a normal state and 1 indicates that the network is in a sub-healthy state.

[0068] It should be noted that this invention does not impose a unique limitation on the number of layers and nodes of the BP neural network. The number of layers and nodes can be adjusted according to the complexity of the data in the detection domain. For example, the number of hidden layers can be increased when the data features are more complex, as long as accurate classification of the network state can be achieved.

[0069] The training dataset needs to be generated based on the historical performance data of the storage cluster, covering complete samples of normal and sub-health states. Specifically, the historical service network performance results and historical active detection response results of the physical network cards of each detection domain are obtained in the past 90 days. The normal state data (positive samples) are the natural operating data when the network is not disturbed, while the sub-health state data (negative samples) are generated by artificially constructing fault scenarios. Specifically, typical sub-health states are created by adjusting network configurations, such as creating a "congestion" scenario by limiting the bandwidth of physical network cards, creating a "packet loss" scenario by setting firewall packet loss rules, and creating a "latency anomaly" scenario by adjusting routing and forwarding policies. Performance data for at least 100 detection cycles are continuously collected under each sub-health scenario to ensure the diversity of negative samples. The collected raw data was then preprocessed: outliers exceeding a reasonable range (such as data where RTT latency suddenly increased to the second level due to hardware failure) were removed; the Min-Max normalization method was used to map all feature parameters to the [0,1] interval to eliminate the impact of dimensional differences on model training; finally, the preprocessed data was divided into training and validation sets in an 8:2 ratio, with the training set used for model parameter learning and the validation set used for evaluating model performance.

[0070] The model training process is based on the BP algorithm, and is mainly divided into two stages: forward propagation and error backpropagation. In the forward propagation stage, the feature data from the training set is input into the initial BP neural network. The output of each layer is calculated using an activation function (preferably the Sigmoid function in this embodiment, but it can be replaced with ReLU functions, etc., depending on the data characteristics; this invention is not limited to this). The final prediction result (0 or 1) is obtained. In the error backpropagation stage, the mean squared error between the prediction result and the actual label (0 for normal, 1 for sub-healthy) is calculated. With the goal of minimizing the error, the weight coefficient matrix and offset of each layer are adjusted backward along the network layers, iteratively updating the model parameters. During training, the accuracy and recall of the model are evaluated using a validation set every 50 iterations. If the accuracy on the validation set is consistently above 95% and the recall is consistently above 92% after 10 consecutive iterations, or if the number of iterations reaches the preset upper limit of 1000, training stops. The model at this point is the final network state detection model for that detection domain. It should be noted that the models for the three detection domains are trained independently and their parameters do not interfere with each other, ensuring that each model only learns the performance features of the corresponding network type and avoids cross-interference.

[0071] Preferably, this embodiment also includes a dynamic model update mechanism: every 30 days, the latest historical performance data (including newly added normal and sub-healthy samples) is automatically acquired, and the models for each detection domain are incrementally trained according to the above training process, updating the weight coefficients and offsets so that the model can adapt to changes in performance characteristics during network operation (such as increased base latency due to hardware aging). Simultaneously, after model deployment, the matching between prediction results and actual network conditions is recorded in real time. If the model prediction accuracy is below 90% for 10 consecutive detection cycles, an emergency update process is triggered, and the model is immediately retrained using recent data to ensure that the model maintains high detection performance over the long term.

[0072] S6: Input the active detection response results and business network performance results into the corresponding state detection model to obtain the network state detection results of the storage cluster.

[0073] In network sub-health detection of centralized all-flash storage systems, single-cycle performance data is easily affected by instantaneous network fluctuations, leading to misjudgments. Relying solely on a single model output lacks fault tolerance and struggles to distinguish between "instantaneous anomalies" and "persistent sub-health." This step employs a two-layer logic of "single-cycle candidate labeling + continuous-cycle cumulative judgment," inputting dual-dimensional performance data into a dedicated model corresponding to the detection domain. Combined with continuous-cycle statistical rules, the final detection result is generated, ensuring both detection sensitivity and avoiding misjudgments caused by instantaneous fluctuations, thus guaranteeing the reliability of network status assessment. This embodiment is still based on a storage cluster containing four controllers and two intelligent hard disk enclosures, using a 10-second detection cycle, maintaining consistency with the detection domain division and model configuration of the aforementioned embodiment.

[0074] In this embodiment, the main controller first performs format standardization processing on the collected active probe response results and service network performance results to ensure that they match the input requirements of the network state detection model for the corresponding detection domain: the six core feature parameters of each physical network card (service packet loss rate, service average RTT latency, service throughput, active probe packet loss rate, active probe average RTT latency, and active probe packet response rate) are converted according to the normalization rules during model training (Min-Max normalization to the [0,1] interval) to eliminate the impact of data dimension differences on the model output. It should be noted that the present invention does not limit the specific method of data standardization. In addition to Min-Max normalization, other methods such as Z-Score normalization can also be used, as long as the format of the input data and the model training data can be kept consistent.

[0075] Subsequently, the main controller inputs the standardized data of each physical network card within the same independent detection domain into the trained BP neural network model corresponding to that detection domain. The model outputs a single-cycle judgment result, i.e., a candidate sub-health label, based on the learned feature mapping relationship. The label value is either "1" or "0": a label value of "1" indicates that the physical network card's two-dimensional performance data in the current cycle meets the sub-health characteristics and is judged as a candidate sub-health network card; a label value of "0" indicates that the performance data in the current cycle is normal and is not marked as a candidate sub-health.

[0076] For example, if the physical network card eth0 of controller 1 in the first detection domain (inter-controller network) has a current period of standardized service packet loss rate of 0.8%, active detection packet loss rate of 1.2%, and average service RTT latency of 45 microseconds, and outputs "1" after inputting it into the model of the first detection domain, then the network card is marked as a candidate for sub-health. If all indicators of physical network card eth1 of controller 2 are lower than the evaluation benchmark in the current period, and the model outputs "0", then it is not marked.

[0077] To distinguish between instantaneous fluctuations and persistent sub-health, this embodiment sets up a continuous periodic cumulative statistical rule: the main controller maintains an independent "continuous candidate sub-health period counter" for each physical network card, taking the time period when it is first marked as a candidate sub-health as the statistical starting point, and incrementing the counter value by 1 for each consecutive period marked as "1"; if the model output of any period is "0" (i.e. not marked as a candidate sub-health), the counter is immediately reset to 0 and the statistics start again.

[0078] In this embodiment, the preset periodicity threshold is set differently based on the network performance requirements of different detection domains: the first detection domain (control room network) has the highest performance requirements, and the preset threshold is 3 consecutive periods; the third detection domain (back-end network) has a preset threshold of 4 consecutive periods; and the second detection domain (replicated active-active network) has relatively lower performance requirements, and the preset threshold is 5 consecutive periods. It should be noted that the above thresholds can be flexibly adjusted according to the stability requirements of the storage system and the tolerance of services to network fluctuations. For example, for financial-grade storage systems with extremely high stability requirements, the threshold for the first detection domain can be adjusted to 2 periods. This invention does not limit this.

[0079] When the number of consecutive candidate sub-health cycles of a physical network interface card (NIC) reaches the preset threshold of the corresponding detection domain, the physical NIC is determined to be in a sub-healthy network state; physical NICs that have not reached the threshold or whose counters have been reset are determined to be in a normal network state.

[0080] For example, if the physical network interface card eth0 of controller 1 in the first detection domain is marked as a candidate for sub-health in the 1st, 2nd, and 3rd consecutive detection cycles (the counter value accumulates to 3), reaching the preset threshold of 3 for the first detection domain, then the network interface card is determined to be in a sub-healthy state. If the network interface card is not marked in the 3rd cycle (outputs "0"), the counter is reset to 0. Even if it is marked again in the 4th and 5th cycles, the count must start from 1 again. In addition, if a physical network interface card switches detection domains due to network configuration changes during the statistical process, the counter will be reset synchronously with the detection domain switch and will start accumulating again according to the preset threshold of the new detection domain, ensuring that the statistical rules match the performance requirements of the detection domain.

[0081] Finally, the main controller integrates the judgment results of all physical network cards to generate the network status detection results of the storage cluster. These results are presented in structured data format, containing core information for each physical network card: physical network card identifier (e.g., "Controller1-eth0"), network type, detection domain, number of consecutive candidate sub-health periods, final judgment status ("normal" or "sub-healthy"), and key abnormal performance indicators (e.g., "active detection packet loss rate consistently 1.2%"). It should be noted that the storage format of the detection results can be tables, JSON, etc.; this invention does not limit this, as long as the status and key information of each network card are clearly presented.

[0082] Preferably, the main controller performs redundancy verification on the detection results: if more than 30% of the physical network cards in the same detection domain are simultaneously determined to be sub-healthy, cross-detection domain correlation analysis will be triggered to check for common problems such as network switch failures and global bandwidth bottlenecks, avoiding the omission of systemic network anomalies; at the same time, the detection results will be synchronized to the local storage of the storage cluster in real time, retaining historical detection records for the most recent 90 days, which will facilitate subsequent network fault tracing and model optimization. This invention does not limit the triggering conditions for redundancy verification or the retention period of historical records, and can be adjusted according to actual operation and maintenance needs.

[0083] This invention further provides a specific embodiment of a network state detection method to illustrate the beneficial effects of the invention. See also: Figure 3 ,in Figure 3 This is an overall flowchart of a network status detection method provided in an embodiment of the present invention.

[0084] This embodiment uses a centralized all-flash storage cluster of a large commercial bank's core business system as an application scenario. The cluster comprises two control frames (each deploying two controllers, for a total of four controllers) and four intelligent disk frames, supporting active-active replication. It needs to handle core business operations such as accounting processing and customer data queries for millions of transactions daily, placing extremely high demands on the storage system's stability, performance, and network anomaly early warning capabilities. The cluster network includes three types: the control network, the active-active replication network, and the backend network, corresponding to physical network cards eth0 / eth1, eth3 / eth4, and eth5 / eth6, respectively. The detection time interval is set to 10 seconds, fully utilizing the technical features of the invention to achieve accurate detection of network sub-health.

[0085] In this embodiment, the network topology information of the storage cluster is first obtained: each controller reads the preset configuration file and parses to obtain the IP address of the controller within the cluster (192.168.1.1-192.168.1.4) and the IP address of the smart disk enclosure (192.168.2.1-192.168.2.4). The controller then reverse-looks up the inter-controller network connection corresponding to eth0 / eth1 (e.g., controller 1 is directly connected to controllers 2-4), and the back-end network connection corresponding to eth5 / eth6 (e.g., controllers 1-2 are directly connected to smart disk enclosures 1-2). The controller then obtains three preset LIFs (LIF1-LIF3) from the replication active-active module. Based on their bound PortGroup-A / B / C, the controller then reverse-looks up the replication active-active network connection corresponding to eth3 / eth4 (communicating with the remote array controller IP 192.168.3.1-192.168.3.2). Finally, complete network topology information is generated, laying the foundation for subsequent detection and domain division.

[0086] Regarding the innovative aspect of dynamically adjusting the number of proactively probed packets for service load, this embodiment is implemented as follows: Based on the historical load data of the cluster over the past 30 days, a single-cycle service load reference value T=1200 (unit: packets / 10 seconds) is calculated, and the service load adjustment factor is... =0.7、 =0.3. Within a certain detection period, controller 1's eth0 (inter-controller network interface card) sends 600 packets and receives 580 packets in the current period, with a cumulative total of 12,000 packets sent and 11,800 packets received. Substituting these values ​​into the service load calculation formula... =(0.7×(600+580)+0.3×(12000+11800)) / (10×2)=(826+6540) / 20=368.3, because <T, the number of resent probe messages N = T - =1200-368.3≈832; while the service load of eth5 (backend network card) of controller 1 in the same period is L=1350 (higher than T), then N=0, and no additional probe packets are sent. This setting solves the problem of misjudgment caused by insufficient data volume in low-load scenarios, and avoids the interference of probe packets to core services in high-load scenarios.

[0087] In the dual-dimensional performance data acquisition phase, this embodiment simultaneously acquires service network performance data and active probe response results: During the aforementioned period, controller 1's eth0 exhibited a service packet loss rate of 0.3%, an average service RTT latency of 42 microseconds, and a service throughput of 120Mbps. After actively sending 832 ICMP probe packets, 815 valid responses were received, with a total response time of 40.75 milliseconds. The calculated active probe packet loss rate was approximately 2.04%, and the average active probe RTT latency was approximately 50 microseconds. Both types of data from all controllers were reported to the main controller within 3 seconds after the end of the period, ensuring comprehensive and real-time data coverage.

[0088] To address the innovation of dividing detection domains and employing dedicated BP neural networks, this embodiment divides eth0 / eth1 into the first detection domain (controller network), eth3 / eth4 into the second detection domain (replicated active-active network), and eth5 / eth6 into the third detection domain (backend network). Each domain trains its own BP neural network. By constructing sub-health scenarios such as congestion, packet loss, and abnormal latency, 5000 sets of positive samples (normal data) and 3000 sets of negative samples (sub-health data) are collected for each domain. After preprocessing, these samples are input into the model for training. The final accuracy rates are 96.8% for the first detection domain, 95.2% for the second, and 97.1% for the third. The standardized two-dimensional data of eth0 is then input into the first detection domain model, outputting a candidate sub-health label "1". If the network card is labeled as a candidate sub-healthy for three consecutive cycles (30 seconds cumulatively), reaching the preset threshold of 3 for the first detection domain, eth0 is determined to be in a sub-healthy network state.

[0089] The application results of this embodiment show that, through the synergistic effect of various technical features, the sub-health state of eth0 (the core anomaly being a persistent packet loss rate higher than 1%) was successfully detected before the business was significantly affected (at which point the upper-layer transaction response time only increased by 5 milliseconds, not exceeding the threshold). An alarm message containing the network interface card identifier, network type, and anomaly indicators was immediately generated and sent to the operations and maintenance platform. Simultaneously, redundant links in the cluster were detected (eth1 of controller 1 can replace eth0 for inter-controller network communication), and isolation operations were automatically executed after the isolation conditions were met. Compared with existing technologies, the technical solution of this invention improves the detection accuracy by approximately 23%, reduces the false positive rate to below 1.5%, and provides an average early warning time of 2-5 minutes for network sub-health, effectively avoiding performance degradation or paralysis of the storage system caused by the deterioration of the sub-health state. This fully verifies the technical advantages and practical value of this invention.

[0090] Another embodiment of the present invention provides a network status detection system; for details, please refer to [link to relevant documentation]. Figure 4 , Figure 4 The diagram shown illustrates a structural block diagram of a network state detection system according to one embodiment of the present invention, comprising: The acquisition module 11 is used to acquire network topology information between controllers in the storage cluster, wherein the network topology information includes each physical network card associated with the network connection relationship of different network types; Evaluation module 12 is used to determine the service cycle load of each physical network card within a preset time period, and dynamically determine the number of active probe packets sent by each physical network card based on the service cycle load. The detection module 13 is used to determine the target network address based on the network topology information and initiate an active detection request to the target network address. The number of active detection requests is determined by the number of corresponding active detection messages sent. The collection module 14 is used to obtain the active probe response result in response to the active probe request, and to obtain the service network performance result within the active probe request period; Module 15 is used to construct various state detection models corresponding to different network types; The detection module 16 is used to input the active detection response result and the service network performance result into the corresponding state detection model to obtain the network state detection result of the storage cluster.

[0091] Another embodiment of the present invention provides a network status detection device, see [link to relevant documentation]. Figure 5 , Figure 5 This is a schematic diagram of a network status detection device provided in an embodiment of the present invention. The network status detection device includes a processor 21, a memory 22, and a computer program stored in the memory 22 and configured to be executed by the processor 21. When the processor 21 executes the computer program, it implements the steps described in the above-described network status detection method embodiment, for example... Figure 1 The steps S1 to S6 described above; or, when the processor 21 executes the computer program, it implements the functions of each module in the above-described device embodiments, such as the acquisition module 11.

[0092] For example, the computer program may be divided into one or more modules, which are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the network status detection device.

[0093] The network status detection device may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will understand that the schematic diagram is merely an example of a network status detection device and does not constitute a limitation on the device. It may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the network status detection device may also include input / output devices, network access devices, buses, etc.

[0094] The processor 21 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 21 is the control center of the network status detection device, connecting various parts of the entire network status detection device through various interfaces and lines.

[0095] The memory 22 can be used to store the computer programs and / or modules. The processor 21 implements various functions of the network status detection device by running or executing the computer programs and / or modules stored in the memory 22 and calling the data stored in the memory 22. The memory 22 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0096] If the modules integrated into the network status detection device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0097] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0098] Accordingly, embodiments of the present invention provide a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform steps in the network status detection method of the above embodiments, for example... Figure 1 Steps S1 to S6 as described above.

[0099] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A network state detection method, characterized in that, include: Obtain network topology information between controllers in the storage cluster, wherein the network topology information includes each physical network interface card associated with network connection relationships of different network types; Determine the service cycle load of each physical network interface card (NIC), and dynamically determine the number of active probe packets sent by each physical NIC based on the service cycle load. Based on the network topology information, the target network address is determined, and an active probe request is initiated to the target network address. The number of active probe requests is determined by the number of corresponding active probe packets sent. Obtain the active probe response result in response to the active probe request, and obtain the service network performance results within the active probe request period; Construct state detection models corresponding to different network types; The active detection response result and the service network performance result are input into the corresponding state detection model to obtain the network state detection result of the storage cluster.

2. The network status detection method as described in claim 1, characterized in that, The network types include control room network, replication dual-active network, and backend network; The process of obtaining network topology information between controllers in the storage cluster includes: Based on the preset configuration information of the storage cluster, the IP address connection relationship between each node in the cluster is obtained, and the corresponding physical network card identifier is determined through the IP address connection relationship, so as to obtain the first connection relationship of the control network and its associated first physical network card. Obtain the preset logical interface of the storage cluster, determine the corresponding physical network card identifier based on the port group information bound to the preset logical interface, and obtain the second connection relationship and associated second physical network card of the replication dual-active network; Based on the preset configuration information of the storage cluster, the IP address connection relationship between the nodes in the cluster and the back-end storage device is obtained. The corresponding physical network card identifier is determined through the IP address connection relationship, and the third connection relationship of the back-end network and the associated third physical network card are obtained. The network topology information of the storage cluster is generated by integrating the first connection relationship, the first physical network card, the second connection relationship, the second physical network card, the third connection relationship, and the third physical network card.

3. The network status detection method as described in claim 1, characterized in that, The step of determining the service cycle load of each physical network interface card (NIC) and dynamically determining the number of active probe packets sent by each physical NIC based on the service cycle load includes: A reference value for the single-cycle service load is determined based on historical load data, and the load of the service cycle is compared with the reference value. If the service cycle load is lower than the reference value, the number of active probe messages sent is generated based on the difference between the service cycle load and the reference value. If the load of the service cycle is not lower than the reference value, no additional number of active probe messages will be generated.

4. The network status detection method as described in claim 1, characterized in that, The step of determining the target network address based on the network topology information and initiating an active probe request to the target network address includes: Using the network to which the physical network card currently belongs as the source address, target network addresses are filtered based on the connection relationships in the network topology information, wherein the target network addresses cover all associated nodes under the network type corresponding to the physical network card; Based on the number of active probe messages sent, probe requests are initiated to each of the target network addresses, and the number of responses and response times of the probe requests are obtained; Based on the number of responses and the response time, the active detection response result is calculated.

5. The network status detection method as described in claim 1, characterized in that, The construction of various state detection models corresponding to different network types includes: All physical network cards corresponding to the same network type are divided into the same independent detection domain, wherein the network performance evaluation benchmark of the independent detection domain is preset according to the characteristics of the corresponding network type; Construct an initial network state detection model corresponding to each independent detection domain, and obtain historical service network performance results and historical active probe response results in the storage cluster to generate a training dataset; The initial network state detection model is trained based on the training dataset to obtain the network state detection model for each independent detection domain.

6. The network status detection method as described in claim 1, characterized in that, The step of inputting the active detection response result and the service network performance result into the corresponding state detection model to obtain the network state detection result of the storage cluster includes: The service network performance results and active detection response results of each physical network card within the same independent detection domain are input into the network status detection model corresponding to the detection domain, and the candidate sub-health markers of each physical network card in the current time period are output. Record the time period during which each physical network card is marked as a candidate sub-healthy, and take the time period of the first time it is marked as the candidate sub-healthy as the statistical starting point, and accumulate the number of consecutive periods in which it is marked as the candidate sub-healthy. If the number of consecutive cycles reaches the preset cycle determination threshold, the physical network card is determined to be in a sub-healthy state; if any cycle during the accumulation process is not marked as a candidate sub-healthy state, the consecutive cycle count is reset. The network status detection results of the storage cluster are generated by integrating the judgment results of all the physical network cards.

7. The network status detection method as described in claim 6, characterized in that, After generating the network status detection results of the storage cluster, the method further includes: Based on the network status detection results, identify the target physical network interface card (NIC) that is in a sub-healthy network state. Generate alarm information containing the identification information, network type, and sub-health characteristic parameters of the target physical network card, and send the alarm information to a preset receiving end; The network redundancy configuration of the storage cluster is detected, and it is determined whether there are redundant links or redundant network cards that can replace the target physical network card. If the redundant link or redundant network card exists and the preset isolation conditions are met, then an isolation operation is performed on the target physical network card.

8. A network status detection system, characterized in that, include: The acquisition module is used to acquire network topology information between controllers in the storage cluster. The network topology information includes each physical network card associated with the network connection relationship of different network types. An evaluation module is used to determine the service cycle load of each physical network interface card (NIC) and dynamically determine the number of active probe packets sent by each physical NIC based on the service cycle load. The detection module is used to determine the target network address based on the network topology information and initiate an active detection request to the target network address. The number of active detection requests is determined by the number of corresponding active detection packets sent. The collection module is used to obtain the active probe response results in response to the active probe request, and to obtain the service network performance results within the active probe request period; The building module is used to construct various state detection models corresponding to different network types. The detection module is used to input the active detection response result and the service network performance result into the corresponding state detection model to obtain the network state detection result of the storage cluster.

9. A network status detection device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the network state detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the device containing the computer-readable storage medium executes the computer program, it implements the network status detection method as described in any one of claims 1 to 7.