Abnormality detection method, detector, analysis platform, electronic equipment and storage medium
By combining a prober and an analysis platform with the TCP handshake process for proactive probing and anomaly detection, the timeliness and accuracy of anomaly detection in data centers are solved, enabling comprehensive network monitoring and fault location.
Patent Information
- Application Number
- CN202511217433.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-19
AI Technical Summary
Existing network monitoring tools are unable to detect anomalies in data centers in a timely, accurate, and effective manner, leading to delays in fault location and resolution.
The probe is used to actively probe based on the TCP handshake process, obtain the probe results and perform anomaly detection, use the analysis platform for granular detection and alarm, and combine the network device topology map to locate the fault.
It enables timely, accurate, and comprehensive anomaly detection and fault location in data centers, improving anomaly detection efficiency and fault location accuracy.
Smart Images

Figure CN121173705A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular to the technical field of network monitoring, monitoring and management of data center, and the like, and especially relates to an anomaly detection method, a probe machine, an analysis platform, an electronic device and a storage medium. BACKGROUND
[0002] In order to ensure that the service server of the data center can provide stable and reliable service, in the prior art, the data center can be monitored to facilitate timely solution when the data center has an anomaly or a fault.
[0003] The existing network monitoring technology can include passive network monitoring and active network monitoring. For example, the traditional passive monitoring technology can include network monitoring of the data center based on the Simple Network Management Protocol (SNMP). The SNMP is the most widely used device management and monitoring protocol in the industry. It obtains performance indicators such as CPU, memory utilization, port traffic, error counters by periodically polling the Management Information Base (MIB) variables of network devices such as switches and routers, and thus realizes network monitoring. SUMMARY
[0004] The present disclosure provides an anomaly detection method, a probe machine, an analysis platform, an electronic device and a storage medium.
[0005] According to an aspect of the present disclosure, an anomaly detection method is provided, comprising:
[0006] obtaining a detection result of a probe machine detecting a service server of a data center in a current time period; when the probe machine detects the service server, active detection is performed based on a handshake process of the Transmission Control Protocol (TCP);
[0007] performing at least one granularity of anomaly detection in the data center based on the detection result of the probe machine detecting the service server in the current time period;
[0008] in response to detecting a target granularity anomaly, issuing an alarm information.
[0009] According to another aspect of the present disclosure, an anomaly detection method is provided, comprising:
[0010] constructing a detection packet of a probe machine detecting a service server of a data center in a current time period;
[0011] based on a handshake procedure of a transmission control protocol, actively probe the service server in the current time period using the probe packet;
[0012] obtain a probe result of the probe machine probing the service server in the current time period;
[0013] send the probe result corresponding to the service server in the current time period to an analysis platform, so that the analysis platform performs at least one granularity of anomaly detection in the data center based on the probe result of the probe machine probing the service server in the current time period; and in response to detecting a target granularity anomaly, issue an alarm information.
[0014] According to still another aspect of the present disclosure, an analysis platform is provided, comprising:
[0015] an obtaining module configured to obtain a probe result of a probe machine probing a service server of a data center in a current time period; the probe machine actively probes the service server based on a handshake procedure of a transmission control protocol when probing the service server;
[0016] an anomaly detection module configured to perform at least one granularity of anomaly detection in the data center based on the probe result of the probe machine probing the service server in the current time period;
[0017] an alarm module configured to issue an alarm information in response to detecting a target granularity anomaly.
[0018] According to still another aspect of the present disclosure, a probe machine is provided, comprising:
[0019] a constructing module configured to construct a probe packet for the probe machine to probe a service server of a data center in a current time period;
[0020] a probing module configured to actively probe the service server in the current time period using the probe packet based on a handshake procedure of a transmission control protocol;
[0021] an obtaining module configured to obtain a probe result of the probe machine probing the service server in the current time period;
[0022] a sending module configured to send the probe result corresponding to the service server in the current time period to an analysis platform, so that the analysis platform performs at least one granularity of anomaly detection in the data center based on the probe result of the probe machine probing the service server in the current time period; and in response to detecting a target granularity anomaly, issue an alarm information.
[0023] According to still another aspect of the present disclosure, there is provided an anomaly detection system, comprising: at least one probe machine, a control center device and an analysis platform;
[0024] Each of the probe machines is connected with the control center device and the analysis platform respectively; each of the probe machines adopts the probe machine of the aspects and any possible implementation manners described above; the analysis platform adopts the analysis platform of the aspects and any possible implementation manners described above.
[0025] According to still another aspect of the present disclosure, there is provided an electronic device, comprising:
[0026] at least one processor; and
[0027] a memory connected with the at least one processor in communication; wherein,
[0028] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the aspects and any possible implementation manners described above.
[0029] According to still another aspect of the present disclosure, there is provided a non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method of the aspects and any possible implementation manners described above.
[0030] According to still another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method of the aspects and any possible implementation manners described above.
[0031] According to the technology of the present disclosure, the anomaly of the data center can be detected in time, accurately and effectively, and the efficiency of anomaly detection can be effectively improved.
[0032] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0033] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:
[0034] Figure 1 is an application architecture diagram of anomaly detection provided by the present disclosure;
[0035] Figure 2 is a network architecture schematic diagram of a data center provided by the present disclosure;
[0036] Figure 3 This is a schematic diagram based on the first embodiment of the present disclosure;
[0037] Figure 4 This is a schematic diagram according to the second embodiment of the present disclosure;
[0038] Figure 5 This is a network device topology diagram provided in an embodiment of the present disclosure;
[0039] Figure 6 This is a schematic diagram according to the third embodiment of the present disclosure;
[0040] Figure 7 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0041] Figure 8 This is a schematic diagram of an application scenario provided according to an embodiment of this disclosure;
[0042] Figure 9 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0043] Figure 10 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0044] Figure 11 This is a schematic diagram according to the seventh embodiment of the present disclosure;
[0045] Figure 12 This is a schematic diagram according to the eighth embodiment of the present disclosure;
[0046] Figure 13 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0047] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0048] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0049] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.
[0050] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0051] Most existing traditional monitoring tools, such as SNMP-based monitoring tools, rely on polling or sampling intervals of seconds or longer, which makes them unable to monitor data center anomalies in a timely, accurate, and effective manner when monitoring networks.
[0052] Figure 1 This is an application architecture diagram for anomaly detection provided in this disclosure. For example... Figure 1 As shown, this anomaly detection scenario may include a detector 100, an analysis platform 200, and a control center device 300.
[0053] The probe 100, also known as a data acquisition node, simulates real user access to detect the business servers 410 in the data center 400. The architecture shown in Figure 1 is merely a schematic diagram of this embodiment and does not limit the location or number of probes deployed in actual applications, or the structure and number of business servers 410 within the data center 400. For example, to improve anomaly detection efficiency, the probe 100 in this embodiment may include one, two, or more.
[0054] The analysis platform 200 is used to receive the detection results uploaded by the detector 100 and to perform anomaly detection.
[0055] The control center device 300 is used to pre-configure information about the data center's business servers, such as the target Internet Protocol (IP) address and available target port identifiers of the business servers to be probed, as well as the source IP address of the probe machine. Specifically, the available target ports of the business servers should avoid the ports used by the business servers to provide business services, so as to avoid affecting the services provided by the business servers during anomaly detection.
[0056] The detector 100 can connect to the control center equipment 300 to obtain information about the detector and the business server that can be detected, thereby enabling the detection of the business server.
[0057] Figure 2 This is a schematic diagram of a data center network architecture provided in this disclosure. For example... Figure 2 As shown, taking the data center using the CLOS network architecture as an example, in practical applications, data centers can also use other network architectures or network architectures that are improved based on the CLOS network architecture, which is not limited here.
[0058] like Figure 2 As shown, the network architecture of a data center, from top to bottom, can include spine switches, point-of-delivery (Pod) switches, and top-of-rack (TOR) switches. Spine switches are one of the core components of the CLOS network architecture, primarily responsible for providing high-speed connectivity between different layers of the data center. They are typically located at the top layer of the network, connecting all Pod switches.
[0059] Pod switches are typically located below Spine switches, forming a smaller aggregation point (Pod). Each Pod can contain multiple server racks, and each rack is connected to the Pod switch via a TOR switch.
[0060] The TOR switch is located at the top of the server rack and is directly connected to the motherboard or other network interface card (NIC) of the business server.
[0061] The switches in this embodiment can be collectively referred to as network devices, and therefore can also be called Spine network devices, Pod network devices, and TOR network devices. Spine network devices, Pod network devices, and TOR network devices are structured according to a top-to-bottom layer hierarchy. Figure 2 The diagram shows the architecture of a three-tier network device.
[0062] Based on the above Figure 1 and Figure 2 The architecture shown describes the technical solution of this disclosure in detail.
[0063] Figure 3 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 3 As shown, this embodiment provides an anomaly detection method, applied in an analysis platform, which may specifically include the following steps:
[0064] S301. Obtain the detection results of the probe machine probing the business servers of the data center within the current time period; when the probe machine probes the business servers, it actively probes based on the handshake process of Transmission Control Protocol (TCP);
[0065] In this embodiment, the current time period can be one cycle per second. In practical applications, other time lengths, such as 2 seconds, can also be set according to requirements, and are not limited here.
[0066] When the probe machine probes the business server, it can use raw sockets to send probe packets with underlying TCP synchronization sequence numbering (SYN) and listen for the business server's response packets. No actual complete TCP connection is established between the probe machine and the business server; only the TCP handshake process is utilized to achieve the probe's detection. After the probe, it can report the results to the analysis platform.
[0067] In this embodiment, when the probe probes the business server within the current time period, it can select an appropriate packet sending frequency according to actual needs. For example, it can select 20 PPS. PPS is an abbreviation for Packets Per Second. In the scenario of this embodiment, it means that when the current time period is 1 second, the packet sending frequency per second can be 20.
[0068] S302. Based on the detection results of the probe machine on the business server within the current time period, perform at least one granularity of anomaly detection within the data center.
[0069] In this embodiment, the granularity can be a division based on the data center structure. Larger granularity results in a wider coverage area for anomaly detection, while smaller granularity results in a smaller coverage area. For example, the entire data center can be considered the largest granularity. Following a hierarchical structure from large to small, the granularity can also be selected at the cluster level, or any network device at any level within a cluster can be considered a granularity. For example... Figure 2 In the network architecture of the data center shown, each level of network device can also be considered as a granularity, such as Spine network device, Pod network device, or TOR network device granularity.
[0070] In practical applications, when data centers adopt other network architectures, such as two-tier or five-tier network architectures, anomaly detection can be performed by selecting each network device at one level as a granularity.
[0071] In this embodiment, based on the detection results of the current time period detector, anomaly detection of any granularity can be performed within the data center, covering any level of detection in the data center's network architecture.
[0072] S303. When an abnormality in the target granularity is detected, an alarm message is issued.
[0073] Specifically, in this embodiment, the alarm information may be in the form of voice alarms, or it may be sent to maintenance personnel in text form using pre-configured contact information. Other forms of alarms may also be used, which are not limited here.
[0074] Taking anomaly detection in a data center as an example, the number of probes and the number of service servers used for detection are not limited. For example, only one probe can be used, located either inside or outside the data center. If there are two or more probes, they can be deployed both inside and outside the data center. In this embodiment, although probes are used to detect service servers, the detection actually checks for anomalies in the network devices or ports along the path from the probe to the service server. Therefore, the selection of the service server for detection does not focus on the service server's business operations. In practical applications, various performance indicators of each service server under each TOR network device, such as response time, throughput, and resource utilization, can be comprehensively considered to select the most suitable service server for detection.
[0075] The anomaly detection method of this embodiment can perform anomaly detection at least one granularity within the data center based on the detection results of the probe machine on the business server within the current time period. When an anomaly at the target granularity is detected, an alarm message is issued, thus achieving anomaly detection within the data center. According to the anomaly detection method of this embodiment, anomalies at any time within the current time period can be detected, enabling timely, accurate, and effective detection of anomalies in the data center, effectively improving the efficiency of anomaly detection.
[0076] Figure 4 This is a schematic diagram based on the second embodiment of this disclosure; the anomaly detection method of this embodiment, in the above... Figure 3 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 4 As shown, the anomaly detection method in this embodiment may specifically include the following steps:
[0077] S401. Obtain the number of packets sent and lost by the probe machine during the current time period when it probes the business server.
[0078] In this step, we take the detection results, which include the number of packets sent and lost by the probe machine probing the business server, as an example. Since packet loss is a very important manifestation of anomalies, in this embodiment, by obtaining the number of packets sent and lost by the probe machine probing the business server within the current time period, we can effectively assist in subsequent anomaly detection and fault location.
[0079] S402. Based on the number of packets sent and lost obtained by the probe machine from probing the business server within the current time period, obtain the number of packets sent and lost at least one granularity in the data center, cluster, and network devices.
[0080] The network device in this embodiment is a device located at the upper layer of the business server in the network architecture, used to assist in the deployment of the business server.
[0081] For example, the network device in this embodiment can be the one described above. Figure 2 The data center network architecture shown is granular at the level of Spine switches, Pod switches, or TOR switches.
[0082] Specifically, statistical methods can be used to count the number of packets sent and lost at each granularity. For example, at the data center granularity, the total number of packets sent and lost when the probe machine probes all service servers within that data center during the current time period can be counted. At the TOR granularity, the total number of packets sent and lost when the probe machine probes all service servers within the TOR coverage area during the current time period can be counted. The principle is the same for counting the number of packets sent and lost for other granularities of network devices, and will not be elaborated here.
[0083] S403. Based on the number of packets sent and lost at each granularity, determine the packet loss rate for the corresponding granularity;
[0084] Specifically, the number of lost packets is divided by the number of sent packets to obtain the corresponding packet loss rate.
[0085] S404. Detect whether the packet loss rate of each granularity is greater than the preset packet loss ratio at the corresponding granularity;
[0086] S405. When the packet loss rate at the target granularity is detected to be greater than the corresponding preset packet loss ratio, an alarm message is issued.
[0087] The preset packet loss ratio in this embodiment can be set according to actual needs to determine the maximum allowable packet loss ratio, such as 1%, 2%, etc. If the packet loss rate at a certain granularity exceeds the corresponding preset packet loss ratio within a time period, it is considered that there is an anomaly within the range covered by that granularity in the data center's network architecture.
[0088] Since the number of network devices covered by different granularities is not the same, the preset packet loss ratios corresponding to different granularities can be different.
[0089] In this embodiment, the number of abnormal target granularities detected in each time period is not limited. Target granularities with a packet loss rate greater than the preset packet loss ratio at the corresponding granularity need to be detected and an alarm is triggered.
[0090] In practical applications, to improve the accuracy of anomaly detection, all levels and all granularities that need to be detected can be pre-configured according to requirements. For example, all clusters need to be detected, and the granularity of each network device at each level in each cluster needs to be detected. This can cover all scenarios in the data center, thereby effectively improving the comprehensiveness and accuracy of anomaly detection.
[0091] In this embodiment, if the packet loss rate of each granularity after detection is not greater than the preset packet loss ratio of the corresponding granularity, it is determined that there is no anomaly in the data center, and the detection can continue to the next time period without any alarm.
[0092] It should be noted that the above anomaly detection uses a single probe as an example to describe the technical solution of this disclosure. In practical applications, at least two probes can be used to probe the same data center. Because different probes are located in different positions, even if different probes probe the same business server, their respective probe paths will be different, and ultimately the network devices detected will also be different. Therefore, based on the detection results of each probe, an alarm should be triggered when an anomaly at the target granularity is detected. By using two or more probes, the detection coverage can be effectively improved, thereby enabling comprehensive network monitoring of the data center.
[0093] S406. Based on the pre-created network device topology map and the five-tuple information of the probe packets lost in the current time period, perform fault location.
[0094] Specifically, since the target granularity of the anomaly has been determined during the aforementioned anomaly detection, it could be an anomaly occurring in the entire data center, a specific cluster within the data center, or within the coverage area of a specific network device within the cluster. Based on this, when this step is implemented, fault location can be performed within the target granularity coverage area using a pre-created network device topology map and the five-tuple information of the lost probe packets in the current time period. This reduces fault location time and improves fault location efficiency.
[0095] Furthermore, the detection results obtained in step S401 may also include the 5-tuple information of the lost probe packets. For example, the 5-tuple information may include: the source IP address of the probe machine, the target IP address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the probe machine used in the probe packet.
[0096] Before step S406, the following steps may also be included:
[0097] (1) Information on network nodes included in the communication path between the probe and each business server in the data center;
[0098] In this embodiment, the network node refers to the network device. The information of the network nodes included in the communication path between the probe and the business server includes not only the network node's identifier, the network node's port, and the communication direction of adjacent network nodes, but also the network node's identifier, port, and the communication direction of adjacent network nodes.
[0099] In this specific implementation, the probe can send probe packets with incrementing Time-To-Live (TTL) to the service server and collect information about each network node along the way, thereby obtaining information about the network nodes included in the communication path between the probe and the service server. This step can also be called the Traceroute process.
[0100] (2) Based on the information of each network node in the communication path between the probe and each business server, construct a network device topology diagram.
[0101] Specifically, by collecting a large number of communication paths, the communication relationships between network nodes within these paths can be obtained, allowing for the construction of a network device topology map. For example, Figure 5 This is a network device topology diagram provided in an embodiment of the present disclosure. Figure 5 This is merely an exemplary network device topology diagram and does not limit the number of network node layers, the number and distribution of network nodes at each layer, or the number and distribution of service servers. In practical applications, other network device topology diagrams can be obtained, which are not limited here.
[0102] The port through which the preceding network node communicates with the following network node in a communication path can be identified in the attribute information of the preceding network node.
[0103] In this embodiment, the constructed network device topology diagram depicts all possible communication paths between the probe and the service server.
[0104] It should be noted that in this embodiment, the network device topology map constructed using steps (1)-(2) above can be constructed once at certain time intervals. This time interval is much longer than the time interval for the probe to probe the service server. Specifically, the length of the time interval for constructing the network device topology map can be set according to the update cycle of the network devices in actual application to ensure the accuracy of the network device topology map. For example, it can be updated and constructed hourly, daily, or according to other time intervals.
[0105] By using the above methods, network device topology diagrams can be constructed accurately and effectively.
[0106] Alternatively, in one embodiment of this disclosure, step 406 may be implemented using the following steps:
[0107] (a) Based on the quintuple information of the probe packets lost within the current time period, obtain the target network nodes that multiple lost probe packets within the target granularity coverage area pass through from the network device topology diagram.
[0108] In this embodiment, the network node refers to a network device.
[0109] (b) Determine the location of the fault based on the packet loss rate of the target network node within the current time period.
[0110] Specifically, based on the five-tuple information of the lost probe packets, the network nodes and ports traversed along the path from the probe machine to the service server can first be determined. Then, since the target granularity of the anomaly has already been determined during the aforementioned anomaly detection, the target network node commonly traversed by multiple lost probe packets is obtained within the target granularity coverage area of the network device topology diagram. The packet loss rate of this target network node in the current time period is then analyzed to accurately locate the fault.
[0111] For example, step (b) can be implemented in the following ways:
[0112] In the first scenario, if the packet loss rate of all ports in the target network node is greater than the preset ratio within the current time period, the target network node is determined to be faulty.
[0113] In this embodiment, the preset ratio can be greater than the preset packet loss ratio. For example, it can be set to 50%, 60%, etc., depending on actual needs. If all ports experience a certain degree of packet loss, the current target network node is considered to be faulty.
[0114] In practical applications, if the underlying network devices of all ports of the target network node fail simultaneously, the packet loss rate of all ports of the target network node will be greater than the preset ratio. However, the probability of the underlying network devices of all ports failing simultaneously is small. In this embodiment, in order to improve the fault detection efficiency, the situation where the underlying network devices of all ports of the target network node fail simultaneously can be ignored.
[0115] In the second scenario, if the packet loss rate of only a specified port in the target network node is greater than a preset ratio within the current time period, the fault location is determined based on the position of the target network node in the network architecture.
[0116] If the packet loss rate of only a specified port of the target network node exceeds a preset percentage, it indicates that all other ports of the target network node are normal. The anomaly of the specified port may be due to a problem with the port itself or a problem with the lower-level network device connected to that port. Further analysis of the target network node's position within the network architecture is needed to pinpoint the location of the fault.
[0117] For example, if the target network node is the lowest-level network device above the business server in the network architecture, the specified port of the target network node is determined to be faulty. Since there are no other network devices below the target network node in this case, the specified port of the target network node can be directly considered to be faulty.
[0118] If the target network node is not the lowest-level network device above the business server in the network architecture, information about the next-layer network node connected through the specified port can be obtained from the network device topology diagram; and the fault location can be determined based on the packet loss rate of the next-layer network node in the current time period.
[0119] Specifically, the packet loss rate of the next-layer network node can be obtained by statistical analysis according to the method described in the above embodiments.
[0120] Based on the packet loss rate of the next-layer network node within the current time period, the fault location is determined. This is equivalent to using the next-layer network node as the target network node in step (b) to determine the fault location. For a detailed explanation of the implementation principle, please refer to the specific implementation method of step (b) above, which will not be repeated here.
[0121] The anomaly detection method in this embodiment performs anomaly detection on at least one granularity of data center, cluster, and network devices within the current time period, and issues an alarm when an anomaly is detected at the target granularity. This enables timely, accurate, and comprehensive detection of anomalies in the data center and timely issuance of alarms, allowing staff to be informed in a timely manner and resolve problems promptly.
[0122] Furthermore, the technical solution of this embodiment can accurately and effectively locate faults when an anomaly is detected. Specifically, fault location can be performed based on a pre-created network device topology map and the five-tuple information of probe packets lost in the current time period, which can effectively improve the accuracy and efficiency of fault location.
[0123] Figure 6 This is a schematic diagram based on the third embodiment of this disclosure; as shown Figure 6 As shown, this embodiment provides an anomaly detection method applied to the detector side, which may specifically include the following steps:
[0124] S601. Construct a probe packet for the prober to probe the business servers of the data center within the current time period;
[0125] In this embodiment, during anomaly detection, the probe needs to send probe packets to the business server. Therefore, it is necessary to first build probe packets that can be sent to the business server on the probe side.
[0126] It should be noted that, referring to the description in the above embodiments, within a time period, the probe can send probe packets to a service server at a frequency of 20 PPS, and each probe packet needs to be constructed.
[0127] S602. Based on the TCP handshake process, probe packets are used to actively probe the business server within the current time period;
[0128] Using raw sockets, the probe sends underlying TCP synchronization sequence number (SYN) packets and listens for responses from the business server. A complete TCP connection is not actually established between the probe and the business server; only the first two steps of the TCP handshake are utilized to enable the probe to actively probe the business server.
[0129] In this embodiment, the probe packet uses a TCP SYN probe packet instead of a traditional Internet Control Message Protocol (ICMP) message. This is because ICMP messages have low priority and are easily affected by firewalls and rate-limiting policies, and their results cannot accurately reflect the performance of the TCP protocol carrying application traffic. TCP SYN probe packets, on the other hand, can efficiently detect transport layer connectivity without incurring the additional system resource overhead of establishing a full TCP connection. Furthermore, without the packet loss retransmission mechanism of a full TCP connection, they can accurately reflect the network's packet loss situation.
[0130] S603. Obtain the detection results of the probe machine probing the business server within the current time period;
[0131] S604. Send the detection results of the business server within the current time period to the analysis platform so that the analysis platform can perform anomaly detection at least one granularity within the data center based on the detection results of the probe on the business server within the current time period; and issue an alarm message when an anomaly at the target granularity is detected.
[0132] The technical solution of this embodiment is similar to that described above. Figure 3 The difference between the technical solutions in the illustrated embodiments is that: Figure 3 The illustrated embodiment describes the technical solution of this disclosure on the analysis platform side, while this embodiment describes the technical solution of this disclosure on the detector side. The specific implementation principle can also be referred to the above. Figure 3 The embodiment shown is described.
[0133] The anomaly detection method in this embodiment actively probes the business server using constructed probe packets within the current time period through a TCP handshake process. It obtains the probe results of the probe on the business server within the current time period and sends them to the analysis platform for anomaly detection. This method can detect the communication path between the probe and the business server in a timely and accurate manner and transmit the detection results to the sub-platform in a timely and accurate manner, thereby enabling timely, accurate and effective anomaly detection.
[0134] Figure 7 This is a schematic diagram based on the fourth embodiment of this disclosure; the anomaly detection method of this embodiment is described above. Figure 6 Based on the technical solutions of the embodiments described above, the technical solutions of this disclosure will be further described in more detail. For example... Figure 7 As shown, the anomaly detection method in this embodiment may specifically include the following steps:
[0135] S701. Obtain the target IP address and target port identifier of the business server, as well as the source IP address of the probe machine, from the control center equipment;
[0136] The control center can be pre-configured with information such as the target IP address and target port identifier of the business server that can be detected, as well as the source IP address of the probe machine.
[0137] Furthermore, it should be noted that in the embodiments of this disclosure, the number of business servers in the data center is very large. Therefore, the technical solution of this embodiment cannot be used to probe all business servers individually, as too many probe packets would place excessive pressure on the probe machine and subsequent analysis. Therefore, in this embodiment, the control center device can select the top N business servers with the highest performance index values from each TOR network device based on at least one of the performance indicators of each business server in the data center, such as response time, throughput, and resource utilization, as representatives to participate in the anomaly detection of this disclosure embodiment.
[0138] The performance metrics of each business server can be obtained by weighted summation based on the server's response time, throughput, and resource utilization. Resource utilization can include at least one of the following: CPU utilization, memory usage, disk I / O, and network bandwidth, without limitation here.
[0139] After identifying the top N service servers with the highest performance metrics, you can then configure all information for each server in the control center, such as target IP addresses and target port identifiers. To avoid probes interfering with normal operations, the target port identifier for each service server cannot be the port on which the service server provides the service.
[0140] In practice, if a business server becomes unstable during the detection process, or if the business server is shut down or restarted, the control center equipment can remove the business server, lower its performance index score, and replace it with a business server that has a higher performance index score.
[0141] Additionally, the source IP address of the probe needs to be configured in the control center, and a list of source port identifiers that the probe can choose from can also be configured.
[0142] S702. According to the preset rules, configure the source port identifier of the probe used by each probe packet to be constructed in the current time period in sequence;
[0143] Referring to the description in the above embodiment, the probe can send TCP SYN probe packets to the business server at a relatively high frequency, such as 20 PPS. This "pulsating" probe frequency is much higher than the second-level sampling of traditional monitoring tools, enabling it to capture millisecond-level network "micro-bursts" that traditional methods cannot detect. At the same time, this frequency is not extremely high. This embodiment adopts a balanced probe frequency that is neither too high nor too low to cover micro-bursts because: 1. Since the business server needs to process business, too high a probe frequency can easily affect the business server's business; 2. Too high a frequency of TCP SYN probe packets will be considered as TCP SYN flooding attacks and will easily be dropped by the business server without a response.
[0144] To increase the number of network devices passing through the path during probing, this embodiment can configure the source port identifiers of each probe packet within the current time period sequentially according to preset rules. For example, the preset rule could be to obtain the source port identifiers sequentially from the source port identifier list in a round-robin manner. Alternatively, it could be to obtain them randomly, but ensure that the source port identifiers of each probe packet within the current time period are not repeated. In practical applications, other specific rules can also be used to configure the source port identifiers of each probe packet, which will not be elaborated here.
[0145] The source port identifier list in this embodiment can be configured with all selectable source port identifiers, such as values from 0 to 2000.
[0146] S703. Based on the source IP address of the prober, the target IP address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the prober used in each probe packet, generate each probe packet of the prober probing the service server in the current time period in sequence.
[0147] In this embodiment, each probe packet includes a 5-tuple of information. Furthermore, within the current time period, the source port identifier of the probe in the 5-tuple of all probe packets sent by the probe to this service server is different; while the other 4-tuples are all the same.
[0148] By using step S702, the source port identifier of probe packets from the same prober targeting the same service server within the same time period can be dynamically modified. This allows for the construction of different five-tuple probe packets, which are then sent to the service server. Different five-tuple probe packets can cover more network devices and more ports in the network structure. For example, using step S703, after multiple time periods of active probing, full coverage of all network devices and ports can be achieved, effectively improving the comprehensiveness and accuracy of anomaly detection. Based on experience and practical verification, a five-tuple count of 10 times the number of ports can cover all ports of all network devices.
[0149] S704. The handshake process based on TPC sends each probe packet to the business server in sequence within the current time period and obtains the response results of each probe packet in sequence.
[0150] For example, for each probe packet, obtain the response message returned by the business server based on the probe packet; the response message includes a business response or a response to terminate the abnormal connection.
[0151] Specifically, for any probe packet, the target port of the business server is the listening port. In this case, the business server will return SYN+ACK and attempt to establish a TCP connection. However, there is no complete TCP connection between the probe and the business server. The probe's kernel protocol stack will automatically return a Reset (RST) packet to the business server, and the business server will close this half-open connection.
[0152] If the target port of the business server is a non-listening port, the business server will return an RST packet to the probe machine.
[0153] In practical applications, packet loss may occur, meaning that after a probe packet is sent, no response is received from the business server. Therefore, if no response message is received from the business server based on the probe packet within a preset time period, the probe packet is considered lost.
[0154] Through the above analysis, the detection results of each detection packet within the current time period can be obtained comprehensively and accurately.
[0155] S705. Based on the response results of each probe packet within the current time period, aggregate the number of packets sent and the number of packets lost in the current time period according to the source IP address of the probe machine and the target IP address of the business server.
[0156] In practical applications, aggregation can be skipped, and the response results of each probe packet can be directly uploaded to the analysis platform, which then analyzes the number of packets sent and lost in each time period. However, the data across the entire network is enormous. To reduce the pressure on the subsequent analysis platform, this step is preferably adopted, where the probe data is pre-aggregated on the probe machine. Aggregating once per second in each time period can effectively reduce the amount of data uploaded and improve the anomaly detection efficiency of the analysis platform.
[0157] S706. Obtain the quintuple information corresponding to the lost probe packet;
[0158] If packet loss occurs, this embodiment will also need to obtain the 5-tuple information corresponding to the lost probe packet. Of course, if no probe packet loss occurs within a certain time period, this information will be empty.
[0159] S707. Send the number of packets sent, the number of packets lost, and the five-tuple information corresponding to the lost probe packets for the current time period to the analysis platform, so that the analysis platform can perform at least one granularity of anomaly detection in the data center based on the probe results of the probe machine probing the business server within the current time period; and issue an alarm message when an anomaly at the target granularity is detected.
[0160] Additionally, it should be noted that in this embodiment, step S705 aggregates the number of packets sent and lost in the current time period as an example. In practical applications, the average latency of each probe packet within the current time period can also be aggregated. Of course, this refers to the average latency of probe packets that have a response, excluding lost probe packets. This average latency can also be used as the probe result and sent to the analysis platform. Correspondingly, the analysis platform can calculate the average latency of each granularity in the current time period according to the statistical method of packet loss rate at each granularity. For example, the average latency of probe packets from probes within the data center probing all business servers within the current time period can be obtained as the average latency at the data center granularity. The average latency of probe packets from probes within a specified cluster probing all business servers within the current time period can also be obtained as the average latency at the cluster granularity. Similarly, the average latency of network devices at any granularity can be obtained. Then, it is further detected whether the average latency of each granularity in the current time period is greater than a preset average latency. This preset average latency can be the maximum allowable latency set based on experience. In this embodiment, the preset average latency set for each granularity can be the same. If the value is greater than the preset average latency, the analysis platform can consider the granularity to be abnormal and issue an alarm. However, for anomalies caused by the average latency exceeding the preset average latency, the analysis platform does not perform further fault localization.
[0161] The anomaly detection method in this embodiment constructs different probe packets from the same prober targeting the same business server within the same time period by dynamically modifying the source port identifier. This effectively expands the port coverage of network devices in the data center, thereby significantly improving the comprehensiveness and accuracy of anomaly detection.
[0162] In this embodiment, the number of packets sent and lost in the current time period can be aggregated according to the source IP address of the probe and the target IP address of the business server, and uploaded to the analysis platform. This can effectively reduce the amount of data transmitted from the probe to the business server, reduce the data processing of the analysis platform, and improve the anomaly detection efficiency of the analysis platform.
[0163] In this embodiment, the five-tuple information corresponding to the lost probe packet can also be sent to the analysis platform, which can effectively assist the analysis platform in fault location when an anomaly is detected, and can effectively improve the accuracy and efficiency of fault location.
[0164] Figure 8 This is a schematic diagram of an application scenario provided according to an embodiment of this disclosure; such as Figure 8As shown, in this application scenario, taking the deployment of two probes as an example, the first probe 801 and the second probe 802 are deployed in the first data center 803 and the second data center 804, respectively. In this way, it can not only cover the network devices across clusters or Pods within the data center, but also cover the network devices from the data center egress to other data centers. It can also partially realize the function of disaster recovery, for example, when the probes across data centers fail, the probes in the data center can still probe the network devices in the data center.
[0165] In practical applications, both detectors can be deployed in one of the data centers to achieve more comprehensive detection of network devices in the same data center.
[0166] This embodiment uses the deployment of two probes as an example. In practical applications, more than two probes can be deployed.
[0167] like Figure 8 As shown, in this application scenario, taking two clusters in the first data center 803, namely the first cluster 803a and the second cluster 803b, as an example, each cluster can include network devices with a three-layer network architecture from top to bottom, namely Spine network devices, Pod network devices, and TOR network devices. In this embodiment, taking the deployment of two business servers under each TOR network device as an example, in actual applications, multiple business servers can be deployed under one TOR network device.
[0168] The network architecture within the second data center 804 can be the same as or different from that of the first data center 803, and is not limited here. The second probe 802 within the second data center 804 is connected to the first analog-to-digital converter (ADC) 8031 within the first data center 803 via a digital converter (DC) 8041. Then, the first ADC 8031 is connected to the second ADC 8032 of the first cluster 803a and the third ADC 8033 of the second cluster 803b, respectively.
[0169] The second ADC8032 is connected to the Spine network device in the first cluster 803a; the third ADC8033 is connected to the Spine network device in the second cluster 803b.
[0170] Based on the above connection relationship, a complete communication path can be formed between the second probe 802 and the business server in the first data center 803.
[0171] Within the first data center 803, the first probe 801 is also connected to the first ADC 8031, forming a complete communication path between the first probe 801 and the business server within the first data center 803. Thus, not only can the first probe 801 located within the first data center 803 probe the business server within the first data center 803, but the second probe 802 located within the second data center 804 can also probe the business server within the first data center 803.
[0172] In this embodiment, the first detector 801 and the second detector 802 can be respectively adopted as described above. Figure 6 or Figure 7 The technical solution of the illustrated embodiment involves active detection. The detection results are then uploaded to the analysis platform 805 for reference. Figure 3 or Figure 4 The technical solution of the illustrated embodiment performs anomaly detection, and when an anomaly is detected, it issues an alarm and further locates the fault.
[0173] In addition, the control center device 806 of this embodiment may also be pre-configured with a list of source IPR addresses and source port identifiers for each probe; it may also be pre-configured with the target IP address and target port identifier of the service server that can be probed, as well as the communication protocol type between the probe and the service server. Based on the above configuration of the control center device 806, the first probe 801 and the second probe 802 can respectively generate probe packets to probe the service server based on the configuration information in the control center device 806. For details, please refer to the description of the relevant method embodiments above, which will not be repeated here.
[0174] use Figure 8 The architecture of the illustrated embodiment can adopt the above-described... Figure 3 , Figure 5 , Figure 6 as well as Figure 7 The technical solutions of the embodiments shown implement the anomaly detection of the embodiments of this disclosure. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0175] The anomaly detection method in this embodiment constructs different probe packets by setting up at least one prober and dynamically modifying the source port identifier of the prober; it actively probes the business server based on the TCP handshake process. This detection method is similar to LiDAR scanning. Given a sufficiently rich set of source port identifiers from the prober, after a certain period of detection, it can comprehensively probe all ports of all network devices within the data center. This detection method can also be called a radar-scanning-style active detection scheme.
[0176] The anomaly detection method of this disclosure can transform network monitoring from "reactive response after the fact" to "proactive prediction before the fact", and is easy to deploy, update and migrate, and has broad application prospects.
[0177] The anomaly detection method of this disclosure can be applied to large-scale cloud services and data center operations.
[0178] In public cloud environments, cloud service providers (CSPs) are required to provide customers with high availability and low latency service level agreements (SLAs). The technical solutions described in this disclosure can be used to continuously and frequently perform "health checks" on the entire data center network, enabling real-time detection of potential performance bottlenecks and faults. This allows for proactive intervention before problems impact customers, effectively ensuring SLA compliance.
[0179] The anomaly detection method of this disclosure does not require the deployment of a probe machine on the business server, thus solving the problem that the business server is not allowed to deploy probe programs without the customer's permission.
[0180] The anomaly detection method of this disclosure can also be used for automated network operation and maintenance and fault diagnosis.
[0181] Traditional troubleshooting is a time-consuming process that relies heavily on manual experience. Network engineers need to manually run various commands such as Ping and Traceroute, and log into devices to view configurations and counters. The anomaly detection method of this disclosure can automate the entire "discovery-inference-location" process, reducing fault location time from tens of minutes to seconds, significantly improving operational efficiency and reducing mean time to repair (MTTR).
[0182] The anomaly detection method of this disclosure can also directly output the fault location results to the automated operation and maintenance platform, triggering corresponding repair processes such as device isolation, traffic migration, and link restart, thereby realizing the network's self-healing capability and further reducing the reliance on manual intervention.
[0183] Figure 9 This is a schematic diagram according to the fifth embodiment of this disclosure; as shown Figure 9 As shown, this embodiment provides an analysis platform 900, including:
[0184] The acquisition module 901 is used to acquire the detection results of the probe machine probing the business server of the data center within the current time period; when the probe machine probes the business server, it actively probes based on the handshake process of the transmission control protocol.
[0185] Anomaly detection module 902 is used to perform anomaly detection at least one granularity within the data center based on the detection results of the probe machine on the business server within the current time period.
[0186] The alarm module 903 is used to issue an alarm message in response to the detection of an abnormality in the target granularity.
[0187] The analysis platform 900 in this embodiment achieves the same implementation principle and technical effect of anomaly detection by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0188] Figure 10 This is a schematic diagram according to the sixth embodiment of this disclosure; as shown Figure 10 As shown, this embodiment provides an analysis platform 1000, in the above... Figure 9 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 10 As shown, the analysis platform 1000 in this embodiment includes the above-mentioned... Figure 9 The modules with the same name and function shown are: acquisition module 1001, anomaly detection module 1002, and alarm module 1003.
[0189] In this embodiment, the acquisition module 1001 is used for:
[0190] Obtain the number of packets sent and lost by the probe machine during the current time period.
[0191] Further optionally, in one embodiment of this disclosure, the anomaly detection module 1002 is used for:
[0192] Based on the number of packets sent and the number of packets lost obtained by the probe machine from probing the service server within the current time period, the number of packets sent and lost at least one granularity in the data center, cluster, and network device is obtained; the network device is a device located above the service server in the network architecture and used to assist in the deployment of the service server.
[0193] Based on the number of packets sent and lost at each granularity, the packet loss rate at the corresponding granularity is determined.
[0194] Detect whether the packet loss rate at each granularity is greater than the preset packet loss ratio at the corresponding granularity.
[0195] Further optionally, in one embodiment of this disclosure, the detection result further includes: 5-tuple information of the probe packets lost within the current time period, the 5-tuple information including: the source Internet Protocol address of the probe, the target Internet Protocol address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the probe used in the probe packet; the source Internet Protocol address of the probe, the target Internet Protocol address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the probe used in the probe packet; further, the analysis platform 1000 also includes:
[0196] The fault location module 1004 is used to locate faults based on a pre-created network device topology map and the five-tuple information of probe packets lost in the current time period.
[0197] Further, optionally, in one embodiment of this disclosure, the analysis platform 1000 further includes:
[0198] The detection module 1005 is used to detect information about network nodes included in the communication path between the detector and each business server in the data center.
[0199] The construction module 1006 is used to construct the network device topology map based on the information of the network nodes in each of the communication paths.
[0200] Further optionally, in one embodiment of this disclosure, the fault location module 1004 is used for:
[0201] Based on the five-tuple information of the probe packets lost within the current time period, the target network node that multiple lost probe packets within the target granularity coverage area are obtained from the network device topology graph.
[0202] The location of the fault is determined based on the packet loss rate of the target network node within the current time period.
[0203] Further optionally, in one embodiment of this disclosure, the fault location module 1004 is used for:
[0204] If the packet loss rate of all ports in the target network node is greater than a preset ratio within the current time period, the target network node is determined to be faulty.
[0205] Further optionally, in one embodiment of this disclosure, the fault location module 1004 is used for:
[0206] If the packet loss rate of only the specified port of the target network node is greater than the preset ratio within the current time period, the fault location is determined based on the position of the target network node in the network architecture.
[0207] Further optionally, in one embodiment of this disclosure, the fault location module 1004 is used for:
[0208] If the target network node is the lowest-level network device above the service server in the network architecture, determine that the specified port of the target network node is faulty.
[0209] Further optionally, in one embodiment of this disclosure, the fault location module 1004 is used for:
[0210] If the target network node is not the lowest-level network device above the service server in the network architecture, obtain information about the next-level network node connected through the specified port from the network device topology diagram.
[0211] The location of the fault is determined based on the packet loss rate of the next-layer network node within the current time period.
[0212] The analysis platform 1000 in this embodiment achieves the same implementation principle and technical effect of anomaly detection by using the above-mentioned modules as the related method embodiments described above. For details, please refer to the description of the related method embodiments described above, which will not be repeated here.
[0213] Figure 11 This is a schematic diagram according to the seventh embodiment of the present disclosure; as shown Figure 7 As shown, this embodiment provides a detector 1100, including:
[0214] Module 1101 is used to construct a probe packet for the prober to probe the business servers of the data center within the current time period.
[0215] The detection module 1102 is used to actively probe the service server using the detection packet within the current time period based on the handshake process of the transmission control protocol.
[0216] The acquisition module 1103 is used to acquire the detection results of the probe machine probing the service server within the current time period;
[0217] The sending module 1104 is used to send the detection results corresponding to the business server within the current time period to the analysis platform, so that the analysis platform can perform anomaly detection at least one granularity within the data center based on the detection results of the probe on the business server within the current time period; and issue an alarm message when an anomaly at the target granularity is detected.
[0218] The detector 1100 in this embodiment achieves the same implementation principle and technical effect of anomaly detection by using the above-mentioned modules as the related method embodiments described above. For details, please refer to the description of the related method embodiments described above, which will not be repeated here.
[0219] Further optionally, in one embodiment of this disclosure, the construction module 1101 is configured to:
[0220] Obtain the target Internet Protocol address and target port identifier of the business server, as well as the source Internet Protocol address of the probe, from the control center equipment;
[0221] According to preset rules, the source port identifier of the probe used by each probe packet to be constructed within the current time period is configured sequentially;
[0222] Based on the source Internet Protocol address of the probe, the target Internet Protocol address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the probe used in each probe packet, each probe packet for the probe to probe the service server within the current time period is generated sequentially.
[0223] Further optionally, in one embodiment of this disclosure, the detection module 1102 is used for:
[0224] Based on the handshake process of the transmission control protocol, each probe packet is sent to the service server in sequence within the current time period; and the response result of each probe packet is obtained in sequence.
[0225] Further optionally, in one embodiment of this disclosure, the acquisition module 1103 is used for:
[0226] For each of the aforementioned probe packets, obtain the response message returned by the service server based on the probe packet; the response message includes a service response or a response terminating the abnormal connection; or
[0227] If no response message is received from the service server based on the probe packet within a preset time period for each probe packet, the probe packet is determined to be lost.
[0228] Further optionally, in one embodiment of this disclosure, the acquisition module 1103 is used for:
[0229] Based on the response results of each probe packet within the current time period, the number of packets sent and the number of packets lost in the current time period are aggregated according to the source Internet Protocol address of the probe and the target Internet Protocol address of the service server.
[0230] It is also used to obtain the quintuple information corresponding to the lost probe packet.
[0231] The detection machine 1100 in the above embodiment achieves the same implementation principle and technical effect of anomaly detection by using the above modules as the related method embodiments. For details, please refer to the description of the related method embodiments, which will not be repeated here.
[0232] Figure 12 This is a schematic diagram based on the eighth embodiment of the present disclosure; as shown Figure 8 As shown, this embodiment provides an anomaly detection system 1200, including: at least one detector 1201, a control center device 1202, and an analysis platform 1203. At least one detector 1201 can be set up inside or outside the data center; if there are two or more detectors 1201, detectors can be set up both inside and outside the data center at the same time.
[0233] Each detector 1201 is connected to the control center equipment 1202 and the analysis platform 1203 respectively; each detector can use Figure 11 The detector and analysis platform 1203 described in the illustrated embodiment can employ the above-mentioned... Figure 9 or Figure 10 The analysis platform described in the illustrated embodiment; further, the above-mentioned... Figures 3-4 ,as well as Figures 6-7 The anomaly detection method of the embodiment shown performs anomaly detection. For details, please refer to the description of the above-mentioned related embodiments, which will not be repeated here.
[0234] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0235] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0236] Figure 13 A schematic block diagram of an example electronic device 1300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0237] like Figure 13As shown, device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1302 or a computer program loaded from storage unit 1308 into random access memory (RAM) 1303. The RAM 1303 may also store various programs and data required for the operation of device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. Input / output (I / O) interface 1305 is also connected to bus 1304.
[0238] Multiple components in device 1300 are connected to I / O interface 1305, including: input unit 1306, such as keyboard, mouse, etc.; output unit 1307, such as various types of monitors, speakers, etc.; storage unit 1308, such as disk, optical disk, etc.; and communication unit 1309, such as network card, modem, wireless transceiver, etc. Communication unit 1309 allows device 1300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0239] The computing unit 1301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the methods of this disclosure described above can be performed. Alternatively, in other embodiments, the computing unit 1301 may be configured to perform the methods described above in this disclosure by any other suitable means (e.g., by means of firmware).
[0240] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0241] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0242] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0243] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0244] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0245] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0246] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0247] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An anomaly detection method, comprising: Obtain the detection results of the probe machine probing the business servers in the data center within the current time period; When the probe probes the service server, it actively probes based on the handshake process of the transmission control protocol. Based on the detection results of the probe on the business server within the current time period, at least one granularity of anomaly detection is performed within the data center; An alarm message is issued in response to the detection of an anomaly in the target granularity.
2. The method according to claim 1, wherein, Obtain the detection results of the probe machine's probes of the data center's business servers within the current time period, including: Obtain the number of packets sent and lost by the probe machine during the current time period.
3. The method according to claim 2, wherein, Based on the detection results of the probe machine on the business server within the current time period, at least one granularity of anomaly detection is performed within the data center, including: Based on the number of packets sent and the number of packets lost obtained by the probe machine from probing the service server within the current time period, the number of packets sent and lost at least one granularity in the data center, cluster, and network device is obtained; the network device is a device located above the service server in the network architecture and used to assist in the deployment of the service server. Based on the number of packets sent and lost at each granularity, the packet loss rate at the corresponding granularity is determined. Detect whether the packet loss rate at each granularity is greater than the preset packet loss ratio at the corresponding granularity.
4. The method according to claim 2 or 3, wherein, The detection results also include: 5-tuple information of probe packets lost within the current time period, the 5-tuple information including: the source Internet Protocol address of the probe, the target Internet Protocol address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the probe used in the probe packet; in response to detecting an abnormal target granularity, the method further includes: Fault location is performed based on a pre-created network device topology map and the 5-tuple information of probe packets lost in the current time period.
5. The method according to claim 4, wherein, Before fault location is performed based on a pre-created network device topology map and the five-tuple information of probe packets lost in the current time period, the method includes: The probe detects information about the network nodes included in the communication paths between the probe and each business server in the data center. Based on the information of the network nodes in each of the communication paths, the network device topology diagram is constructed.
6. The method according to claim 4, wherein, Based on a pre-created network device topology map and the five-tuple information of probe packets lost in the current time period, fault location is performed, including: Based on the five-tuple information of the probe packets lost within the current time period, the target network node that multiple lost probe packets within the target granularity coverage area are obtained from the network device topology graph. The location of the fault is determined based on the packet loss rate of the target network node within the current time period.
7. The method according to claim 6, wherein, Based on the packet loss rate of the target network node within the current time period, the fault location is determined, including: If the packet loss rate of all ports in the target network node is greater than a preset ratio within the current time period, the target network node is determined to be faulty.
8. The method according to claim 6, wherein, Determining the fault location based on the packet loss rate of the target network node within the current time period also includes: If the packet loss rate of only the specified port of the target network node is greater than the preset ratio within the current time period, the fault location is determined based on the position of the target network node in the network architecture.
9. The method according to claim 8, wherein, Based on the location of the target network node in the network architecture, the fault location is determined, including: If the target network node is the lowest-level network device above the service server in the network architecture, determine that the specified port of the target network node is faulty.
10. The method according to claim 8, wherein, Based on the location of the target network node in the network architecture, the fault location is determined, including: If the target network node is not the lowest-level network device above the service server in the network architecture, obtain information about the next-level network node connected through the specified port from the network device topology diagram. The location of the fault is determined based on the packet loss rate of the next-layer network node within the current time period.
11. An anomaly detection method, comprising: Construct a probe packet for the prober to probe the business servers in the data center within the current time period; Based on the handshake process of the Transmission Control Protocol, the probe packet is used to actively probe the service server within the current time period; Obtain the detection results of the probe machine's probe on the service server within the current time period; The system sends the detection results corresponding to the business server within the current time period to the analysis platform, so that the analysis platform can perform anomaly detection at least at one granularity within the data center based on the detection results of the probe on the business server within the current time period; and issues an alarm message when an anomaly at the target granularity is detected.
12. The method according to claim 11, wherein, Construct a probe packet for the prober to probe the business servers in the data center within the current time period, including: Obtain the target Internet Protocol address and target port identifier of the business server, as well as the source Internet Protocol address of the probe, from the control center equipment; According to preset rules, the source port identifier of the probe used by each probe packet to be constructed within the current time period is configured sequentially; Based on the source Internet Protocol address of the probe, the target Internet Protocol address and target port identifier of the service server, the determined communication protocol type, and the source port identifier of the probe used in each probe packet, each probe packet for the probe to probe the service server within the current time period is generated sequentially.
13. The method according to claim 12, wherein, Based on the handshake process of the Transmission Control Protocol, the probe packet is used to actively probe the service server within the current time period, including: Based on the handshake process of the transmission control protocol, each probe packet is sent to the service server in sequence within the current time period; and the response result of each probe packet is obtained in sequence.
14. The method according to claim 13, wherein, The response results of each probe packet are obtained sequentially, including: For each of the aforementioned probe packets, obtain the response message returned by the service server based on the probe packet; the response message includes a service response or a response terminating the abnormal connection; or If no response message is received from the service server based on the probe packet within a preset time period for each probe packet, the probe packet is determined to be lost.
15. The method according to claim 14, wherein, Obtain the detection results of the probe machine's probe against the service server within the current time period, including: Based on the response results of each probe packet within the current time period, the number of packets sent and the number of packets lost in the current time period are aggregated according to the source Internet Protocol address of the probe and the target Internet Protocol address of the service server. It also includes: obtaining the quintuple information corresponding to the lost probe packet.
16. An analysis platform, comprising: The acquisition module is used to acquire the detection results of the probe machine probing the business servers of the data center within the current time period; When the probe probes the service server, it actively probes based on the handshake process of the transmission control protocol. An anomaly detection module is used to perform anomaly detection at least one granularity within the data center based on the detection results of the probe machine on the business server within the current time period. The alarm module is used to issue alarm information in response to the detection of anomalies at the target granularity.
17. A detector, comprising: The building module is used to construct a probe package for the prober to probe the business servers in the data center within the current time period; The detection module is used to actively probe the service server using the detection packet within the current time period, based on the handshake process of the transmission control protocol. The acquisition module is used to acquire the detection results of the probe machine probing the business server within the current time period; The sending module is used to send the detection results corresponding to the business server within the current time period to the analysis platform, so that the analysis platform can perform anomaly detection at least one granularity within the data center based on the detection results of the probe on the business server within the current time period; and issue an alarm message when an anomaly at the target granularity is detected.
18. An anomaly detection system, comprising: At least one detector, control center equipment, and analysis platform; Each of the aforementioned detectors is connected to the control center equipment and the analysis platform respectively; each of the aforementioned detectors is the detector described in claim 17; the analysis platform is the analysis platform described in claim 15.
19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.
21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-15.