Node anomaly detection method and device, equipment and medium

By acquiring the time window running data of nodes in a clustered system, calculating the mapping relationship coefficient, and combining it with the risk threshold for anomaly detection, the problem of insufficient accuracy in node anomaly detection in existing technologies is solved, achieving more efficient anomaly state identification and system stability assurance.

CN121690974APending Publication Date: 2026-03-17INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511931701.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies have poor accuracy in detecting node anomalies and cannot effectively guarantee the reliable operation of clustered systems.

Method used

By acquiring the operational data of each node in the clustered system at each time window, calculating the mapping relationship coefficient using historical and current operational data, and combining it with risk thresholds to perform anomaly detection, the abnormal state of the nodes is determined.

Benefits of technology

It improves the accuracy of node anomaly detection, enabling more accurate identification of abnormal node states and ensuring the stability and reliability of clustered systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121690974A_ABST
    Figure CN121690974A_ABST
Patent Text Reader

Abstract

The invention discloses a node anomaly detection method and device, equipment and a medium. The method comprises the following steps: acquiring operation data of nodes in the cluster system in each time window; according to the operation data of each time window, respectively determining a risk threshold value of each time window; calculating a mapping relation coefficient of the time window according to the operation data from the initial time window to the time window and the operation data of the time window; determining a coefficient range according to the mapping relation coefficient of each historical time window; the historical time window is located before the current time window; and according to the coefficient range and the mapping relation coefficient of the current time window, performing anomaly detection on the operation state of the node in the current time window to obtain an anomaly detection result. According to the embodiment of the invention, the accuracy of the abnormal detection result of the node can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of anomaly detection technology, and in particular to a method, apparatus, equipment and medium for node anomaly detection. Background Technology

[0002] With the development of technologies such as cloud computing and distributed systems, more and more business systems are deployed in multi-node environments as clustered systems. The overall availability and stability of the system largely depend on the operating status of each node. To ensure business continuity, operations and maintenance personnel typically need to continuously monitor the operating parameters of each node and assess whether the node is in an abnormal state, thereby enabling timely fault location and resource scheduling. Therefore, how to accurately and timely detect the operating status of each node in a clustered system has become a fundamental technology for ensuring the reliable operation of clustered systems.

[0003] Currently, existing technologies obtain anomaly detection results for nodes by setting fixed thresholds and comparing them with the node's operating parameters.

[0004] However, existing technologies have poor accuracy in detecting anomalies in nodes. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for detecting node anomalies. The embodiments of this invention can improve the accuracy of node anomaly detection results.

[0006] In a first aspect, embodiments of the present invention provide a node anomaly detection method, the method comprising:

[0007] Obtain the running data of nodes in a clustered system in each time window;

[0008] Based on the operational data for each time window, the risk threshold for each time window is determined.

[0009] Based on the running data from the initial time window to this time window, and the running data of this time window, calculate the mapping relationship coefficient of this time window;

[0010] The range of coefficients is determined based on the mapping relationship coefficients of each historical time window; the historical time window is located before the current time window.

[0011] Based on the coefficient range and the mapping relationship coefficients to the current time window, anomaly detection is performed on the running status of the node within the current time window, and the anomaly detection results are obtained.

[0012] Secondly, embodiments of the present invention also provide a node anomaly detection device, the device comprising:

[0013] The data acquisition module is used to acquire the running data of nodes in the clustered system in each time window;

[0014] The risk threshold determination module is used to determine the risk threshold for each time window based on the operational data of each time window.

[0015] The mapping relationship coefficient calculation module is used to calculate the mapping relationship coefficient of the time window based on the running data from the initial time window to the current time window, and the running data of the current time window.

[0016] The coefficient range determination module is used to determine the coefficient range based on the mapping relationship coefficients of each historical time window; the historical time window is located before the current time window.

[0017] The anomaly detection module is used to detect anomalies in the running status of nodes within the current time window based on the coefficient range and the mapping relationship coefficients to the current time window, and obtain the anomaly detection results.

[0018] Thirdly, embodiments of the present invention also provide a node anomaly detection device, the node anomaly detection device comprising:

[0019] At least one processor; and

[0020] A memory that is communicatively connected to at least one processor; wherein,

[0021] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the node anomaly detection method of any embodiment of the present invention.

[0022] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions, which are used to cause a processor to execute the node anomaly detection method of any embodiment of the present invention.

[0023] The technical solution of this invention, by acquiring the operational data of each node in a clustered system for each time window, can simultaneously utilize historical and current operational data during anomaly detection, providing a more complete time-dimensional reference for subsequent risk assessment. By combining the risk threshold and operational data from multiple time periods to calculate the mapping relationship coefficient, the relationship between operational data and the risk threshold can be quantified into a unified coefficient index. By detecting whether the mapping relationship coefficient of the current time window falls within the coefficient range, the degree of deviation between the current time period and the historical normal range can be used to determine whether a node is abnormal, thereby achieving node anomaly detection based on a comprehensive evaluation of historical and current data. This solves the technical problem of poor accuracy in node anomaly detection in existing technologies, thereby improving the accuracy of node anomaly detection results.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart of a node anomaly detection method provided in an embodiment of the present invention;

[0027] Figure 2 A flowchart of a node anomaly detection method provided in an embodiment of the present invention;

[0028] Figure 3 This is a schematic diagram of the structure of a node anomaly detection device provided in an embodiment of the present invention;

[0029] Figure 4 This is a schematic diagram of the structure of a node anomaly detection device provided in an embodiment of the present invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] In the technical solutions of the embodiments of the present invention, the acquisition, storage and application of operational data, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0033] Figure 1 This is a flowchart illustrating a node anomaly detection method provided in an embodiment of the present invention. This embodiment is applicable to situations requiring the detection of node operational status. The method can be executed by a node anomaly detection device, which can be implemented in hardware and / or software.

[0034] See Figure 1 The node anomaly detection method shown includes:

[0035] S101. Obtain the running data of nodes in the cluster system in each time window.

[0036] In this context, a clustered system refers to a computing system as a whole, consisting of multiple independently operable computing nodes working collaboratively. A clustered system typically includes multiple nodes, such as multiple servers, virtual machines, or container instances. For example, a clustered system can be a distributed service cluster deployed in a data center, or it can be a distributed storage system, a distributed database cluster, or a microservice cluster.

[0037] In this context, a node refers to the smallest computing unit in a clustered system that possesses independent computing capabilities and can be deployed and run applications independently. Nodes are the direct targets of anomaly detection methods, and anomaly detection results are generated for each node, thus supporting the identification and location of specific abnormal nodes within the cluster. For example, a single physical server can be a node, as can a single virtual machine or container instance.

[0038] In this context, operational data can refer to operational parameters that characterize the operational status of a node within a certain time period. For example, the operational data of a node within a certain time period could be average input / output latency, processor utilization, or memory usage.

[0039] S102. Based on the operational data of each time window, determine the risk threshold for each time window.

[0040] The risk threshold can refer to a pre-set or determined threshold for operational data within a specific time period of a node. The risk threshold serves as a boundary value to distinguish between normal operating conditions and potential abnormal conditions. By defining a risk boundary for operational data, the risk threshold indicates a high risk of anomalies during that time period when the operational data approaches or exceeds this boundary.

[0041] S103. Based on the running data from the initial time window to this time window, and the running data of this time window, calculate the mapping relationship coefficient of this time window.

[0042] Optionally, the mapping coefficient for the time window can be calculated by taking the average running data from the initial time window to that time window, as well as the running data within that time window. The average running data can refer to the average value of the running parameters of a certain running data over multiple time periods. By averaging the running data over multiple time periods, the impact of occasional fluctuations in a single time period can be eliminated, resulting in more stable running characteristics.

[0043] The mapping relationship coefficient can be a coefficient calculated based on the risk threshold of the time window and the operational data from the initial time window to that time window. The mapping relationship coefficient is used to quantitatively characterize the mapping relationship and deviation of the operational mean relative to the risk threshold. As an intermediate quantity connecting the operational mean and the risk threshold, the mapping relationship coefficient normalizes complex multi-time-period, multi-parameter relationships into a coefficient with a unified scale. For example, the mapping relationship coefficient can be the ratio of the operational mean to the risk threshold, such as the result of operational mean / risk threshold; or it can be the coefficient value of the operational mean and the risk threshold after being transformed by a linear or nonlinear function, where the specific function form can be configured according to different indicators and different systems.

[0044] S104. Determine the coefficient range based on the mapping relationship coefficients of each historical time window; the historical time window is located before the current time window.

[0045] The coefficient range refers to the range of values ​​determined based on the mapping relationship coefficients for each historical time window. The coefficient range characterizes the fluctuation range of the mapping relationship coefficients under normal operating conditions. It allows us to obtain the variation range of these coefficients under normal operating conditions, reflecting the reasonable range of the mapping relationship coefficients under normal operating conditions.

[0046] S105. Based on the coefficient range and the mapping relationship coefficient to the current time window, perform anomaly detection on the running status of the node within the current time window and obtain the anomaly detection result.

[0047] This involves performing anomaly detection on the node's operational status within the current time window. The anomaly detection result can refer to the detection conclusion given for the operational status of a specific node in a clustered system during the current time period. The anomaly detection result is used to indicate whether the node is abnormal during the current time period.

[0048] As can be seen, in this embodiment, by acquiring the operational data of each node in the clustered system for each time window, both historical and current operational data can be used simultaneously during anomaly detection, providing a more complete time-dimensional reference for subsequent risk assessment. By combining the risk threshold and operational data from multiple time periods to calculate the mapping coefficient, the relationship between operational data and the risk threshold can be quantified into a unified coefficient index. By detecting whether the mapping coefficient of the current time window falls within the coefficient range, the degree of deviation between the current time period and the historical normal range can be used to determine whether a node is abnormal, thereby achieving node anomaly detection based on a comprehensive evaluation of historical and current data. This solves the technical problem of poor accuracy in node anomaly detection in existing technologies, thereby improving the accuracy of node anomaly detection results.

[0049] In an optional embodiment, Figure 2 The flowchart of a node anomaly detection method provided in this embodiment of the invention refines the step of "calculating the mapping relationship coefficient of the time window based on the running data from the initial time window to the time window and the running data of the time window" into "obtaining a preset mapping relationship; substituting the risk threshold of the time window and the running data from the initial time window to the time window into the mapping relationship, and calculating the mapping relationship coefficient of the time window", so as to improve the node anomaly detection operation.

[0050] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.

[0051] See Figure 2 The node anomaly detection method shown includes:

[0052] S201. Obtain the running data of nodes in the cluster system in each time window.

[0053] S202. Based on the operational data of each time window, determine the risk threshold for each time window.

[0054] S203. Obtain the preset mapping relationship.

[0055] The preset mapping relationship can refer to a pre-determined functional relationship. This preset mapping relationship is used to map the risk threshold of the current time window and the average operating values ​​of all previous time windows to the corresponding mapping coefficients for that time period. The preset mapping relationship serves as the basis for calculating the mapping coefficients for each time period, ensuring a clear constraint relationship between the mapping coefficients and the risk threshold and the average operating values.

[0056] S204. Substitute the risk threshold of the time window and the running data from the initial time window to the time window into the mapping relationship, and calculate the mapping relationship coefficient of the time window.

[0057] S205. Determine the coefficient range based on the mapping relationship coefficients of each historical time window; the historical time window is located before the current time window.

[0058] S206. Based on the coefficient range and the mapping relationship coefficients to the current time window, perform anomaly detection on the running status of the node within the current time window and obtain the anomaly detection results.

[0059] As can be seen, in this embodiment, by obtaining the preset mapping relationship, a unified calculation rule can be adopted when calculating the mapping relationship coefficients of each time period, which is conducive to ensuring that the coefficients of different time periods are comparable and consistent. By substituting the risk threshold of the time window and the running data from the initial time window to the time window into the mapping relationship, the mapping relationship coefficient of the time window can be calculated, and the relationship between the risk threshold and the running average can be quantified into the mapping relationship coefficient of the time period.

[0060] In some embodiments, the risk threshold of the time window and the running data from the initial time window to the time window are substituted into the mapping relationship to calculate the mapping relationship coefficient of the time window, including:

[0061] The running average is calculated based on the running data from the initial time window to that time window;

[0062] Substitute the risk threshold and operating average of the time window into the mapping relationship to calculate the mapping relationship coefficient for the time window.

[0063] As can be seen, in this embodiment, by calculating the average operating value based on the operating data from the initial time window to the current time window, the operating conditions of multiple historical time windows can be comprehensively considered in the coefficient calculation of the current time window. This helps to reduce the impact of occasional fluctuations in a single time window on the mapping relationship coefficient. By substituting the risk threshold and the average operating value of the time window into the mapping relationship to calculate the mapping relationship coefficient of the time window, the relationship between the operating level and the risk boundary can be transformed into the coefficient quantification result of the time window. This is beneficial for subsequent anomaly detection of the node operating status corresponding to the time window based on the coefficient range.

[0064] In some embodiments, the risk threshold for each time window is determined based on the operational data for each time window, including:

[0065] Based on the operational data of the time window and the preset risk threshold relationship, the risk threshold for each time window is calculated.

[0066] The preset risk threshold is a pre-determined functional relationship between operational data and risk thresholds. It is used to calculate the risk threshold for the i-th time period based on the operational data for that time period, thus establishing a calculable mapping relationship between the risk threshold and the operational data.

[0067] As can be seen, in this embodiment, the risk threshold for the i-th time period is calculated based on the relationship between the running data of the i-th time period and the preset risk threshold. The preset relationship can be used to quantify and map the running data, which is beneficial to automatically obtain the risk threshold for each time period with a unified rule.

[0068] In some embodiments, the anomaly detection result includes: normal or abnormal;

[0069] Based on the coefficient range and the mapping coefficients to the current time window, anomaly detection is performed on the node's operating status within the current time window, yielding anomaly detection results, including:

[0070] When the mapping coefficient of the current time window is within the coefficient range, or when the mapping coefficient of the current time window is equal to the two boundary values ​​of the coefficient range, the abnormal detection result of the node in the current time window is determined to be normal.

[0071] When the mapping coefficient of the current time window is outside the coefficient range, the abnormal detection result of the node within the current time window is determined to be abnormal.

[0072] As can be seen, in this embodiment, by determining the abnormal detection result of the node as abnormal when the mapping relationship coefficient of the current time window falls outside the coefficient range, the node can be marked as abnormal when the current coefficient deviates significantly from the historical normal range, thereby realizing automatic detection of abnormal node status based on the degree of coefficient deviation.

[0073] In some embodiments, after performing anomaly detection on the node's running state within the current time window and obtaining the anomaly detection result, the method further includes:

[0074] Based on the anomaly detection results of the operating status of each node within the current time window, and the degree of system impact of each node, anomaly detection is performed on the operating status of the clustered system within the current time window to obtain system anomaly detection results.

[0075] The degree of system impact can refer to a weighted or quantitative indicator used to characterize the impact of a node anomaly on the overall operation of the entire clustered system. The degree of system impact is used for weighted aggregation of anomaly detection results across multiple nodes. The degree of system impact is typically a numerical weight, such as a weight coefficient between 0 and 1.

[0076] As can be seen, in this embodiment, when the mapping coefficient of the current time window is within the coefficient range or equal to the boundary value of the coefficient range, the abnormal detection result of the node is determined to be normal. The entire coefficient range obtained from historical statistics can be taken as the normal interval. When the mapping coefficient of the current time window is outside the coefficient range, the abnormal detection result of the node is determined to be abnormal. When the current coefficient deviates significantly from the historical normal range, the node can be marked as abnormal, thereby realizing abnormal detection based on the degree of coefficient deviation.

[0077] In some embodiments, runtime data includes: input / output latency, processor utilization, memory utilization, packet loss rate, or packet error rate.

[0078] Input / output latency refers to the time delay experienced by a node from initiating a request to completing a response when performing input / output operations (such as disk read / write and network requests) within a certain period of time. It is used to characterize the performance status of a node at the input and output stream levels.

[0079] Processor utilization refers to the proportion of a node's processors that are occupied within a certain period of time, reflecting the computing load of the node during that period.

[0080] Among them, memory utilization rate can refer to the proportion of memory resources occupied by a node to the total available memory resources within a certain period of time, which is used to characterize the memory load of the node during that period of time.

[0081] Packet loss rate refers to the proportion of data packets lost during network data packet transmission within a certain period of time to the total number of data packets sent or to be received, and is used to reflect the status of the node in terms of network transmission reliability.

[0082] Among them, the error rate can refer to the proportion of data packets with checksum errors, format errors, or content errors received by a node within a certain period of time, out of the total number of received data packets. It is used to characterize the node's status in terms of network data correctness.

[0083] As can be seen, in this embodiment, by incorporating input / output latency, processor utilization, memory utilization, packet loss rate, or packet error rate as part of the runtime data, the node's runtime status can be characterized from multiple dimensions such as performance, resource consumption, and network transmission quality. This is beneficial for reflecting the actual runtime status of the node more comprehensively during subsequent anomaly detection.

[0084] In an optional embodiment, in a cloud platform environment, detecting abnormal states of nodes in the cloud platform can be achieved through the following steps:

[0085] First, based on the actual operating status of the cloud platform, the alarm status of each indicator, and the proportion of the indicator's impact on the business, the important indicators affecting the stable operation of the cloud platform are determined. These important indicators can be further divided into the following categories: such as input stream latency and output stream latency, processor utilization, memory usage, packet loss rate, and packet error rate.

[0086] Taking a key metric as an example, the user attention threshold for this metric is β. If there are no relevant alarms on the current cloud platform, the peak values ​​of this metric for the first day for all nodes are obtained by calling the cloud platform interface {A1, B1, C1, D1...}, where A, B, C, and D represent nodes, and 1 represents the first day. Taking node A as an example, the default correspondence between A1 and β is A1=f(β). The relationship model between Ai, β, and the initial value is determined by analyzing Ai, β, and the initial values: Ai=f(a, b, β).

[0087] As the number of days increases, obtain the average value of the set {A1...An} of Ai for the previous n days at that node, and determine the average values ​​of a and b. Subsequently, obtain this indicator for each node at regular intervals every day, and calculate the values ​​of a and b. When a is between the maximum and minimum values ​​of ai at that node, and b is between the maximum and minimum values ​​of bi at that node, then the indicator for that node is considered normal.

[0088] Figure 3 This invention provides a schematic diagram of a node anomaly detection device. This invention is applicable to situations requiring the detection of node operational status. The device can execute node anomaly detection methods and can be implemented in hardware and / or software.

[0089] See Figure 3 The node anomaly detection device shown includes: a data acquisition module 301, a risk threshold determination module 302, a mapping relationship coefficient calculation module 303, a coefficient range determination module 304, and an anomaly detection module 305, wherein...

[0090] The data acquisition module 301 is used to acquire the running data of nodes in the cluster system in each time window;

[0091] The risk threshold determination module 302 is used to determine the risk threshold for each time window based on the running data of each time window.

[0092] The mapping relationship coefficient calculation module 303 is used to calculate the mapping relationship coefficient of the time window based on the running data from the initial time window to the time window and the running data of the time window.

[0093] The coefficient range determination module 304 is used to determine the coefficient range based on the mapping relationship coefficients of each historical time window; the historical time window is located before the current time window.

[0094] The anomaly detection module 305 is used to perform anomaly detection on the running status of nodes within the current time window based on the coefficient range and the mapping relationship coefficient to the current time window, and obtain the anomaly detection result.

[0095] The technical solution of this invention, by acquiring the operational data of each node in a clustered system for each time window, can simultaneously utilize historical and current operational data during anomaly detection, providing a more complete time-dimensional reference for subsequent risk assessment. By combining the risk threshold and operational data from multiple time periods to calculate the mapping relationship coefficient, the relationship between operational data and the risk threshold can be quantified into a unified coefficient index. By detecting whether the mapping relationship coefficient of the current time window falls within the coefficient range, the degree of deviation between the current time period and the historical normal range can be used to determine whether a node is abnormal, thereby achieving node anomaly detection based on a comprehensive evaluation of historical and current data. This solves the technical problem of poor accuracy in node anomaly detection in existing technologies, thereby improving the accuracy of node anomaly detection results.

[0096] In some embodiments, in calculating the mapping coefficient of a time window based on the running data from the initial time window to that time window, and the running data of that time window, the mapping coefficient calculation module 303 is specifically used for:

[0097] Obtain the preset mapping relationship;

[0098] Substitute the risk threshold of the time window and the running data from the initial time window to the time window into the mapping relationship, and calculate the mapping relationship coefficient of the time window.

[0099] In some embodiments, in calculating the mapping coefficient of the time window by substituting the risk threshold of the time window and the running data from the initial time window to the time window into the mapping relationship, the mapping coefficient calculation module 303 is specifically used for:

[0100] The running average is calculated based on the running data from the initial time window to that time window;

[0101] Substitute the risk threshold and operating average of the time window into the mapping relationship to calculate the mapping relationship coefficient for the time window.

[0102] In some embodiments, in determining the risk threshold for each time window based on the operational data of each time window, the risk threshold determination module 302 is specifically used for:

[0103] Based on the operational data of the time window and the preset risk threshold relationship, the risk threshold for each time window is calculated.

[0104] In some embodiments, the anomaly detection result includes: normal or abnormal;

[0105] In terms of detecting anomalies in the node's operating status within the current time window based on the coefficient range and the mapping coefficients to the current time window, and obtaining the anomaly detection results, the anomaly detection module 305 is specifically used for:

[0106] When the mapping coefficient of the current time window is within the coefficient range, or when the mapping coefficient of the current time window is equal to the two boundary values ​​of the coefficient range, the abnormal detection result of the node in the current time window is determined to be normal.

[0107] When the mapping coefficient of the current time window is outside the coefficient range, the abnormal detection result of the node within the current time window is determined to be abnormal.

[0108] In some embodiments, the node anomaly detection device further includes, in order to perform anomaly detection on the running state of a node within the current time window and obtain anomaly detection results:

[0109] The system anomaly detection result determination module is used to perform anomaly detection on the operating status of the clustered system within the current time window based on the anomaly detection results of the operating status of each node within the current time window and the degree of system impact of each node, and obtain the system anomaly detection result.

[0110] In some embodiments, runtime data includes: input / output latency, processor utilization, memory utilization, packet loss rate, or packet error rate.

[0111] The node anomaly detection device provided in this embodiment of the invention can execute the node anomaly detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of executing the node anomaly detection method.

[0112] Figure 4 This is a schematic diagram of the structure of a node anomaly detection device provided in an embodiment of the present invention.

[0113] like Figure 4 As shown, the node anomaly detection device 400 includes at least one processor 401 and a memory, such as a read-only memory (ROM) 402 or a random access memory (RAM) 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the ROM 402 or loaded into the RAM 403 from storage unit 408. The RAM 403 can also store various programs and data required for the operation of the node anomaly detection device 400. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 408 is also connected to the bus 404.

[0114] Multiple components in the node anomaly detection device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless transceiver, etc. The communication unit 409 allows the node anomaly detection device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0115] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as node anomaly detection methods.

[0116] In some embodiments, the node anomaly detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the node anomaly detection device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by processor 401, one or more steps of the node anomaly detection method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to execute the node anomaly detection method by any other suitable means (e.g., by means of firmware).

[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0120] To provide interaction with the user, the systems and techniques described herein can be implemented on an operational detection device, which includes: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the node anomaly detection device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0122] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0123] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0124] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method of node anomaly detection, the method comprising: The method comprises: obtaining running data of nodes in a cluster system in each time window; determining a risk threshold of each time window according to the running data of each time window respectively; calculating a mapping relationship coefficient of the time window according to the running data from an initial time window to the time window and the running data of the time window; determining a coefficient range according to the mapping relationship coefficients of historical time windows; the historical time windows are located before the current time window; performing abnormality detection on the running state of the nodes in the current time window according to the coefficient range and the mapping relationship coefficient of the current time window, to obtain an abnormality detection result.

2. The method of claim 1, wherein, The calculation of the mapping relationship coefficient of the time window according to the running data from the initial time window to the time window and the running data of the time window comprises: obtaining a preset mapping relationship; substituting the risk threshold of the time window and the running data from the initial time window to the time window into the mapping relationship to calculate the mapping relationship coefficient of the time window.

3. The method of claim 2, wherein, The calculation of the mapping relationship coefficient of the time window according to the running data from the initial time window to the time window and the running data of the time window comprises: calculating a running mean value according to the running data from the initial time window to the time window; substituting the risk threshold of the time window and the running mean value into the mapping relationship to calculate the mapping relationship coefficient of the time window.

4. The method of claim 1, wherein, The determination of the risk threshold of each time window according to the running data of each time window comprises: calculating the risk threshold of each time window according to the running data of the time window and a preset risk threshold relationship.

5. The method of claim 1, wherein, The abnormality detection result comprises: normal or abnormal. The abnormality detection on the running state of the nodes in the current time window according to the coefficient range and the mapping relationship coefficient of the current time window to obtain the abnormality detection result comprises: when the mapping relationship coefficient of the current time window is within the coefficient range, or the mapping relationship coefficient of the current time window is equal to two boundary values of the coefficient range, determining that the abnormality detection result of the nodes in the current time window is normal; when the mapping relationship coefficient of the current time window is outside the coefficient range, determining that the abnormality detection result of the nodes in the current time window is abnormal.

6. The method of claim 1, wherein, After the abnormality detection on the running state of the nodes in the current time window to obtain the abnormality detection result, the method further comprises: performing abnormality detection on the running state of the cluster system in the current time window according to the abnormality detection results of the running states of the nodes in the current time window and the system influence degrees of the nodes, to obtain a system abnormality detection result.

7. The method of claim 1, wherein, The running data comprises: input / output delay, processor usage, memory usage, packet loss rate or packet error rate.

8. A node anomaly detection apparatus characterized by comprising: The method comprises: a data acquisition module configured to obtain running data of nodes in a cluster system in each time window; a risk threshold determination module configured to determine a risk threshold of each time window according to the running data of each time window respectively; The mapping relationship coefficient calculation module is configured to calculate the mapping relationship coefficient of the time window according to the operation data from an initial time window to the time window and the operation data of the time window; The coefficient range determination module is configured to determine a coefficient range according to the mapping relationship coefficients of each historical time window, the historical time window being located before the current time window; The anomaly detection module is configured to perform anomaly detection on the operation state of the node in the current time window according to the coefficient range and the mapping relationship coefficient of the current time window, and obtain an anomaly detection result.

9. A node anomaly detection device, characterized by, The node anomaly detection device comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the node anomaly detection method in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to enable the processor to implement the node anomaly detection method in any one of claims 1-7 when executed.