Anomaly detection method and apparatus

By distributing indicator data to multiple alarm computing nodes through forwarding nodes, the problem of inaccurate detection results caused by the limited resources of a single alarm computing node is solved, achieving more efficient and accurate anomaly detection.

CN122431995APending Publication Date: 2026-07-21HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-01-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

When there are a large number of computing nodes, the resources of a single alarm computing node are limited, resulting in inaccurate anomaly detection results.

Method used

The forwarding node sends multiple metric data to at least two target alarm computing nodes, enabling them to process the data together, reducing the load on individual alarm computing nodes, and distributing the data according to alarm rules matched by the observed object tags and metric identifiers.

Benefits of technology

The accuracy of anomaly detection results has been improved by rationally allocating resources and reducing the processing load of individual alarm computing nodes, ensuring sufficient resources to process indicator data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431995A_ABST
    Figure CN122431995A_ABST
Patent Text Reader

Abstract

The application discloses an abnormality detection method and device, and relates to the technical field of computers. Compared with forwarding multiple index data to a single alarm computing node, in the application, the forwarding node sends part of the multiple index data to each of at least two target alarm computing nodes, so that the at least two target alarm computing nodes jointly process the multiple index data by using resources supported by the at least two target alarm computing nodes. In this way, the number of index data that needs to be processed by a single alarm computing node is reduced, the inaccuracy of an abnormality detection result caused by limited resources of the single alarm computing node is avoided, sufficient resources are provided for the alarm computing node to process index data, and the accuracy of the abnormality detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an anomaly detection method and apparatus. Background Technology

[0002] When a compute node's operational status is abnormal, one or more of its metrics will also show anomalies. The anomaly detection system can determine the compute node's operational status using forwarding nodes and multiple alarm compute nodes. Specifically, the anomaly detection system sets different alarm rules on different alarm compute nodes. When it's necessary to determine whether a compute node meets a specified alarm rule, the anomaly detection system's forwarding nodes forward the compute node's metric data to the alarm compute node responsible for processing that specific alarm rule. This alarm compute node then determines the compute node's operational status based on its metric data and the specified alarm rule.

[0003] In the above process, the anomaly detection system uses a single alarm computing node to determine whether the node's metric data meets the alarm rules. When there are a large number of computing nodes, the limited resources provided by each node may lead to inaccurate detection results. Summary of the Invention

[0004] This application provides an anomaly detection method and apparatus to solve the problem of inaccurate detection results caused by limited resources of a single alarm computing node.

[0005] Firstly, this application provides an anomaly detection method. This method can be executed by a forwarding node, which is communicatively connected to multiple alarm computing nodes. Each of the multiple alarm computing nodes has at least one alarm rule set. An alarm rule is set on at least two alarm computing nodes. An alarm rule includes: an indicator name, an observation object label, and an alarm policy. The method provided in the first aspect includes: the forwarding node acquiring multiple indicator data. Each indicator data includes: an observation object label, an indicator identifier, and an indicator value. The observation object label is used to indicate the observation object. The multiple indicator data includes multiple observation object labels. The forwarding node determines at least two target alarm computing nodes corresponding to the multiple observation object labels and / or the indicator identifiers of the multiple indicator data. Wherein, the at least two target alarm computing nodes are alarm computing nodes among the multiple alarm computing nodes that have deployed a first alarm rule. The indicator identifier of the first alarm rule matches the indicator identifier of the multiple indicator data. The observation object label of the first alarm rule includes the observation object label of the multiple indicator data. The forwarding node sends the multiple indicator data to the at least two target alarm computing nodes. The system includes at least two target alarm calculation nodes, namely a first target alarm calculation node and a second target alarm calculation node. Multiple indicator data points are included, including first indicator data and second indicator data. The first indicator data has a first observation object label. The second indicator data has a second observation object label. The first indicator data is sent to the first target alarm calculation node. The second indicator data is sent to the second target alarm calculation node.

[0006] Compared to a forwarding node forwarding multiple metric data points to a single alarm computing node, in the first aspect of this application, the forwarding node sends a portion of the multiple metric data points to each of at least two target alarm computing nodes, enabling the at least two target alarm computing nodes to jointly process the multiple metric data points using their supported resources. This reduces the amount of metric data points that a single alarm computing node needs to process, avoids inaccurate anomaly detection results due to resource constraints of a single alarm computing node, ensures sufficient resources for alarm computing nodes to process metric data, and thus improves the accuracy of anomaly detection results.

[0007] The aforementioned forwarding node can be a single computing device, a cluster of computing devices including multiple computing devices, a component of a computing device such as a processor, chip, or chip system, or a logic module or software that can implement all or part of the functions of a computing device.

[0008] In one possible implementation, a forwarding node sends multiple indicator data to at least two target alarm computing nodes. This includes: the forwarding node determining a first observation object label corresponding to a first target alarm computing node and a second observation object label corresponding to a second target alarm computing node based on multiple observation object labels and the node identifiers of the at least two target alarm computing nodes. The forwarding node sends indicator data including the first observation object label to the first target alarm computing node, and sends indicator data including the second observation object label to the second target alarm computing node. In this way, the forwarding node sends indicator data including different observation object labels to different target alarm computing nodes. The target alarm computing nodes process the indicator data including the corresponding observation object label from the multiple indicator data, without needing to process indicator data including other observation object labels. This reduces the amount of indicator data that a single alarm computing node needs to process, ensuring sufficient resources for a single alarm computing node to implement alarm rules, and thus ensuring the accuracy of the anomaly detection results.

[0009] In another possible implementation, the forwarding node determines the first observation object label corresponding to the first target alarm computing node and the second observation object label corresponding to the second target alarm computing node based on multiple observation object labels and the node identifiers of at least two target alarm computing nodes. This includes: the forwarding node generating multiple first hash values ​​based on the multiple observation object labels, with each hash value corresponding to one observation object label; the forwarding node generating at least two second hash values ​​based on the node identifiers of at least two target alarm computing nodes, with each second hash value corresponding to one target alarm computing node; and the forwarding node determining the first observation object label corresponding to the first target alarm computing node and the second observation object label corresponding to the second target alarm computing node based on the multiple first hash values ​​and at least two second hash values. Here, the first observation object label and the second observation object label are multiple observation object labels. Thus, the forwarding node determines the correspondence between the target alarm computing node and the observation object label based on the first hash value of the observation object label and the second hash value of the target alarm computing node. The calculation logic is simple, the complexity is low, and the efficiency of anomaly detection is improved.

[0010] In another possible implementation, the first target alarm calculation node also has a second alarm rule. Multiple indicator data are sent to at least two target alarm calculation nodes, including: if the indicator identifier of the second alarm rule matches the indicator identifier of the first indicator data, and the observation object label of the second alarm rule includes the observation object label of the first indicator data, then the forwarding node sends the first indicator data to the first target alarm calculation node in a specified format. Thus, when multiple alarm rules of the first target alarm calculation node use the same indicator data, the forwarding node sends only one set of indicator data to the first target alarm calculation node. This reduces the bandwidth resources consumed in sending indicator data and improves resource utilization.

[0011] In another possible implementation, the specified format includes a first field and a second field. The first field carries the rule identifiers of the first and second alarm rules, while the second field carries the first metric data. In this way, the forwarding node carries different information in different fields, allowing the alarm calculation node to parse the corresponding fields as needed to obtain the target information, thus improving the efficiency of obtaining the target information.

[0012] In another possible implementation, the forwarding node also communicates with the management node. The method further includes: the forwarding node acquiring at least one third alarm rule. A third alarm rule includes: an observed object label, an indicator identifier, and an alarm policy. The forwarding node determines the amount of resources required to implement at least one third alarm rule and at least one alarm rule. If the amount of resources is greater than or equal to the amount of resources that multiple alarm computing nodes can provide, the forwarding node sends a scaling request to the management node. The scaling request indicates whether to increase the number of alarm computing nodes or the amount of resources that multiple alarm computing nodes can provide. This ensures that multiple alarm computing nodes can provide sufficient resources to implement at least one new third alarm rule and at least one new alarm rule, thereby ensuring the accuracy of anomaly detection results.

[0013] In another possible implementation, at least one third alarm rule includes M observation object labels and N indicator identifiers. At least one alarm rule includes m observation object labels and n indicator identifiers. The forwarding node determines the resource quantity required to implement at least one third alarm rule and at least one alarm rule, including: the forwarding node determining the M observation object labels and the matching observation object labels included in the m observation object labels, as well as the non-matching observation object labels. The forwarding node determines the N indicator identifiers and the matching indicator identifiers included in the n indicator identifiers, as well as the non-matching indicator identifiers. The forwarding node determines the resource quantity based on the matching observation object labels, non-matching observation object labels, matching indicator identifiers, and non-matching indicator identifiers. Thus, the forwarding node determines the resource quantity required to implement at least one new third alarm rule and at least one alarm rule based on the matching observation object labels, non-matching observation object labels, matching indicator identifiers, and non-matching indicator identifiers. The forwarding node uses a classification method to determine the resource quantity required to implement at least one new third alarm rule and at least one alarm rule, improving the accuracy of the determined required resource quantity.

[0014] In another possible implementation, each of the multiple alarm calculation nodes is configured with the same alarm rules. With identical alarm rules for each node, when a forwarding node receives metric data, it can determine the alarm calculation node that processes the metric data based on the observed object labels in the metric data. This shortens the time required to determine the alarm calculation node and improves anomaly detection efficiency.

[0015] Secondly, this application provides an anomaly detection method. This method can be executed by a first computing node among multiple alarm computing nodes. The first alarm computing node has a first alarm rule set. The first alarm rule includes: an observation object label, an indicator identifier, and an alarm policy. The anomaly detection method provided in this second aspect of the application includes: the first alarm computing node receiving indicator data. The indicator data includes: an observation object label, an indicator identifier, and an indicator value. The observation object label is used to indicate the observation object. The observation object label corresponds to the first alarm computing node. The observation object label of the first alarm rule includes the observation object label of the indicator data. The indicator identifier of the first alarm rule matches the indicator identifier of the indicator data. The first alarm computing node uses the first alarm rule to perform anomaly detection on the indicator data and obtain an anomaly detection result.

[0016] In the second aspect of this application, the observation object labels of the indicator data received by the first alarm calculation node correspond to the first alarm calculation node. In this way, the alarm calculation node processes indicator data including the corresponding observation object labels, without having to process indicator data including other observation object labels. This reduces the amount of indicator data that a single alarm calculation node needs to process, and provides a guarantee that a single alarm calculation node can provide sufficient resources to implement alarm rules, thereby improving the accuracy of anomaly detection results.

[0017] The aforementioned first alarm computing node can be a single computing device, a cluster of computing devices including multiple computing devices, a component of a computing device such as a processor, chip, or chip system, or a logic module or software that can implement all or part of the functions of a computing device.

[0018] In another possible implementation, the first alarm calculation node is also configured with other alarm rules.

[0019] In another possible implementation, the other alarm rules include a second alarm rule. The second alarm rule includes: an observation object label, an indicator identifier, and an alarm policy. The anomaly detection method provided in the second aspect further includes: if the observation object label of the second alarm rule includes the observation object label of the indicator data, and the indicator identifier of the second alarm rule matches the indicator identifier of the indicator data, then the first alarm computing node uses a first storage space to store the indicator data. The first storage space supports access from a first process and a second process. The first process is the process implementing the first alarm rule. The second process is the process implementing the second alarm rule. Thus, by using a first storage space that supports access from both the first and second processes to store the indicator data, the first alarm computing node reduces the number of copies of indicator data that need to be stored and lowers the storage resources occupied by storing the indicator data.

[0020] In another possible implementation, the first alarm computing node uses a first alarm rule to perform anomaly detection on the indicator data and obtain anomaly detection results. This includes: the first alarm computing node retrieving indicator data from a first storage space based on a first address. The first address is the address of the first storage space. The first alarm computing node also uses the first alarm rule to perform anomaly detection on the indicator data and obtain anomaly detection results. Thus, when implementing the first alarm rule, retrieving indicator data from the first address of the first alarm computing node reduces the number of copies of indicator data sent by the forwarding nodes to the first alarm computing node, saving communication resources between the forwarding nodes and the first alarm computing node.

[0021] In another possible implementation, the first alarm computing node receives a scaling command. The scaling command instructs the first alarm computing node to increase the number of alarm computing nodes and / or the amount of resources they can provide. In response to the scaling command, the first alarm computing node increases the number of alarm computing nodes and / or the amount of resources they can provide. This ensures that the alarm computing nodes have sufficient resources to detect anomalies in the indicator data, thereby ensuring the accuracy of the detection results.

[0022] In another possible implementation, each of the multiple alarm calculation nodes is configured with the same alarm rules. With identical alarm rules for each node, when a forwarding node receives metric data, it can determine the alarm calculation node that processes the metric data based on the observed object labels in the metric data. This shortens the time required to determine the alarm calculation node and improves anomaly detection efficiency.

[0023] Thirdly, this application provides an anomaly detection apparatus. The apparatus includes modules for performing the anomaly detection method described in the first aspect or any possible design of the first aspect.

[0024] Fourthly, this application provides an anomaly detection apparatus. The apparatus includes modules for performing the anomaly detection method of the second aspect or any possible design of the second aspect.

[0025] Fifthly, this application provides a processor. The processor includes an interface circuit and a control circuit. The interface circuit is used to acquire multiple index data and cooperate with the control circuit to implement the operational steps of the method in the first aspect or any possible design of the first aspect. Alternatively, the interface circuit is used to receive index data and cooperate with the control circuit to implement the operational steps of the method in the second aspect or any possible design of the second aspect.

[0026] Sixthly, this application provides a computing device. The computing device includes at least one processor and a memory, the memory being used to store a set of computer instructions; when the processor executes the set of computer instructions as an execution device in the first aspect or any possible implementation of the first aspect, it executes the operation steps of the anomaly detection method in the first aspect or any possible implementation of the first aspect, or executes the operation steps of the anomaly detection method in the second aspect or any possible implementation of the second aspect.

[0027] In a seventh aspect, this application provides a computing device cluster. The computing device cluster includes at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the at least one memory to cause the computing device cluster to perform operational steps of the anomaly detection method in the first aspect or any possible design of the first aspect, or to cause the computing device cluster to perform operational steps of the anomaly detection method in the second aspect or any possible design of the second aspect.

[0028] Eighthly, this application provides a computer-readable storage medium. It includes: computer software instructions; when executed in a computing device, the computer software instructions cause the computing device to perform operational steps of the method as described in the first aspect or any possible implementation thereof, or cause the computing device to perform operational steps of the method as described in the second aspect or any possible implementation thereof.

[0029] Ninthly, this application provides a computer program product. When the computer program product is run on a computer cluster, it causes the computing device cluster to perform the operation steps of the method as described in the first aspect or any possible implementation of the first aspect, or causes the computing device to perform the operation steps of the method as described in the second aspect or any possible implementation of the second aspect.

[0030] The beneficial effects of aspects three through nine above can be described with reference to the implementation of any of aspects one or two, and will not be repeated here. Based on the implementations provided in the above aspects, this application can be further combined to provide more implementations. Attached Figure Description

[0031] Figure 1 This application provides a schematic diagram of the architecture of an anomaly detection system.

[0032] Figure 2 A schematic diagram of the structure of a computing device provided in this application;

[0033] Figure 3 This application provides a schematic diagram of the structure of a chip;

[0034] Figure 4 A flowchart illustrating an anomaly detection method provided in this application;

[0035] Figure 5 A flowchart illustrating the correspondence between target alarm calculation nodes and observed object labels provided in this application;

[0036] Figure 6 A comparison chart showing the sending of multiple metrics data to the forwarding node;

[0037] Figure 7 A schematic diagram of a specified format provided for this application;

[0038] Figure 8 A comparison chart of the indicator data sent by the forwarding nodes;

[0039] Figure 9 An example diagram of the evaluation of a new alarm rule provided in this application;

[0040] Figure 10 A flowchart illustrating an anomaly detection method provided in this application;

[0041] Figure 11 This application provides a schematic diagram of the structure of an anomaly detection device.

[0042] Figure 12 This application provides a schematic diagram of the structure of a computing device cluster;

[0043] Figure 13 This is a schematic diagram of the connection between computing devices provided in this application. Detailed Implementation

[0044] This application provides an anomaly detection method. Instead of a forwarding node forwarding multiple metric data points to a single alarm computing node, this application involves the forwarding node sending a portion of the multiple metric data points to each of at least two target alarm computing nodes. This allows the at least two target alarm computing nodes to jointly process the multiple metric data points using their provided resources. This reduces the amount of metric data points that a single alarm computing node needs to process, avoids inaccurate anomaly detection results due to resource constraints of a single alarm computing node, ensures sufficient resources for alarm computing nodes to process metric data, and ultimately improves the accuracy of anomaly detection results.

[0045] To ensure clarity and conciseness in the description of the following embodiments, the relevant terms that may be involved in this application will be explained first.

[0046] A time series (TS) is a sequence of data that arranges the values ​​of the same indicator in chronological order of their occurrence. In some possible examples, a time series can also be called a timeline.

[0047] Alarm rules are rules used to determine whether the value of a specified indicator of an observed object meets the specified alarm strategy.

[0048] Anomaly detection refers to the process of determining whether the value of a specified indicator of an observed object meets the alarm rules. If the value of the specified indicator of the observed object meets the alarm rules, the computing device provides alarm information.

[0049] The above text has explained the relevant terms that may be involved in this application. The following text describes the prior art involved in this application.

[0050] The anomaly detection system sets an alarm rule on an alarm calculation node, and uses this node to determine whether the value of a specified indicator for all observed objects meets the alarm rule. In this process, as the number of observed objects increases, the limitations of the resources provided by a single alarm calculation node may lead to inaccurate detection results.

[0051] Based on this, this application provides an anomaly detection method, which can be applied to... Figure 1 Anomaly detection system. Figure 1 This application provides an architecture diagram of an anomaly detection system, as shown below. Figure 1 As shown, the anomaly detection system 100 includes a forwarding node 110 and multiple alarm computing nodes 120. The forwarding node 110 and the multiple alarm computing nodes 120 are communicatively connected. This communication connection can be wired or wireless. Wired communication can be Ethernet, fiber optic, or various Peripheral Component Interconnect Express (PCIe) buses installed within the anomaly detection system 100 to connect the forwarding node 110 and the multiple alarm computing nodes 120. Wireless communication can be the Internet, Wireless Local Area Network (WLAN), or Ultra Wide Bandwidth (UWB) technology.

[0052] Forwarding node 110 acquires multiple indicator data and determines at least two target alarm computing nodes based on the observation object labels of the multiple indicator data, and sends the multiple indicator data to the at least two target alarm computing nodes. Forwarding node 110 can be a single computing device, a cluster of computing devices including multiple computing devices, or a component of a computing device, such as a processor, chip, or chip system, etc., which is not limited in this application. When forwarding node 110 is a single computing device, the computing device can be a common computing device, such as a server, personal computer, desktop computer, mobile phone, etc.

[0053] Alarm computing node 120 is used to receive indicator data and perform anomaly detection on the indicator data based on alarm rules to obtain anomaly detection results. Alarm computing node 120 can be a single computing device, a cluster of computing devices including multiple computing devices, or a component of a computing device, such as a processor, chip, or chip system; this application does not limit this. When alarm computing node 120 is a single computing device, this computing device can be a common computing device, such as a server, personal computer, desktop computer, or mobile phone.

[0054] Each of the multiple alarm calculation nodes 120 can be configured with at least one alarm rule, and an alarm rule can be configured on at least two alarm calculation nodes. Multiple alarm calculation nodes 120 can be configured with alarm rules in various ways, which are explained below for different scenarios.

[0055] In scenario a, each of the multiple alarm calculation nodes 120 has the same alarm rules set.

[0056] In this scenario, each alarm calculation node 120 is configured with at least one identical alarm rule. For example, the anomaly detection system 100 includes alarm calculation node 1 and alarm calculation node 2. Both alarm calculation node 1 and alarm calculation node 2 are configured with alarm rule 1. Alternatively, the anomaly detection system 100 includes alarm calculation node 1 and alarm calculation node 2. Both alarm calculation node 1 and alarm calculation node 2 are configured with alarm rules 1 through N, where N ≥ 2.

[0057] In scenario b, at least one alarm calculation node 120 exists among the multiple alarm calculation nodes 120, and the alarm rules set by this at least one alarm calculation node 120 are different from the alarm rules set by the other alarm calculation nodes 120.

[0058] In this case, the fact that the alarm rules set by the at least one alarm computing node 120 are different from the alarm rules set by other alarm computing nodes 120 can mean that the alarm rules set by the at least one alarm computing node 120 are completely different from the alarm rules set by other alarm computing nodes 120, or it can mean that the alarm rules set by the at least one alarm computing node 120 are partially different from the alarm rules set by other alarm computing nodes 120. This application does not limit this.

[0059] The alarm rules set by at least one alarm calculation node 120 can be completely different from the alarm rules set by other alarm calculation nodes 120. Specifically, this means that the alarm rules set by at least one alarm calculation node 120 do not contain any alarm rules identical to those set by other alarm calculation nodes 120. For example, the anomaly detection system 100 includes alarm calculation nodes 1 to 4. Alarm calculation nodes 1 and 2 set alarm rules 1 to N, where N ≥ 2. Alarm calculation nodes 3 and 4 set alarm rule N+1.

[0060] The alarm rules set by at least one alarm computing node 120 may differ from those set by other alarm computing nodes 120. This could mean that the alarm rules set by the at least one alarm computing node 120 include the alarm rules set by other alarm computing nodes 120 and other alarm rules, or it could mean that the alarm rules set by other alarm computing nodes 120 include the alarm rules set by the at least one alarm computing node 120 and other alarm rules, and so on. For example, the anomaly detection system 100 includes alarm computing nodes 1 to 4. Alarm computing nodes 1 and 2 set alarm rules 1 to N, and alarm computing nodes 3 and 4 set alarm rules 1 to N-1, where N ≥ 3 and is an integer.

[0061] In some possible scenarios, the aforementioned alarm rules can be preset or set according to the needs of actual applications, and this application does not limit this. For example, if it is necessary to perform alarm detection on multiple computing nodes running Task 1, one or more alarm rules can be set to detect the running status of these multiple computing nodes according to the needs of actual applications.

[0062] In some possible scenarios, the anomaly detection system 100 may further include a management node 130. The management node 130 is communicatively connected to the forwarding node 110 and multiple alarm computing nodes 120. The management node 130 receives expansion requests from the forwarding node 110 and, in response to the expansion requests, sends expansion commands to the alarm computing nodes. Similarly, the management node 130 may be a single computing device, a cluster of computing devices including multiple computing devices, or a component of a computing device, such as a processor, chip, or chip system; this application does not limit this. When the management node 130 is a single computing device, this computing device may be a common computing device, such as a server, personal computer, desktop computer, or mobile phone.

[0063] The aforementioned forwarding node 110, alarm computing node 120, and management node 130 can be implemented using computing devices. These computing devices can be common computing devices, such as personal computers, desktop computers, mobile phones, servers, etc. For example, Figure 2 A schematic diagram of the structure of a computing device provided in this application, such as... Figure 2 As shown, the computing device 200 includes a communication interface 214, a processor 211, and a memory 212. The communication interface 214 is used to communicate with other devices located outside the computing device 200. For example, when the computing device 200 is used to implement the function of the forwarding node 110, the computing device 200 interacts with the alarm computing node 120 through the communication interface 214. Specifically, the computing device 200 acquires multiple indicator data through the communication interface 214 and forwards the multiple indicator data to the alarm computing node 120 corresponding to the observation object tag of the indicator data. As another example, when the computing device 200 is used to implement the function of the alarm computing node 110, the computing device 200 interacts with the forwarding node 110 through the communication interface 214. Specifically, the computing device 200 receives indicator data through the communication interface 214, obtains processing results based on the acquired indicator data (such as the observation object's working status being abnormal / normal as indicated by the observation object tag), and feeds back the processing results through the communication interface 214. This communication interface 214 can be an input / output (I / O) interface. For example, when the computing device 200 is used to implement the function of the alarm management node 130, the computing device 200 interacts with the forwarding node 110 and the alarm computing node 120 through the communication interface 214. Specifically, the computing device 200 receives the expansion request from the forwarding node 110 through the communication interface 214, and the computing device 200 sends an expansion command to the alarm computing node based on the received expansion request.

[0064] Processor 211 is the core of computing device 200 for both computation and control. It may include: a central processing unit (CPU), a specific integrated circuit, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In practical applications, computing device 200 may also include multiple processors. Processor 211 may include one or more processor cores. An operating system and other software programs are installed in processor 211, enabling it to access memory 212 and various peripheral component interconnect (PCIe) devices.

[0065] The processor 211 is connected to the memory 212 via bus 216. Bus 216 can be a double data rate (DDR) bus or other types of bus. Memory 212 is the main memory of computing device 200. Memory 212 is typically used to store various running software in the operating system. To improve the access speed of processor 211, memory 212 needs to have the advantage of high access speed. In traditional computer devices, dynamic random access memory (DRAM) is usually used as memory 212. In addition to DRAM, memory 212 can also be other random access memories, such as static random access memory (SRAM). Alternatively, memory 212 can also be read-only memory (ROM). For example, read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). This embodiment does not limit the number or type of memory 212.

[0066] In some possible scenarios, in order to persistently store data (such as indicator data, anomaly detection results, etc.), the computing device 200 may also be equipped with a data storage system 213, which may be located outside the computing device 200 (e.g., Figure 2 As shown, the data storage system 213 exchanges data with the computing device 200 via a network. Optionally, the data storage system 213 can also be located inside the host, such as by exchanging data with the processor 211 via the bus 216. In this case, the data storage system 213 manifests as a hard disk.

[0067] For example, Figure 2 The processor 211 in the text can be implemented through a chip, such as... Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a chip provided in this application. For example, the chip 300 includes a core 301, a CPU 302, a system buffer 303, an input / output (I / O) device 305, and a DDR 306.

[0068] The CPU 302 is used to accept tasks (such as compression tasks, decompression tasks, and anomaly detection tasks) and call core 301 to execute the tasks. When chip 300 has multiple cores 301, the CPU 302 is also used for scheduling tasks. For example, the CPU 302 can be implemented by an ARM processor, which is small in size, low in power consumption, uses a 34-bit reduced instruction set, and has simple and flexible addressing. Of course, in some implementations, the CPU 302 can also be implemented by other processors.

[0069] Core 301 provides the computing power required for tasks such as anomaly detection. In one optional configuration, core 301 includes a load / store unit (LSU), a cube computing unit, a scalar computing unit, a vector computing unit, and a buffer. The LSU loads data to be processed (such as index data) and stores processed data. It also manages read / write operations between different buffers within the core and performs format conversions. The cube computing unit provides the core computing power for matrix multiplication. The scalar computing unit is a single-instruction single-data (SISD) processor, meaning it processes only one data item (usually an integer or floating-point number) at a time. The vector computing unit, also known as an array processor, is a processor capable of directly manipulating an array or vector for computation. The number of buffers may be one or more. For example, this buffer mainly refers to the level 1 buffer (L1 buffer). The buffer is used to temporarily store some data that the core 301 needs to use repeatedly, thereby reducing read and write operations from the bus. In addition, the implementation of certain data format conversion functions also requires the source data to be located in the buffer. The system buffer 303 mainly refers to the level 2 buffer, which is used to temporarily store the input data, intermediate results, or final results that have passed through the chip.

[0070] DDR 306 is an off-chip memory that can be replaced by high-bandwidth memory (HBM) or other off-chip memory. Located between the chip and external memory, DDR 306 overcomes the access speed limitations of shared memory read / write operations in computing resource sharing.

[0071] The I / O device 305 included in chip 300 refers to the hardware that performs data transmission, or the device that interfaces with the I / O interface. Common I / O devices 305 include network cards, printers, keyboards, and mice. All external storage devices can also be used as I / O devices 305, such as hard drives, floppy disks, and optical discs.

[0072] Core 301, CPU 302, system buffer 303, I / O devices 305, and DDR 306 are connected via a bus. The bus may include a pathway for transmitting information between the aforementioned components (such as CPU 302 and system buffer 303). In addition to a data bus, the bus may also include a power bus, control bus, and status signal bus. However, for clarity, the bus may be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. For example, core 301 can access these I / O devices 305 via a PCIe bus. Core 301 is connected to system buffer 303 via a DDR bus. Here, different system buffers 303 may use different data buses to communicate with core 301; therefore, the DDR bus can also be replaced with other types of data buses. This application embodiment does not limit the bus type.

[0073] For example, after CPU 302 loads the data (such as indicator data) to be processed by the anomaly detection task into DDR 306, the LSU in core 301 reads (loads) the data from DDR 306, processes the data, and obtains the processing result (such as the abnormal / normal working status of the observed object indicated by the observed object tag). After obtaining the processing result, the LSU then loads (stores) the processing result into DDR 306, and the network interface card feeds back the processing result.

[0074] It is understood that the structures illustrated in the above embodiments do not constitute a specific limitation on the anomaly detection system, computing device, or chip. In other embodiments, the anomaly detection system may include more or fewer components (e.g., the anomaly detection system may also include other forwarding nodes, and all forwarding nodes receive multiple indicator data and forward the multiple indicator data to the alarm computing node corresponding to the observed object tag), and the computing device and chip may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0075] The above text combined Figures 1 to 3 The anomaly detection system applying the anomaly detection method, as well as the computing devices and chips that can be used to implement forwarding nodes, alarm computing nodes 120, etc., are described below. Figures 1 to 3The content shown provides a detailed description of the anomaly detection method provided in this application.

[0076] Figure 4 This is a flowchart illustrating an anomaly detection method provided in this application. This anomaly detection method can be applied to an anomaly detection system, which can include a forwarding node and at least one alarm calculation node 120. The anomaly detection system can employ... Figure 1 The architecture described. For a more detailed description of the anomaly detection system, please refer to the above text. Figures 1 to 3 The relevant descriptions will not be repeated here.

[0077] This section focuses on the anomaly detection system's capabilities. Figure 1 The anomaly detection method provided in this application is illustrated using the described architecture as an example. This anomaly detection method can be executed by a forwarding node, such as... Figure 4 As shown, the method may include the following steps S410 to S450.

[0078] S410, the forwarding node obtains multiple indicator data.

[0079] Each of the multiple indicator data sets includes: an observation object label, an indicator identifier, and an indicator value. The multiple indicator data sets include multiple observation object labels.

[0080] Observation object tags are used to identify the observed object. Observation object tags may include, but are not limited to: the observation object type identifier (ID), the observation object name, the geographical location of the observation object, etc. The observation object type identifier indicates the type of the observation object, which includes, but is not limited to: processor, hard drive, network card, router, etc. The geographical location of the observation object includes, but is not limited to: the region where the observation object is located, the availability zone (azone, az) where the observation object is located, the microservice to which the observation object belongs, etc. For example, observation object tags may include: the region where the observation object is located, the availability zone where the observation object is located, and the observation object type. For example, the observation object tag could be Region = Region 1, Availability Zone = Availability Zone 1, Observation Object Type = Processor.

[0081] The object of observation can include a single device or a cluster of devices, which is not limited in this application. The multiple devices constituting a device cluster can be devices performing the same task or devices located in the same geographical location, which is not limited in this application. For example, the object of observation can refer to a host, a cluster of multiple hosts, a microservice, etc. For instance, when the object of observation is a host, the host's Internet Protocol (IP) address and the geographical location of the object can be used as the object label. Similarly, when the object of observation is a cluster, the cluster's physical address information and the cluster's name can be used as the object label. Furthermore, when the object of observation is a microservice, the microservice's physical address information and the microservice's name can be used as the object label.

[0082] Indicator identifiers are used to indicate indicators, including but not limited to: indicator name, indicator code, etc. Indicators may include, but are not limited to: temperature, processor utilization, processor idle rate, storage resource utilization, storage resource idle rate, etc. Depending on the needs of actual applications, indicators may also include other content, such as humidity, etc., which are not limited in this application.

[0083] In some possible scenarios, arranging the values ​​of the same indicator for the same observed object in chronological order yields a time series or a timeline. For example, arranging the temperature values ​​of host 1 in chronological order yields a time series or a timeline. Similarly, arranging the temperature values ​​of each host from host 1 to host n in chronological order yields n time series or n timelines, with one timeline corresponding to one host.

[0084] Based on the above explanation of the various contents of the indicator data, the following section uses S420 to explain the process by which the forwarding node determines the target alarm calculation node for each indicator data among multiple indicator data.

[0085] S420, the forwarding node determines at least two target alarm calculation nodes corresponding to the indicator identifiers of multiple observation object labels and / or multiple indicator data.

[0086] Specifically, at least two target alarm calculation nodes are alarm calculation nodes among multiple alarm calculation nodes that have deployed the first alarm rule. The indicator identifier of the first alarm rule matches the indicator identifiers of multiple indicator data. The observation object label of the first alarm rule includes the observation object labels of multiple indicator data.

[0087] The forwarding node stores the alarm rules set by each of the multiple alarm computing nodes, as well as the at least two target alarm computing nodes that process multiple indicator data based on multiple observation object tags and / or indicator identifiers of multiple indicator data. The forwarding node can store the alarm rules set by each of the multiple alarm computing nodes in various forms, including but not limited to lists. For example, the forwarding node stores a rule list, which includes the alarm rules set by each of the multiple alarm computing nodes.

[0088] Depending on the different ways alarm rules are deployed across multiple alarm calculation nodes, the at least two target alarm calculation nodes determined by the forwarding node will be different. The following describes the different scenarios.

[0089] Scenario A: Multiple alarm computing nodes are deployed with different alarm rules.

[0090] The deployment of different alarm rules among multiple alarm computing nodes can mean that at least one of the multiple alarm computing nodes has alarm rules set differently from those set by the other alarm computing nodes. The other alarm computing nodes are those other than the at least one alarm computing node. In this case, the forwarding node can determine at least two target alarm computing nodes from the multiple alarm computing nodes to process the multiple indicator data based on the multiple observation object tags and indicator identifiers included in the multiple indicator data.

[0091] In this scenario, the target alarm calculation node determined by the forwarding node can be all or part of multiple alarm calculation nodes. For example, the multiple alarm calculation nodes may include alarm calculation nodes 1 to n. The alarm rules may include alarm rule 1 to alarm rule N. Alarm calculation nodes 1 to n-1 set alarm rules 1 to N-1, and alarm calculation node n sets alarm rules 1 to N. If the indicator identifier of alarm rule N matches the indicator identifiers of multiple indicator data, and the observation object label of alarm rule N includes multiple observation object labels of multiple indicator data, then the forwarding node determines alarm calculation nodes 1 to n as the target alarm calculation nodes for implementing alarm rule N. In other words, in this scenario, the forwarding node determines all of the multiple alarm calculation nodes as the target alarm calculation node. If the indicator identifiers of alarm rules 1 to alarm rule (N-1) match the indicator identifiers of multiple indicator data, and the observation object labels of alarm rules 1 to alarm rule (N-1) include multiple observation object labels of multiple indicator data, then the forwarding node determines alarm calculation nodes 1 to alarm calculation nodes (n-1) as the target alarm calculation nodes for implementing alarm rules 1 to alarm rule (N-1). In other words, in this case, the forwarding node determines a portion of the multiple alarm calculation nodes as the target alarm calculation nodes.

[0092] The forwarding node can perform the following process to determine the target alarm calculation node for processing the indicator data. Specifically, the first alarm calculation node among multiple alarm calculation nodes has a first alarm rule set. The first alarm rule includes a first observation object label and a first indicator identifier. The forwarding node obtains the observation object label and indicator identifier of the first indicator data. If the first indicator identifier of the first alarm rule is the same as the indicator identifier of the first indicator data, and the first observation object label of the first alarm rule includes the observation object label of the first indicator data, then the forwarding node determines the first alarm calculation node as the first target alarm calculation node for processing the first indicator data.

[0093] If the observation object indicated by the observation object label of the first indicator data is included by the observation object indicated by the first observation object label of the first alarm rule, then the first observation object label of the first alarm rule can be considered to include the observation object label of the first indicator data. Similarly, if the observation object indicated by the first observation object label of the first alarm rule is the same as the observation object indicated by the observation object label of the first indicator data, then the first observation object label of the first alarm rule can also be considered to include the observation object label of the first indicator data.

[0094] For example, alarm computing node 1 among multiple alarm computing nodes is configured with alarm rule 1. Alarm rule 1 includes an observation object label of "ip=xx.xx.xx.*", and the indicator identifier indicates processor utilization. Multiple indicator data include indicator data 1. The observation object label of indicator data 1 is "ip=xx.xx.xx.1", and the indicator identifier indicates processor utilization. In this case, the forwarding node determines that the observation object label "ip=xx.xx.xx.*" of alarm rule 1 includes the observation object label "ip=xx.xx.xx.1" of indicator data 1, and the indicator identifier of alarm rule 1 is the same as the indicator "processor utilization" indicated by the indicator identifier of indicator data 1. The forwarding node then determines that alarm computing node 1 is the target alarm computing node for processing indicator data 1.

[0095] In scenario B, multiple alarm computing nodes are deployed with the same alarm rules.

[0096] The alarm rules deployed on each of the multiple alarm calculation nodes are identical. The forwarding node determines at least two target alarm calculation nodes from among the multiple alarm calculation nodes to process the multiple metric data based on the multiple observation object labels of the multiple metric data. Specifically, the forwarding node can determine at least two target alarm calculation nodes to process the multiple metric data based on the correspondence between the multiple observation object labels of the multiple metric data and the node identifiers of the multiple alarm calculation nodes. For the process by which the forwarding node determines the correspondence between the node identifiers of the target alarm calculation nodes and the observation object labels of the metric data, please refer to the descriptions in S51 to S53 below, which will not be repeated here.

[0097] For example, multiple alarm calculation nodes include alarm calculation node 1 to alarm calculation node n. Each alarm calculation node from alarm calculation node 1 to alarm calculation node n is configured with alarm rule 1 to alarm rule N. Where n ≥ 2, N ≥ 2, and both n and N are integers. Multiple indicator data include indicator data 1 and indicator data 2. Indicator data 1 includes observation object label 1 and indicator identifier 1, and indicator data 2 includes observation object label 2 and indicator identifier 1. The indicator identifier of alarm rule 1 matches indicator identifier 1, and the observation object label of alarm rule 1 includes observation object label 1 and observation object label 2. The node identifier of alarm calculation node 1 corresponds to observation object label 1, and the node identifier of alarm calculation node 2 corresponds to observation object label 2. In this case, the forwarding node determines alarm calculation node 1 and alarm calculation node 2 as target alarm calculation nodes.

[0098] The above section explained the target alarm calculation node determined by the forwarding node under different scenarios. The following section explains the process of the forwarding node forwarding multiple indicator data to the target alarm calculation node in conjunction with S430.

[0099] S430, the forwarding node sends multiple indicator data to at least two target alarm calculation nodes.

[0100] The system includes at least two target alarm calculation nodes, namely a first target alarm calculation node and a second target alarm calculation node. Multiple indicator data points are included, including first indicator data and second indicator data. The first indicator data has a first observation object label. The second indicator data has a second observation object label. The first indicator data is sent to the first target alarm calculation node. The second indicator data is sent to the second target alarm calculation node.

[0101] A forwarding node can determine a scheme for sending multiple indicator data to at least two target alarm computing nodes using various methods. For example, the forwarding node can use a hash algorithm to determine the scheme based on the node identifier of the target alarm computing node and the observation object label of the indicator data. Specifically, the forwarding node uses a hash algorithm to determine the correspondence between the node identifier of each of the at least two target alarm computing nodes and each observation object label included in the multiple observation object labels of the multiple indicator data. The forwarding node then sends the indicator data that includes the observation object label corresponding to the node identifier of the target alarm computing node to the target alarm computing node. The forwarding node can use the process described in S51 to S53 below to determine the correspondence between the node identifier of the target alarm computing node and the observation object label of the indicator data. Please refer to the following text for a detailed description, which will not be repeated here.

[0102] For example, at least two target alarm calculation nodes include target alarm calculation node 1 to target alarm calculation node 3. Multiple indicator data include indicator data 1 to indicator data 6. The observation object labels for indicator data 1 to indicator data 6 are observation object label 1 to observation object label 6, respectively. The forwarding node uses a hash algorithm to determine the correspondence between the node identifier of target alarm calculation node 1 and observation object label 1 and observation object label 2, the node identifier of target alarm calculation node 2 and observation object label 3 and observation object label 4, and the node identifier of target alarm calculation node 3 and observation object label 5 and observation object label 6. In this scenario, the forwarding node sends indicator data 1 and indicator data 2 to target alarm calculation node 1, indicator data 3 and indicator data 4 to target alarm calculation node 2, and indicator data 5 and indicator data 6 to target alarm calculation node 3.

[0103] The above describes the process of a forwarding node sending multiple indicator data to at least two target alarm calculation nodes. The following example illustrates the process of a target alarm calculation node processing indicator data, using the example of a forwarding node sending first indicator data (including the first observation object label) to the first target alarm calculation node among at least two target alarm calculation nodes. Specifically, the first target alarm calculation node can execute the following steps S440 and S450.

[0104] Accordingly, S440 includes: the first target alarm calculation node receiving the first indicator data.

[0105] S450, the first target alarm calculation node uses the first alarm rule to perform anomaly detection on the first indicator data and obtain the anomaly detection result.

[0106] The first target alarm calculation node can utilize the alarm strategy of the first alarm rule to perform anomaly detection on the first indicator data and obtain anomaly detection results. Furthermore, if the anomaly detection results indicate that the indicator value of the first indicator data is abnormal, the first target alarm calculation node can provide alarm information. This alarm information includes, but is not limited to: alarm sound, alarm light, alarm interface, etc.

[0107] For example, the observation object label of the first alarm rule is "ip=xx.xx.xx.*", the indicator identifier indicates the processor utilization rate, and the alarm policy is to trigger an alarm when the processor utilization rate is greater than or equal to 85%. The observation object label of the first indicator data is "ip=xx.xx.xx.1", the indicator identifier indicates the processor utilization rate, and the indicator value is 90%. In this case, the first target alarm computing node detects that the indicator value of the first indicator data meets the alarm policy, and the first target alarm computing node provides an alarm sound.

[0108] The above example illustrates the process of the forwarding node and the first target alarm calculation node cooperating in processing the first indicator data, using the setting of the first alarm rule by the first target alarm calculation node as an example. Depending on the needs of actual applications, the first target alarm calculation node can also set other alarm rules. These other alarm rules can be one or multiple alarm rules; this application does not limit this. For example, if the first target alarm calculation node also sets a second alarm rule, which includes the observed object label, indicator identifier, and alarm strategy, the first target alarm calculation node can implement the second alarm rule using the process described above. Please refer to the above description for relevant details, which will not be repeated here.

[0109] The above section explained the process of forwarding indicator data by forwarding nodes from S410 to S460, and the process of target alarm calculation nodes processing indicator data. The following section will combine... Figure 5This application explains the process of determining the correspondence between target alarm calculation nodes and observed object labels.

[0110] Figure 5 The flowchart provided in this application illustrates a method for determining the correspondence between target alarm calculation nodes and observed object labels, as shown below. Figure 5 As shown, the process includes S51 to S53.

[0111] S51, the forwarding node determines the first observation object label corresponding to the first target alarm calculation node and the second observation object label corresponding to the second target alarm calculation node based on multiple observation object labels and the node identifiers of at least two target alarm calculation nodes.

[0112] Forwarding nodes can use various methods to determine the first observation object label corresponding to the first target alarm computing node and the second observation object label corresponding to the second target alarm computing node, such as hash algorithms, consistent hashing algorithms, etc. The following example illustrates the process by which a forwarding node uses a consistent hashing algorithm to determine the first observation object label corresponding to the first target alarm computing node and the second observation object label corresponding to the second target alarm computing node. Specifically, this process can include the following ① to ③.

[0113] ①The forwarding node generates multiple first hash values ​​based on the labels of multiple observed objects.

[0114] Each hash value corresponds to a label of an observed object.

[0115] For example, the observed object refers to a host, and indicator data 1 includes the observed object label 1. The observed object label 1 includes the region where the observed object is located, the availability zone where the observed object is located, and the IP address of the observed object, as well as region = region 1, availability zone = availability zone 1, and IP = IP1. In this case, the forwarding node can generate hash value 1 based on "region 1, availability zone 1, and IP1".

[0116] In some possible scenarios, where the observation object label includes multiple contents, the forwarding node can sort these multiple contents in a specified order and generate a hash value based on the sorting result. For example, if the observation object label includes region, az, and ip, the forwarding node can sort the contents of the observation object label according to region, az, and ip, and generate a hash value based on the sorting result.

[0117] ②The forwarding node generates at least two second hash values ​​based on the node identifiers of at least two target alarm calculation nodes.

[0118] One second hash value corresponds to one target alarm calculation node.

[0119] ③ The forwarding node determines the first observation object label corresponding to the first target alarm calculation node and the second observation object label corresponding to the second target alarm calculation node based on multiple first hash values ​​and at least two second hash values.

[0120] Among them, the first observation object label and the second observation object label belong to multiple observation object labels.

[0121] In this scenario, the forwarding node can construct a hash ring based on at least two second hash values ​​generated from the node identifiers of at least two target alarm computing nodes. The forwarding node uses a hash algorithm to determine the second hash value corresponding to each of the multiple first hash values, thereby identifying the target alarm computing node corresponding to the observation object label. For example, the forwarding node determines that the node label of the first target alarm computing node corresponds to the first observation object label, and the node label of the second target alarm computing node corresponds to the second observation object label. For the process of the forwarding node determining the second hash value corresponding to the first hash value, please refer to the description of general techniques; it will not be repeated here.

[0122] S52, the forwarding node sends multiple indicator data, including the first observation object label, to the first target alarm calculation node.

[0123] S53, the forwarding node sends multiple indicator data, including the second observation object label, to the second target alarm calculation node.

[0124] The above section explained the process of determining the correspondence between the target alarm calculation node and the observed object label using steps S51 to S53. The following section will further explain... Figure 6 The process of forwarding nodes sending multiple indicator data is explained.

[0125] Figure 6 Send a comparison chart of multiple metric data to the forwarding node, such as Figure 6 As shown in (a), in a typical technology, an alarm rule is set on an alarm calculation node. The forwarding node sends all indicator data matching that alarm rule to that alarm calculation node. For example, if the forwarding node obtains indicator data 1 and indicator data 2, and if the indicator identifiers of indicator data 1 and indicator data 2 are the same as the first alarm rule set on the first alarm calculation node, and the observation object label of the first alarm rule includes the observation object labels of indicator data 1 and indicator data 2, then the forwarding node sends indicator data 1 and indicator data 2 to the first alarm calculation node. Figure 6As shown in (b), in this application, an alarm rule is set on at least two alarm calculation nodes. A forwarding node sends all indicator data matching the alarm rule to the at least two alarm calculation nodes. An alarm calculation node receives a portion of the indicator data from all the indicator data. Thus, the at least two alarm calculation nodes implement the alarm rule, improving the accuracy of anomaly detection results. For example, if a first alarm rule is set on both the first and second alarm calculation nodes, the forwarding node determines that the observation object label of indicator data 1 matches the node identifier of the first alarm calculation node, and the observation object label of indicator data 2 matches the node identifier of the second alarm calculation node. In this case, the forwarding node forwards indicator data 1 to the first alarm calculation node and indicator data 2 to the second alarm calculation node.

[0126] The above text combined Figure 6 The process of forwarding multiple metric data points by a forwarding node is illustrated comparatively. In some possible scenarios, to save data transmission resources, when at least two alarm rules set on an alarm calculation node use the same metric data, the forwarding node can send the same metric data to that alarm calculation node in a specified format. Furthermore, the alarm calculation node can store the same metric data in a shared storage space that supports access from multiple processes, thereby reducing the storage space required to store the metric data.

[0127] The first target alarm computing node sets a first alarm rule and a second alarm rule. Multiple indicator data include the first indicator data. If the indicator identifier of the first alarm rule matches the indicator identifier of the first indicator data, and the observation object label of the first alarm rule includes the observation object label of the first indicator data; and the indicator identifier of the second alarm rule matches the indicator identifier of the first indicator data, and the observation object label of the second alarm rule includes the observation object label of the first indicator data—that is, both the first and second alarm rules use the first indicator data—then the forwarding node sends the first indicator data to the first target alarm computing node in a specified format. The first target alarm computing node receives the first indicator data and stores it in a first storage space. The first storage space supports access by a first process and a second process. The first process is the process by which the first target alarm computing node implements the first alarm rule. The second process is the process by which the first target alarm computing node implements the second alarm rule. When the first target alarm computing node implements the first alarm rule, it retrieves the first indicator data from the first storage space based on a first address. This first address is the address of the first storage space. And the forwarding nodes use the first alarm rule to perform anomaly detection on the first indicator data and obtain the anomaly detection results.

[0128] The above example illustrates how the first target alarm computing node implements alarm rules using processes, explaining the process by which it acquires first indicator data and uses the first alarm rules to perform anomaly detection on the first indicator data to obtain anomaly detection results. In some possible scenarios, the first target alarm computing node can also implement alarm rules using threads. In this case, the first storage space supports access by both a first thread and a second thread. The first thread is the thread that implements the first alarm rule for the first target alarm computing node. The second thread is the thread that implements the second alarm rule for the first target alarm computing node. The first target alarm computing node also acquires the first indicator data from the first storage space based on a first address. This first address is the address of the first storage space. Finally, the forwarding node uses the first alarm rules to perform anomaly detection on the first indicator data and obtains the anomaly detection results.

[0129] In some possible situations, such as Figure 7 As shown, Figure 7 This application provides a schematic diagram of a specified format, which includes a first field and a second field. The first field carries the rule identifier of each of the at least two alarm rules, and the second field carries the same set of indicator data. For example, if both the first alarm rule and the second alarm rule use indicator data 1, the specified format includes field 1 and field 2. Field 1 carries the rule identifier of the first alarm rule and the rule identifier of the second alarm rule, and field 2 carries indicator data 1.

[0130] The above text explained the process of forwarding indicator data in a specified format by the forwarding nodes. The following section will combine this with... Figure 7 The process of providing forwarding index data in this application is compared with the process of providing forwarding index data in conventional technology.

[0131] Figure 8 A comparison chart of the metric data sent to the forwarding nodes, such as Figure 8 As shown in (a), in a typical technology, an alarm rule is set on an alarm computing node, and a forwarding node sends indicator data associated with that alarm rule to that alarm computing node. For example, alarm computing node 1 sets alarm rule 1, alarm computing node 2 sets alarm rule 2, and alarm computing node 3 sets alarm rule 3. Indicator data 1 includes an observation object label 1 and an indicator identifier, and indicator data 2 includes an observation object label 2 and an indicator identifier. The observation object labels for alarm rules 1 to 3 include observation object label 1 and observation object label 2, and the indicator identifiers for alarm rules 1 to 3 are the same as the indicator identifiers for indicator data 1 and indicator data 2. In this case, the forwarding node sends indicator data 1 and indicator data 2 to alarm computing nodes 1 to 3.

[0132] And such as Figure 8 As shown in (b), in this application, alarm calculation nodes 1 to 3 are all configured with alarm rules 1 to 3. Indicator data 1 includes observation object label 1 and indicator identifier, and indicator data 2 includes observation object label 2 and indicator identifier. The observation object label of alarm rule 1 includes observation object label 1 and observation object label 2, and the indicator identifier of alarm rule 1, indicator identifier of indicator data 1, and indicator identifier of indicator data 2 are the same. The observation object label of alarm rule 2 includes observation object label 1 and observation object label 2, and the indicator identifier of alarm rule 2, indicator identifier of indicator data 1, and indicator identifier of indicator data 2 are the same. The observation object label of alarm rule 3 includes observation object label 1 and observation object label 2, and the indicator identifier of alarm rule 3, indicator identifier of indicator data 1, and indicator identifier of indicator data 2 are the same. That is, alarm rules 1 to 3 all use indicator data 1 and indicator data 2. If a forwarding node determines that observation object label 1 corresponds to alarm calculation node 1, the forwarding node determines that observation object label 2 corresponds to alarm calculation node 2. In this scenario, the forwarding node can send indicator data 1 to alarm calculation node 1 and indicator data 2 to alarm calculation node 2 using a specified format. The specified format sent by the forwarding node to alarm calculation node 1 includes fields 11 and 12. Field 11 carries the rule identifiers for alarm rule 1, alarm rule 2, and alarm rule 3. Field 12 carries indicator data 1. Similarly, the specified format sent by the forwarding node to alarm calculation node 2 includes fields 21 and 22. Field 21 carries the rule identifiers for alarm rule 1, alarm rule 2, and alarm rule 3. Field 22 carries indicator data 2.

[0133] The above example illustrates the process of a forwarding node sending the same set of indicator data to an alarm computing node using a specified format, with two alarm rules set by the first target alarm computing node using the same indicator data. In some possible scenarios, a set of indicator data can be used by more alarm rules set by an alarm computing node, such as when the first target alarm computing node also sets a third alarm rule. The rule identifier of the third alarm rule matches the indicator identifier of the first indicator data, and the observation object label of the third alarm rule includes the observation object label of the first indicator data. In this case, the forwarding node sends the first indicator data to the alarm computing node in a specified format. This specified format includes a first field and a second field. The first field carries the rule identifiers of the first, second, and third alarm rules. The second field carries the first indicator data.

[0134] The above text describes the process of a forwarding node sending indicator data to a target alarm calculation node in a specified format, using the example of a first target alarm calculation node corresponding to a first observation object label and a second target alarm calculation node corresponding to a second observation object label.

[0135] Depending on the needs of the actual application, it may be necessary to add new alarm rules. If the resources required for adding new alarm rules and at least one alarm rule (or existing alarm rule) set by each alarm computing node in multiple alarm computing nodes are greater than or equal to the resources supported by multiple alarm computing nodes, it may affect the performance of the anomaly detection system and lead to inaccurate detection results. Based on this, the resources required for adding new alarm rules and existing rules can be evaluated before adding new alarm rules to detect the working status of the observed object. And it can be determined whether to expand the capacity based on the evaluated resources and the resources supported by multiple alarm computing nodes. Specifically, the anomaly detection system can use the following (1) to (6) to implement the above process.

[0136] (1) The forwarding node obtains at least one third alarm rule.

[0137] One of the third alarm rules includes: observation object label, indicator identifier, and alarm strategy.

[0138] (2) The forwarding node determines the amount of resources required to implement at least one third alarm rule and at least one alarm rule.

[0139] Among them, at least one third alarm rule (i.e., a newly added alarm rule) includes M observation object labels and N indicator identifiers, and at least one alarm rule (i.e., an existing alarm rule) includes m observation object labels and n indicator identifiers.

[0140] The forwarding node can determine the resource requirements for implementing at least one third alarm rule and at least one alarm rule using the following process: Specifically, the forwarding node determines M observation object labels and m matching observation object labels, as well as non-matching observation object labels. The forwarding node determines N indicator identifiers and n matching indicator identifiers, as well as non-matching indicator identifiers. The forwarding node then determines the resource requirements based on the matching observation object labels, non-matching observation object labels, matching indicator identifiers, and non-matching indicator identifiers.

[0141] (3) If the resource quantity is greater than or equal to the resource quantity supported by multiple alarm computing nodes, the forwarding node sends an expansion request to the management node.

[0142] The expansion request is used to indicate the number of alarm computing nodes to be increased or the amount of resources that multiple alarm computing nodes can support.

[0143] (4) The management node receives the expansion request from the forwarding node and sends expansion commands to multiple alarm computing nodes in response to the expansion request.

[0144] The expansion command is used to indicate the number of alarm computing nodes to be increased and / or the amount of resources that multiple alarm computing nodes can provide.

[0145] The following example illustrates the process by which an alarm computing node responds to an expansion request by sending an expansion command to the first target alarm computing node among multiple alarm computing nodes.

[0146] (5) The first target alarm computing node receives the expansion command.

[0147] (6) In response to the expansion command, the first target alarm computing node, together with other alarm computing nodes, implements at least one alarm rule implemented by the first target alarm computing node / or increases the amount of resources supported and provided by the first target alarm computing node.

[0148] In some possible scenarios, the forwarding node can provide a resource assessment page. This resource assessment page includes prompts for inputting at least one third-party alarm rule.

[0149] In some possible scenarios, forwarding nodes can employ multiple modules to implement the above evaluation process. For example, a forwarding node may use a matcher, a rule resource evaluation module, and an evaluation timer to implement the evaluation process. The management node uses an alarm rule management service module to receive expansion requests from the resource evaluation module and to send expansion commands to the expansion service module.

[0150] Figure 9 An example diagram for evaluating a new alarm rule provided in this application is shown below. Figure 9As shown in (a), when the time recorded by the evaluation timer meets the requirements, the matcher matches the newly added alarm rules and existing alarm rules to determine the matching and non-matching observation object labels included in the M observation object labels of the newly added alarm rules and the m observation object labels of the existing alarm rules, as well as the matching and non-matching indicator labels included in the N indicator identifiers of the newly added alarm rules and the n indicator identifiers of the existing alarm rules. The matcher sends the evaluation results to the rule resource evaluation module. The rule resource evaluation module receives the evaluation results and determines whether the resource amount required by the newly added alarm rules and existing alarm rules is greater than or equal to the resource amount supported by multiple alarm computing nodes. If the rule resource evaluation module determines that the resource amount required by the newly added alarm rules and existing alarm rules is greater than or equal to the resource amount supported by multiple alarm computing nodes, the rule resource evaluation module sends a scaling request to the management node. The alarm rule management service module of the management node responds to the scaling request by sending a scaling command to the scaling service module. The expansion service module receives and responds to expansion commands, performing expansion on multiple alarm computing nodes, such as increasing the number of alarm computing nodes and / or the amount of resources that the multiple alarm computing nodes can provide. The expansion service module also sends expansion feedback to the alarm rule management service module. If the expansion service module successfully increases the number of alarm computing nodes and / or the amount of resources that the multiple alarm computing nodes can provide, it can send a success feedback to the alarm rule management service module. Conversely, if the expansion service module fails to increase the number of alarm computing nodes and / or the amount of resources that the multiple alarm computing nodes can provide, it can send a failure feedback to the alarm rule management service module.

[0151] In some possible scenarios, forwarding nodes can also provide a resource assessment page. This resource assessment page includes a navigation area and an online status area. The navigation area includes navigation information to instruct the user to add at least one third-party alert rule. The online status area includes an assessment information display area and an assessment result display area. The assessment information display area displays the status of the added at least one third-party alert rule. These statuses include: resource assessment, expansion status, and successful online deployment. The assessment result display area displays assessment information for the added at least one third-party alert rule and an existing alert rule. This assessment information includes: rule name, total number of matched timelines, timelines covered by existing rules, number of new computation timelines, estimated new processor resources, estimated new memory resources, whether expansion is needed, assessment start time, assessment end time, etc. Figure 9 As shown in (b), the newly added alarm rule is for global processor monitoring. The forwarding node evaluates the newly added global processor monitoring and obtains the following results: Figure 9The evaluation results are shown in (b). These results include: "Rule Name: Processor Global Monitoring," "Total Number of Matched Timelines: 1 million," "Existing Rule Coverage Timelines: 300,000," "Number of New Computation Timelines: 700,000," "Estimated New Processor Resources: xx processors," "Estimated New Memory Resources: yy megabytes (MB)," "Expansion Required: Yes," "Evaluation Start Time: xx year x month x day x hour x minute x second," and "Evaluation End Time: xx year x month x day x hour (x+2) minute (x+15) second." The statement "Total Number of Matched Timelines: 1 million" indicates that there are 1 million processors requiring monitoring. "Existing Rule Coverage Timelines: 300,000" indicates that the existing alarm rules cover 300,000 observed objects. "Number of New Computation Timelines: 700,000" indicates that the processor global monitoring generated 700,000 observed objects. "Estimated new processor resources: xx processors" indicates that xx processors are needed to implement global monitoring of new alarm rules and existing alarm rules. "Estimated new memory resources: yyMB" indicates that yyMB of memory resources are needed to implement global monitoring of new alarm rules and existing alarm rules. "Evaluation start time: xx year x month x day x hour x minute x second" indicates that the evaluation started at xx year x month x day x hour x minute x second. "Evaluation end time: xx year x month x day x hour (x+2) minute (x+15) second" indicates that the evaluation ended at xx year x month x day x hour (x+2) minute (x+15) second.

[0152] Figure 9 It describes how, when new alarm rules need to be added, the forwarding nodes evaluate the amount of resources required for adding new alarm rules and existing alarm rules. This ensures that the amount of resources provided by multiple alarm calculation nodes is greater than or equal to the amount of resources required to implement new alarm rules and existing alarm rules, thereby ensuring the accuracy of the working status of the observed objects detected by the anomaly detection system.

[0153] The above description uses a forwarding node as an example to illustrate the anomaly detection method provided in this application. In some possible scenarios, depending on the needs of the actual application, the anomaly detection system may also include multiple forwarding nodes. Furthermore, the forwarding nodes and alarm calculation nodes can employ multiple modules to implement the functions described in the above method embodiments. For example, the forwarding node may use an indicator rule matcher, an indicator calculation routing module, and a rule resource evaluation module to implement the functions implemented by the forwarding node in the above method embodiments; the alarm calculation node may use a rule calculation manager and an indicator cache management module to implement the functions implemented by the alarm calculation node; and the management node may use an alarm rule management service module to implement the functions implemented by the management node. The process by which the forwarding node, alarm calculation node, and management node implement the functions described in the above method embodiments using the aforementioned modules is explained below. Figure 10 The relevant description. In some possible cases, the indicator rule matcher can also be simply called a matcher.

[0154] Figure 10 A flowchart illustrating an anomaly detection method provided in this application is shown below. Figure 10 As shown, the forwarding node receives multiple indicator data. The forwarding node uses an indicator rule matcher to determine the alarm rules that will use the multiple indicator data. The forwarding node also uses an indicator calculation routing module to determine a sending scheme for sending the multiple indicator data to multiple calculation nodes, and sends the multiple indicator data to the multiple alarm calculation nodes based on the determined sending scheme. An alarm calculation node (e.g., the first alarm calculation node) receives indicator data (e.g., the first indicator data) from the multiple indicator data that includes the observation object label corresponding to that alarm calculation node. The first alarm calculation node receives the first indicator data. If the rule calculation manager of the first alarm calculation node indicates that the first indicator data will be used by multiple alarm rules, the first alarm calculation node stores the first indicator data in the first storage space. The first alarm calculation node can also use an indicator cache manager to manage the first indicator data. Based on the alarm rules set thereon, the first alarm calculation node uses multiple rule calculators in the rule calculation manager to determine whether the working status of the observation object from which the first indicator data has been collected is abnormal. If the first alarm calculation node determines that the working status of the observation object is abnormal, the first alarm calculation node provides alarm information. When new alarm rules need to be added, the forwarding node can use the rule resource assessment module to evaluate the resource requirements for the new and existing alarm rules. If the assessment results determine that the resource requirements for the new and existing alarm rules are greater than or equal to the resource capacity provided by multiple alarm computing nodes, the rule resource assessment module sends a scaling request to the management node. The management node receives the scaling request using the alarm rule management service module and issues a scaling command to increase the number of alarm computing nodes and / or the resource capacity provided by multiple alarm computing nodes.

[0155] Compared to a forwarding node forwarding multiple metric data points to a single alarm computing node, this application involves the forwarding node sending a portion of the multiple metric data points to each of at least two target alarm computing nodes. This allows the at least two target alarm computing nodes to jointly process the multiple metric data points using their provided resources. This reduces the amount of metric data points that a single alarm computing node needs to process, avoids inaccurate anomaly detection results due to resource constraints of a single alarm computing node, ensures sufficient resources for alarm computing nodes to process metric data, and ultimately improves the accuracy of anomaly detection results.

[0156] It is understood that, in order to achieve the functions in the above embodiments, the computing device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0157] The above text combines Figures 1 to 10 The anomaly detection method provided according to this embodiment is described in detail below, and will be combined with Figure 11 This describes the anomaly detection device provided according to this embodiment.

[0158] Figure 11 This application provides a schematic diagram of the structure of an anomaly detection device, as shown below. Figure 11 As shown, the anomaly detection device 1100 includes a transceiver module 1110 and a processing module 1120. The anomaly detection device can be used to implement the functions of a forwarding node or a first target alarm calculation node.

[0159] when Figure 11 When the anomaly detection device shown is used to implement the function of a forwarding node, the transceiver module 1110 is used to: acquire multiple indicator data. Each indicator data includes: an observation object label, an indicator identifier, and an indicator value. The observation object label is used to indicate the observation object, and the multiple indicator data includes multiple observation object labels. The processing module 1120 is used to: determine at least two target alarm calculation nodes corresponding to the multiple observation object labels and / or the indicator identifiers of the multiple indicator data. The at least two target alarm calculation nodes are alarm calculation nodes among the multiple alarm calculation nodes that have deployed a first alarm rule. The indicator identifier of the first alarm rule matches the indicator identifier of the multiple indicator data. The observation object label of the first alarm rule includes the observation object label of the multiple indicator data. The processing module 1120 is also used to: send the multiple indicator data to the at least two target alarm calculation nodes. The at least two target alarm calculation nodes include a first target alarm calculation node and a second target alarm calculation node. The multiple indicator data includes first indicator data and second indicator data. The first indicator data has a first observation object label. The second indicator data has a second observation object label. The first indicator data is sent to the first target alarm calculation node. The second indicator data is sent to the second target alarm calculation node.

[0160] In some possible scenarios, processing module 1120 is specifically configured to: determine the first observation object label corresponding to the first target alarm calculation node and the second observation object label corresponding to the second target alarm calculation node based on multiple observation object labels and the node identifiers of at least two target alarm calculation nodes. Processing module 1120 is also specifically configured to: send multiple indicator data, including the first observation object label, to the first target alarm calculation node. Processing module 1120 is also specifically configured to: send multiple indicator data, including the second observation object label, to the second target alarm calculation node.

[0161] In some possible scenarios, processing module 1120 is further specifically configured to: generate multiple first hash values ​​based on multiple observation object tags. Each hash value corresponds to one observation object tag. Processing module 1120 is further specifically configured to: generate at least two second hash values ​​based on the node identifiers of at least two target alarm calculation nodes. Each second hash value corresponds to one target alarm calculation node. Processing module 1120 is further specifically configured to: determine the first observation object tag corresponding to the first target alarm calculation node and the second observation object tag corresponding to the second target alarm calculation node based on the multiple first hash values ​​and at least two second hash values. Wherein, the first observation object tag and the second observation object tag belong to multiple observation object tags.

[0162] In some possible scenarios, the first target alarm calculation node is also configured with a second alarm rule. If the indicator identifier of the second alarm rule matches the indicator identifier of the first indicator data, and the observation object label of the second alarm rule includes the observation object label of the first indicator data, then the first indicator data is sent to the first target alarm calculation node in a specified format.

[0163] In some possible scenarios, the specified format includes a first field and a second field. The first field is used to carry the rule identifier of the first alarm rule and the rule identifier of the second alarm rule, and the second field is used to carry the first indicator data.

[0164] In some possible scenarios, the forwarding node also communicates with the management node. The transceiver module 1110 is further configured to: acquire at least one third alarm rule. A third alarm rule includes: an observation object label, an indicator identifier, and an alarm policy. The processing module 1120 is further configured to: determine the amount of resources required to implement at least one third alarm rule, and the amount of resources required for at least one alarm rule. If the amount of resources is greater than or equal to the amount of resources supported by multiple alarm computing nodes, the processing module 1120 is further configured to: send a scaling request to the management node. The scaling request is used to indicate the increase in the number of alarm computing nodes or the amount of resources supported by multiple alarm computing nodes.

[0165] In some possible scenarios, at least one third alarm rule includes M observation object labels and N indicator identifiers. At least one alarm rule includes m observation object labels and n indicator identifiers. Processing module 1120 is specifically used to: determine the matching observation object labels and non-matching observation object labels included in the M observation object labels and m observation object labels. Processing module 1120 is also specifically used to: determine the matching indicator identifiers and non-matching indicator identifiers included in the N indicator identifiers and n indicator identifiers. Processing module 1120 is also specifically used to: determine the resource quantity based on the matching observation object labels, non-matching observation object labels, matching indicator identifiers, and non-matching indicator identifiers.

[0166] In some possible scenarios, the alarm rules are set to be the same for each of the multiple alarm calculation nodes.

[0167] In some possible scenarios, the anomaly detection device 1110 may further include a display module 1130. The display module 1130 is configured to: provide an alarm rule configuration interface. The processing module 1120 generates alarm rules in response to operations performed on the alarm rule configuration interface.

[0168] For more details on the transceiver module 1110, processing module 1120, and display module 1130, please refer to the relevant description of the forwarding node in the anomaly detection method above. It will not be repeated here.

[0169] when Figure 11 When the anomaly detection device shown is used to implement the function of a forwarding node, the transceiver module 1110 is used to: receive indicator data. The indicator data includes: an observation object label, an indicator identifier, and an indicator value. The observation object label is used to indicate the observation object, and the observation object label corresponds to the first alarm calculation node. The observation object label of the first alarm rule includes the observation object label of the indicator data, and the indicator identifier of the first alarm rule matches the indicator identifier of the indicator data. The processing module 1120 is used to: perform anomaly detection on the indicator data using the first alarm rule to obtain anomaly detection results.

[0170] In some possible scenarios, the first alarm calculation node may also have other alarm rules configured.

[0171] In some possible scenarios, other alarm rules include a second alarm rule. The second alarm rule includes: an observation object label, an indicator identifier, and an alarm policy. If the observation object label of the second alarm rule includes the observation object label of the indicator data, and the indicator identifier of the second alarm rule matches the indicator identifier of the indicator data, then the processing module 1120 is further configured to: store the indicator data in a first storage space. The first storage space supports access by a first process and a second process. The first process is the process that implements the first alarm rule, and the second process is the process that implements the second alarm rule.

[0172] In some possible scenarios, processing module 1120 is specifically used to: retrieve indicator data from the first storage space based on the first address. The first address is the address of the first storage space. Processing module 1120 is also specifically used to: perform anomaly detection on the indicator data using the first alarm rule, and obtain anomaly detection results.

[0173] In some possible scenarios, the transceiver module 1110 is further configured to: receive a capacity expansion command. The capacity expansion command is used to indicate the increase in the number of alarm computing nodes or the amount of resources that the multiple alarm computing nodes can provide. The processing module 1120 is further configured to: in response to the capacity expansion command, increase the number of alarm computing nodes and / or the amount of resources that the multiple alarm computing nodes can provide.

[0174] In some possible scenarios, the alarm rules are set to be the same for each of the multiple alarm calculation nodes.

[0175] For more details on the transceiver module 1110 and the processing module 1120, please refer to the description of the first target alarm calculation node in the anomaly detection method above. It will not be repeated here.

[0176] When the anomaly detection device 1100 corresponds to the steps performed by the forwarding node and the first target alarm calculation node in the anomaly detection method described in the embodiments of this application, the above and other operations and / or functions of each module in the anomaly detection device 1100 are respectively to implement the method flow performed by the computing device in the foregoing figures.

[0177] It is worth noting that if the above-mentioned anomaly detection devices are implemented through software modules, for example, the software modules can be provided to users through a cloud service subscription model, and users can choose different subscription levels according to their needs; or, for example, the software modules can also provide enterprise-level customized services with professional domain customization, interface personalization and extended functions according to the needs of users or enterprises.

[0178] In addition, the anomaly detection device 1100 provided in this application can also be provided to users as an value-added service, and this application does not limit this.

[0179] The anomaly detection device in this application embodiment can also be implemented in hardware, such as a computing device, chip, or processor. For specific implementation details regarding computing devices, please refer to [reference needed]. Figure 2 For a description of the chip and processor's specific implementation, please refer to [link / reference]. Figure 3 The description of that will not be repeated here.

[0180] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a computing device. Of course, the processor and storage medium can also exist as discrete components in a network device or terminal device.

[0181] This application also provides a computing device cluster. The computing device cluster includes at least one computing device, which may be a server with a display. In some embodiments, the computing device may also be a terminal device such as a desktop computer, laptop computer, or smartphone.

[0182] like Figure 12 As shown, Figure 12 This application provides a schematic diagram of a computing device cluster, which includes at least one computing device 200. The memory 212 of one or more computing devices 200 in the computing device cluster may store the same instructions for executing anomaly detection methods.

[0183] In some possible implementations, the memory 212 of one or more computing devices 200 in the computing device cluster may also store partial instructions for executing the anomaly detection method. In other words, a combination of one or more computing devices 200 can jointly execute the instructions for executing the anomaly detection method.

[0184] It should be noted that the memory 212 in different computing devices 200 within the computing device cluster can store different instructions, each used to execute a portion of the computing device's functions. That is, the instructions stored in the memory 212 of different computing devices 200 can implement the functions of one or more units in the transceiver module 1110, processing module 1120, and display module 1130.

[0185] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN). Figure 13 One possible implementation is shown. For example... Figure 13 As shown, Figure 13 This application provides a schematic diagram of a connection between computing devices, where two computing devices 200A and 200B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the instructions stored in the memory 212 of computing device 200A can implement the functions of the transceiver module 1110 and the display module 1130. Simultaneously, the instructions stored in the memory 212 of computing device 200B can implement the functions of the processing module 1120.

[0186] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform an anomaly detection method.

[0187] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform an anomaly detection method.

[0188] This application also provides a chip. The chip includes an interface circuit and a control circuit. The interface circuit is used to acquire indicator data, and the control circuit is used to implement the functions of a forwarding node or a first target alarm calculation node in the anomaly detection method.

[0189] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).

[0190] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An anomaly detection method, characterized in that, The method is executed by a forwarding node, which is communicatively connected to multiple alarm calculation nodes. Each of the multiple alarm calculation nodes has at least one alarm rule configured, and an alarm rule is configured on at least two alarm calculation nodes. An alarm rule includes: an indicator identifier, an observed object label, and an alarm policy. The method includes: Acquire multiple indicator data; each indicator data includes: observation object label, indicator identifier, and indicator value, wherein the observation object label is used to indicate the observation object, and the multiple indicator data includes multiple observation object labels; Determine at least two target alarm calculation nodes corresponding to the multiple observation object labels and / or the indicator identifiers of the multiple indicator data; Wherein, the at least two target alarm calculation nodes are alarm calculation nodes among the plurality of alarm calculation nodes that have deployed the first alarm rule, the indicator identifier of the first alarm rule matches the indicator identifier of the plurality of indicator data, and the observation object label of the first alarm rule includes the observation object label of the plurality of indicator data; Send the plurality of indicator data to the at least two target alarm calculation nodes; The at least two target alarm calculation nodes include a first target alarm calculation node and a second target alarm calculation node. The multiple indicator data include first indicator data and second indicator data. The first indicator data has a first observation object label, and the second indicator data has a second observation object label. The first indicator data is sent to the first target alarm calculation node, and the second indicator data is sent to the second target alarm calculation node.

2. The method according to claim 1, characterized in that, Sending the multiple indicator data to the at least two target alarm calculation nodes includes: Based on the multiple observation object labels and the node identifiers of the at least two target alarm calculation nodes, determine the first observation object label corresponding to the first target alarm calculation node and the second observation object label corresponding to the second target alarm calculation node; Send the index data, which includes the label of the first observed object, among the multiple index data to the first target alarm calculation node; Send the index data, which includes the label of the second observed object, from the multiple index data to the second target alarm calculation node.

3. The method according to claim 2, characterized in that, The step of determining the first observation object label corresponding to the first target alarm calculation node and the second observation object label corresponding to the second target alarm calculation node based on the plurality of observation object labels and the node identifiers of the at least two target alarm calculation nodes includes: Based on the multiple observation object labels, multiple first hash values ​​are generated; each hash value corresponds to one observation object label. Based on the node identifiers of the at least two target alarm calculation nodes, at least two second hash values ​​are generated; one second hash value corresponds to one target alarm calculation node. Based on the plurality of first hash values ​​and the at least two second hash values, determine the first observation object label corresponding to the first target alarm calculation node and the second observation object label corresponding to the second target alarm calculation node; wherein, the first observation object label and the second observation object label belong to the plurality of observation object labels.

4. The method according to any one of claims 1-3, characterized in that, The first target alarm calculation node is also configured with a second alarm rule. Sending the multiple indicator data to the at least two target alarm calculation nodes includes: If the indicator identifier of the second alarm rule matches the indicator identifier of the first indicator data, and the observation object label of the second alarm rule includes the observation object label of the first indicator data, then the first indicator data is sent to the first target alarm calculation node in the specified format.

5. The method according to claim 4, characterized in that, The specified format includes a first field and a second field. The first field is used to carry the rule identifier of the first alarm rule and the rule identifier of the second alarm rule, and the second field is used to carry the first indicator data.

6. The method according to any one of claims 1-5, characterized in that, The forwarding node also has a communication connection with the management node, and the method further includes: Obtain at least one third alarm rule; a third alarm rule includes: observation object label, indicator identifier, and alarm policy; Determine the amount of resources required to implement the at least one third alarm rule, and the amount of resources required for the at least one alarm rule; If the amount of resources is greater than or equal to the amount of resources that the plurality of alarm computing nodes can provide, a scaling request is sent to the management node; the scaling request is used to indicate increasing the number of alarm computing nodes or the amount of resources that the plurality of alarm computing nodes can provide.

7. The method according to claim 6, characterized in that, The at least one third alarm rule includes M observation object labels and N indicator identifiers, and the at least one alarm rule includes m observation object labels and n indicator identifiers. Determining the amount of resources required to implement the at least one third alarm rule, and the at least one alarm rule, includes: Determine the matching observation object labels and the non-matching observation object labels included in the M observation object labels and the m observation object labels; Determine the N indicator identifiers and the matching indicator identifiers and non-matching indicator identifiers included in the N indicator identifiers; The resource quantity is determined based on the matching observation object labels, the non-matching observation object labels, the matching indicator identifiers, and the non-matching indicator identifiers.

8. The method according to any one of claims 1-7, characterized in that, The alarm rules are set to be the same for each of the multiple alarm calculation nodes.

9. An anomaly detection method, characterized in that, The method is executed by a first alarm calculation node among multiple alarm calculation nodes. The first alarm calculation node is configured with a first alarm rule, which includes: an observation object label, an indicator identifier, and an alarm strategy. The method includes: Receive indicator data; the indicator data includes: observation object label, indicator identifier, and indicator value. The observation object label is used to indicate the observation object. The observation object label corresponds to the first alarm calculation node. The observation object label of the first alarm rule includes the observation object label of the indicator data. The indicator identifier of the first alarm rule matches the indicator identifier of the indicator data. Using the first alarm rule, anomaly detection is performed on the indicator data to obtain anomaly detection results.

10. The method according to claim 9, characterized in that, The first alarm calculation node is also configured with other alarm rules.

11. The method according to claim 10, characterized in that, The other alarm rules include a second alarm rule, which includes: an observation object label, an indicator identifier, and an alarm strategy. The method further includes: If the observation object label of the second alarm rule includes the observation object label of the indicator data, and the indicator identifier of the second alarm rule matches the indicator identifier of the indicator data, then the indicator data is stored in the first storage space. The first storage space supports access by a first process and a second process. The first process is the process that implements the first alarm rule, and the second process is the process that implements the second alarm rule.

12. The method according to claim 11, characterized in that, The step of using the first alarm rule to perform anomaly detection on the indicator data and obtaining anomaly detection results includes: The indicator data is obtained from the first storage space according to the first address; the first address is the address of the first storage space; The first alarm rule is used to perform anomaly detection on the indicator data to obtain anomaly detection results.

13. The method according to any one of claims 9-12, characterized in that, The method further includes: Receive an expansion command; the expansion command is used to instruct the addition of a second alarm computing node to the plurality of alarm computing nodes and / or increase the amount of resources supported by at least two of the plurality of alarm computing nodes; In response to the expansion command, the number of alarm computing nodes and / or the amount of resources supported by multiple alarm computing nodes are increased.

14. The method according to any one of claims 9-13, characterized in that, The alarm rules are set to be the same for each of the multiple alarm calculation nodes.

15. An anomaly detection device, characterized in that, The device includes: The transceiver module is used to: acquire multiple indicator data; each indicator data includes: an observation object label, an indicator identifier, and an indicator value, wherein the observation object label is used to indicate the observation object, and the multiple indicator data includes multiple observation object labels; The processing module is used to: determine at least two target alarm calculation nodes corresponding to the multiple observation object labels and / or the indicator identifiers of the multiple indicator data; Wherein, the at least two target alarm calculation nodes are alarm calculation nodes among the plurality of alarm calculation nodes that have deployed the first alarm rule, the indicator identifier of the first alarm rule matches the indicator identifier of the plurality of indicator data, and the observation object label of the first alarm rule includes the observation object label of the plurality of indicator data; The processing module is further configured to: send the plurality of indicator data to the at least two target alarm calculation nodes; The at least two target alarm calculation nodes include a first target alarm calculation node and a second target alarm calculation node. The multiple indicator data include first indicator data and second indicator data. The first indicator data has a first observation object label, and the second indicator data has a second observation object label. The first indicator data is sent to the first target alarm calculation node, and the second indicator data is sent to the second target alarm calculation node.

16. An anomaly detection device, characterized in that, The device includes: The transceiver module is used to: receive indicator data; the indicator data includes: observation object label, indicator identifier, and indicator value, the observation object label is used to indicate the observation object, the observation object label corresponds to the first alarm calculation node, the observation object label of the first alarm rule includes the observation object label of the indicator data, and the indicator identifier of the first alarm rule matches the indicator identifier of the indicator data. The processing module is used to: perform anomaly detection on the indicator data using the first alarm rule, and obtain anomaly detection results.

17. A processor, characterized in that, The processor includes an interface circuit and a control circuit; the interface circuit is used to acquire index data and, in cooperation with the control circuit, execute the method of any one of claims 1-8, or, in cooperation with the control circuit, execute the method of any one of claims 9-14.

18. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, and each computing device includes a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-8, or to cause the cluster of computing devices to perform the method as described in any one of claims 9-14.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions; when the computer instructions are executed in a computing device, the computing device performs the method of any one of claims 1-8, or when the computer instructions are executed in a computing device, the computing device performs the method of any one of claims 9-14.

20. A computer program product, characterized in that, When the computer program product is run in a computing device, the computing device performs the method of any one of claims 1-8, or the computing device performs the method of any one of claims 9-14.