Object index anomaly monitoring method and device, electronic equipment and storage medium

By applying the Bayesian algorithm and the cross-validation method of dynamic monitoring resource allocation weights in the cloud platform, the accuracy and timeliness of abnormal monitoring of virtual machines or containers in the cloud management platform are solved, and efficient anomaly detection is achieved.

CN120763002APending Publication Date: 2025-10-10JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510897081.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

When monitoring virtual machines or containers, existing cloud management platforms have problems such as incomplete indicators, high missed detection rates, and lack of correlation in indicator collection, resulting in the inability to detect anomalies in a timely manner and low accuracy.

Method used

Based on the historical operating indicator data of multiple objects, the Bayesian algorithm is used to determine the prior probability and conditional probability of abnormal events of the indicators. Combined with the dynamic monitoring resource allocation weights, cross-validation is achieved to ensure comprehensive testing of key indicators and timely detection of anomalies.

Benefits of technology

It improves the accuracy of determining virtual machine or container anomalies, reduces the missed detection rate, shortens the abnormal response delay, and improves the monitoring efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763002A_ABST
    Figure CN120763002A_ABST
Patent Text Reader

Abstract

The invention discloses an object index abnormity monitoring method and device, electronic equipment and a storage medium, and relates to the technical field of servers, and the method comprises the following steps: according to historical operation index data corresponding to each index corresponding to each object in a plurality of objects and abnormal data in the historical operation index data corresponding to each index, determining the abnormal data in the historical operation index data corresponding to each index; determining a historical abnormal event prior probability corresponding to each index and abnormal condition probabilities corresponding to other indexes except the jth index in the plurality of indexes, and further determining a historical abnormal event posterior probability of the jth index; and when the historical abnormal event posterior probability of the jth index is greater than a preset probability anomaly threshold corresponding to the jth index, determining that the jth index of the ith object is abnormal. According to the method, the abnormality of the virtual machine or the container can be found in time, the accuracy of determining the abnormality of the virtual machine or the container is improved, the missing test rate is reduced, and the abnormality response delay of the whole system is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server technology, and in particular to a method, device, electronic device, and storage medium for monitoring abnormalities in object indicators. Background Art

[0002] With the development of big data, virtualization technology and containerization technology, cloud management platforms have integrated multiple modules to uniformly schedule and manage the underlying computing, storage, network, security and other resources of multiple virtual machines or multiple containers, realizing dynamic business changes and intelligent resource management. By effectively monitoring and flexibly scheduling large-scale hardware resources, they ensure the security and reliability of user data, greatly improve IT operation and maintenance efficiency while improving the utilization rate of the entire data center resources, and reduce the maintenance cost of the data center.

[0003] The monitoring module in existing cloud management platforms tracks the behavior of each virtual machine or container. Therefore, it needs to collect metrics from other modules within the cloud management platform. However, if a single module modifies data, some metrics may be missed, resulting in an incomplete and high under-detection rate. Furthermore, the collection of metrics by a single module and their transmission to the monitoring module can lead to a lack of correlation between metrics collected across modules, making it difficult to detect virtual machine or container anomalies in a timely manner and resulting in low accuracy in identifying virtual machine or container anomalies. Summary of the Invention

[0004] The present application provides a method, device, electronic device and storage medium for monitoring object indicator anomalies, so as to at least solve the problem in the related art that virtual machine or container anomalies cannot be detected in time, or the accuracy of determining virtual machine or container anomalies is low.

[0005] The present application provides an object indicator abnormality monitoring method, which is applied to a cloud platform, and the cloud platform is used to monitor multiple operating indicator data of each object in a plurality of objects. The method includes: determining the prior probability of historical abnormal events corresponding to each indicator based on the historical operating indicator data corresponding to each indicator corresponding to each of the multiple objects, and the abnormal data in the historical operating indicator data corresponding to each indicator; determining the prior probability of historical abnormal events corresponding to the jth indicator of the i-th object based on the prior probability of historical abnormal events corresponding to the jth indicator, the historical operating indicator data corresponding to the hth indicator other than the jth indicator in the multiple indicators, and the abnormal data in the historical operating indicator data; The abnormal conditional probability that an abnormality exists in the indicator and the j-th indicator is abnormal at the same time, where the h-th indicator is an indicator other than the j-th indicator among the multiple indicators; after obtaining the abnormal conditional probabilities corresponding to the other indicators among the multiple indicators except the j-th indicator, determine the posterior probability of the historical abnormal event of the j-th indicator according to the prior probability of the historical abnormal event corresponding to each indicator and the abnormal conditional probabilities corresponding to the other indicators among the multiple indicators except the j-th indicator; when it is determined that the posterior probability of the historical abnormal event of the j-th indicator is greater than the preset probability abnormal threshold corresponding to the j-th indicator, it is determined that the j-th indicator of the ith object is abnormal.

[0006] The present application also provides an object indicator abnormality monitoring device, which is applied to a cloud platform, and the cloud platform is used to monitor multiple operating indicator data of each object in a plurality of objects. The device includes: a priori probability determination module, which is used to determine the priori probability of historical abnormal events corresponding to each indicator based on the historical operating indicator data corresponding to each indicator of each object in the plurality of objects, and the abnormal data in the historical operating indicator data corresponding to each indicator; a conditional probability determination module, which is also used to determine the priori probability of historical abnormal events corresponding to the jth indicator of the ith object based on the priori probability of historical abnormal events corresponding to the jth indicator of the ith object, the historical operating indicator data corresponding to the hth indicator other than the jth indicator in the plurality of indicators, and the abnormal data in the historical operating indicator data. The abnormal conditional probability that an h-th indicator is abnormal and an abnormal condition occurs in the j-th indicator at the same time, wherein the h-th indicator is an indicator other than the j-th indicator among the multiple indicators; the posterior probability determination module is further used to, after obtaining the abnormal conditional probabilities corresponding to the other indicators among the multiple indicators except the j-th indicator, determine the posterior probability of the historical abnormal event of the j-th indicator according to the prior probability of historical abnormal events corresponding to each indicator and the abnormal conditional probabilities corresponding to the other indicators among the multiple indicators except the j-th indicator; the processing module is used to determine that the j-th indicator of the ith object is abnormal when it is determined that the posterior probability of the historical abnormal event of the j-th indicator is greater than the preset probability abnormal threshold corresponding to the j-th indicator.

[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned object indicator abnormality monitoring methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned object indicator abnormality monitoring methods are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned object indicator abnormality monitoring methods when executed by a processor.

[0010] Through the present application, based on the historical operating indicator data corresponding to each indicator corresponding to each object in a plurality of objects, and the abnormal data in the historical operating indicator data corresponding to each indicator, the prior probability of historical abnormal events corresponding to each indicator and the abnormal conditional probabilities corresponding to other indicators in a plurality of indicators except the j-th indicator are determined, and then the posterior probability of historical abnormal events of the j-th indicator is determined; and when the posterior probability of historical abnormal events of the j-th indicator is greater than the preset probability abnormality threshold corresponding to the j-th indicator, it is determined that the j-th indicator of the ith object is abnormal.

[0011] When determining whether the j-th indicator of the ith object is abnormal, the j-th indicator of the ith object is not considered alone, but is determined by cross-validation of the historical operating indicator data corresponding to each indicator of each object in multiple objects and the abnormal data. Therefore, this method can solve the problem of relying on a single indicator and a fixed threshold for alarm, which makes it difficult to timely detect cross-anomalies between indicators. Based on the correlation between the indicators, virtual machine or container anomalies can be discovered in a timely manner, thereby improving the accuracy of determining virtual machine or container anomalies, reducing the missed detection rate, and shortening the abnormal response delay of the entire system. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 A topological diagram of the object indicator anomaly monitoring system provided in an embodiment of the present application;

[0014] Figure 2 A flowchart of a method for monitoring abnormal object indicators provided in an embodiment of the present application;

[0015] Figure 3 A flowchart of another method for monitoring abnormal object indicators provided in an embodiment of the present application;

[0016] Figure 4 A structural block diagram of a device for monitoring abnormal object indicators provided in an embodiment of the present application;

[0017] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0020] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0021] The embodiment of the present application is applied to a scenario where a cloud platform monitors multiple operating indicator data of each of multiple objects, wherein the objects may be virtual machines or containers.

[0022] In the related art, when monitoring the operating indicator data of multiple objects, current cloud platforms mainly use a single indicator that is required to be collected by each module in advance to monitor multiple objects. As a result, the monitoring data of other modules obtained by the monitoring module is incomplete, resulting in the omission of some functional indicators. In addition, the indicator collection of each module lacks correlation, which is prone to omissions due to interface incompatibility or communication delays; relying on a single indicator and a fixed threshold for alarms makes it difficult to detect cross-indicator anomalies in a timely manner; and poor adaptability in dynamic scenarios. For example, when switching between day and night modes, the adjustment of monitoring parameters may omit key indicator items, resulting in the inability to detect virtual machine or container anomalies in a timely manner, or low accuracy in determining virtual machine or container anomalies.

[0023] In order to solve the above problems, this application proposes a method for monitoring object indicator anomalies, which is based on dynamically determining the monitoring resource allocation weight of each object and the indicator cross-validation mechanism to ensure that all situations of key indicators are effectively tested, so that virtual machine or container anomalies can be discovered in time, and the accuracy of determining virtual machine or container anomalies can be improved.

[0024] Below is Figure 1 Taking the object indicator abnormality monitoring system shown as an example, the method provided in the embodiment of the present application is described. Figure 1 It is only a schematic diagram and does not constitute a limitation on the applicable scenarios of the technical solution provided in this application.

[0025] like Figure 1 As shown, Figure 1 This is a topology diagram of the object indicator anomaly monitoring system provided in an embodiment of the present application. Figure 1 In the embodiment, the object indicator abnormality monitoring system 100 may include a cloud platform 101 , a first object 102 and a second object 103 .

[0026] The cloud platform 101 can be any device with computing and communication capabilities, such as a cloud server, a tower server, a rack server, or a blade server.

[0027] Cloud platform 101 may include a monitoring module and multiple other modules. The monitoring module is used to obtain operational indicator data corresponding to multiple indicators collected by multiple other modules. The multiple other modules may include computing modules, network modules, and storage modules. The computing module is used to collect processor utilization of the object; the network module is used to collect network latency of the object; and the storage module is used to collect storage input and output latency of the object.

[0028] In the embodiment of the present application, the first object 102 and the second object 103 may be virtual machines or containers.

[0029] Figure 1 The object indicator anomaly monitoring system 100 shown is for example only and is not intended to limit the technical solution of this application. Those skilled in the art should understand that in the specific implementation process, the object indicator anomaly monitoring system 100 may also include other virtual machines or containers without limitation.

[0030] The embodiment of the present application provides a method for monitoring abnormality of object indicators, which is applied to Figure 1 The cloud platform shown, such as Figure 2 As shown, Figure 2 The following is a flow chart of a method for monitoring abnormal object indicators provided in an embodiment of the present application, wherein the method comprises the following steps:

[0031] S201, determining, according to historical running index data corresponding to each index of each object in the plurality of objects and abnormal data in the historical running index data corresponding to each index, a historical abnormal event prior probability corresponding to each index.

[0032] The index can be a monitoring data index of an infrastructure layer (such as a processor and a memory), a platform layer (such as a container state), or an application layer (such as an interface call) of a virtual machine or a container. For example, the index can be a processor utilization rate, a memory usage rate, a workload, a storage usage rate, a network delay, or a storage input / output delay. Optionally, the index can also include other more indexes of the virtual machine or the container, which are not limited.

[0033] The abnormal data is running index data of the index that is greater than a threshold value corresponding to the index. For example, the threshold value corresponding to the processor utilization rate is 85%. When the running index data of the processor utilization rate is 95%, the running index data of the processor utilization rate is abnormal data.

[0034] The historical abnormal event prior probability corresponding to each index is a probability of abnormal data in the historical running index data corresponding to each index in a historical time period.

[0035] In an example, the abnormal data in the historical running index data corresponding to each index and a ratio between the historical running index data corresponding to each index are calculated to determine the historical abnormal event prior probability corresponding to each index.

[0036] Optionally, before S201 is performed, the cloud platform can integrate a computing module, a network module, and a storage module through a unified interface protocol (such as a wireless / wired communication protocol) to collect running index data of multi-dimensional indexes in real time. For example, the cloud platform can collect running index data of indexes of a processor, a memory, a disk, and the like of a system through an agent (Agent) or an application programming interface (Application Programming Interface, API). Optionally, the cloud platform can also integrate a software development kit (Software Development Kit, SDK) or expose a standard interface (such as / metrics), or support a hyper text transfer protocol (HyperText Transfer Protocol, HTTP) or a remote procedure call protocol (Google Remote Procedure Call Protocol, GRPC) protocol to collect running index data of indexes of an application layer.

[0037] It can be understood that the running index data of the multi-dimensional indexes is collected in real time through the unified interface protocol, which can avoid missing the indexes of the system, so that the indexes obtained are more comprehensive and the missing rate is reduced.

[0038] S202, based on the prior probability of historical abnormal events corresponding to the j-th indicator of the i-th object and the historical operating indicator data corresponding to the h-th indicator except the j-th indicator among multiple indicators and the abnormal data in the historical operating indicator data, determine the abnormal conditional probability that the h-th indicator is abnormal and the j-th indicator is abnormal at the same time.

[0039] Among them, the hth indicator is the indicator other than the jth indicator among the multiple indicators.

[0040] In one example, the ratio of abnormal data in the historical operating indicator data corresponding to the h-th indicator and the historical operating indicator data is calculated, and the product of the ratio and the prior probability of historical abnormal events corresponding to the j-th indicator of the ith object is calculated to determine the abnormal conditional probability that the h-th indicator is abnormal and the j-th indicator is abnormal at the same time.

[0041] S203, after obtaining the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among the multiple indicators, determine the posterior probability of the historical abnormal event of the j-th indicator according to the prior probability of the historical abnormal event corresponding to each indicator and the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among the multiple indicators.

[0042] The posterior probability of a historical abnormal event for the jth indicator is the probability that the jth indicator is abnormal and all other indicators except the jth indicator are abnormal at the same time.

[0043] In one example, after obtaining the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among multiple indicators, based on the Bayesian algorithm, the posterior probability of the historical abnormal event of the j-th indicator is determined according to the prior probability of the historical abnormal event corresponding to each indicator and the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among multiple indicators.

[0044] The posterior probability of historical abnormal events for the jth indicator is determined based on the Bayesian algorithm and is expressed as follows:

[0045] P(A|X)=P(X|A)*P(A) / P(X)

[0046] Where P(A|X) is the posterior probability of a historical abnormal event for the jth indicator. P(X|A) is the product of the conditional probabilities of abnormal events for each of the multiple indicators except the jth indicator. P(A) is the prior probability of a historical abnormal event for the jth indicator. P(X) is the product of the prior probabilities of abnormal events for each of the multiple indicators except the jth indicator.

[0047] S204 , when it is determined that the posterior probability of the historical abnormal event of the j-th indicator is greater than the preset probability abnormality threshold corresponding to the j-th indicator, it is determined that the j-th indicator of the i-th object is abnormal.

[0048] The preset probability abnormality threshold may be set according to actual conditions. For example, the preset probability abnormality threshold may be 0.2.

[0049] Based on the above Figure 2 The method shown determines the prior probability of historical abnormal events corresponding to each indicator and the abnormal conditional probabilities corresponding to each indicator except the j-th indicator among the multiple indicators based on the historical operating indicator data corresponding to each indicator of each object in the multiple objects and the abnormal data in the historical operating indicator data corresponding to each indicator, and then determines the posterior probability of historical abnormal events for the j-th indicator; and when the posterior probability of historical abnormal events for the j-th indicator is greater than the preset probability abnormality threshold corresponding to the j-th indicator, it is determined that the j-th indicator of the i-th object is abnormal.

[0050] When determining whether the j-th indicator of the ith object is abnormal, the j-th indicator of the ith object is not considered alone, but is determined by cross-validation of the historical operating indicator data corresponding to each indicator of each object in multiple objects and the abnormal data. Therefore, this method can solve the problem of relying on a single indicator and a fixed threshold for alarm, which makes it difficult to timely detect cross-anomalies between indicators. Based on the correlation between the indicators, virtual machine or container anomalies can be discovered in a timely manner, thereby improving the accuracy of determining virtual machine or container anomalies, reducing the missed detection rate, and shortening the abnormal response delay of the entire system.

[0051] Furthermore, after determining that the jth indicator of the i-th object is abnormal, the embodiment of the present application provides another object indicator abnormality monitoring method, such as Figure 3 As shown, Figure 3 A flowchart of another method for monitoring abnormal object indicators provided in an embodiment of the present application is provided. The method for monitoring abnormal object indicators includes the following steps:

[0052] S301 : Determine a monitoring resource allocation weight of the q th object according to historical operating indicator data corresponding to each indicator corresponding to the q th object.

[0053] The qth object is any object among the multiple objects except the i-th object.

[0054] Among them, the object with a larger monitoring resource allocation weight has a higher priority in the cross-validation.

[0055] In some optional embodiments, when the indicators include processor utilization, network delay, and storage input / output delay, the monitoring resource allocation weight of the qth object is determined according to historical running indicator data corresponding to the indicators corresponding to the qth object, and is expressed by the following expression:

[0056] W q =(α*A q +β*B q ) / (γ*D q )

[0057] wherein W q is the monitoring resource allocation weight of the qth object; α is a first adjustment coefficient corresponding to the processor utilization; β is a second adjustment coefficient corresponding to the network delay; γ is a third adjustment coefficient corresponding to the storage input / output delay; A q is the historical running indicator data of the processor utilization of the qth object; B q is the historical running indicator data of the network delay of the qth object; and D q is the historical running indicator data of the storage input / output delay of the qth object.

[0058] It can be understood that the greater the monitoring resource allocation weight of an object, the more monitoring resources required by the object, i.e., the greater the risk of failure of the object, and the object needs to be monitored in priority. Therefore, determining the monitoring resource allocation weight of each object can reduce the probability of missing measurement.

[0059] S302, obtaining a plurality of actual running indicator data respectively corresponding to each object in the plurality of objects except the ith object.

[0060] For example, the cloud platform obtains the plurality of actual running indicator data respectively corresponding to each object in the plurality of objects except the ith object according to an acquisition order respectively corresponding to each object in the plurality of objects except the ith object.

[0061] In one example, before obtaining the plurality of actual running indicator data respectively corresponding to each object in the plurality of objects except the ith object, the cloud platform can generate an order sequence of the plurality of objects according to a descending order of the monitoring resource allocation weight respectively corresponding to each object in the plurality of objects except the ith object.

[0062] The order sequence is used to indicate an acquisition order of the plurality of actual running indicator data respectively corresponding to each object in the plurality of objects except the ith object.

[0063] It can be understood that by obtaining multiple actual operation indicator data corresponding to multiple objects except the i-th object in sorted order, the indicator data of objects with high failure risks can be obtained first, thereby reducing the failure risk of the entire system and reducing the missed detection rate.

[0064] S303, after obtaining the monitoring resource allocation weights and multiple actual operation indicator data corresponding to each of the multiple objects except the i-th object, determine the compensation error rate of the i-th object based on the historical operation indicator data corresponding to the j-th indicator corresponding to the i-th object, the monitoring resource allocation weights and multiple actual operation indicator data corresponding to each of the multiple objects except the i-th object.

[0065] The compensation error rate is used to indicate the failure risk of the i-th object.

[0066] Determine the compensation error rate of the i-th object, which is expressed by the following expression:

[0067]

[0068] Among them, S i is the compensation error rate of the i-th object; W k Assign weights to the kth monitoring resources corresponding to each of the multiple objects except the i-th object; P k is the kth actual operating index data corresponding to each of the multiple objects except the i-th object; δ is the compensation coefficient; Q i The historical operating indicator data corresponding to the j-th indicator of the i-th object.

[0069] Optional, W k It can also be the kth weight corresponding to each of the objects except the i-th object among the multiple objects distributed in inverse proportion based on the Euclidean distance. i It can also be the historical average operating indicator data corresponding to the j-th indicator corresponding to the i-th object.

[0070] S304 : Determine whether to perform a task migration operation on the task corresponding to the jth indicator in the i-th object according to the compensation error rate.

[0071] In one example, the cloud platform detects whether the compensation error rate is greater than a preset compensation error rate; if it is detected that the compensation error rate is greater than the preset compensation error rate, it determines to perform a task migration operation on the task corresponding to the jth indicator in the i-th object.

[0072] The preset compensation error rate may be set according to actual needs. For example, the preset compensation error rate may be 5%.

[0073] As can be understood, if the compensation error rate is detected to be less than the preset compensation error rate, it is determined that the characteristics of the neighboring objects are highly correlated with those of the i-th object, indicating a low probability of failure. No operation is required on the i-th object, ensuring monitoring continuity. If the compensation error rate is detected to be greater than the preset compensation error rate, indicating a high probability of failure, it is necessary to perform a task migration operation on the task corresponding to the j-th indicator in the i-th object. The compensation error rate is calculated by constructing a joint analysis model algorithm based on multi-source data, comparing the deviation between the actual operating parameters and the fused indicators, and iteratively optimizing the monitoring strategy.

[0074] Optionally, when each indicator includes processor utilization, network delay, and storage input and output delay, obtain the first historical operating indicator data corresponding to the processor utilization of the i-th object, the second historical operating indicator data corresponding to the network delay, and the third historical operating indicator data corresponding to the storage input and output delay; normalize the first historical operating indicator data, the second historical operating indicator data, and the third historical operating indicator data respectively to obtain first data, second data, and third data; obtain the first weight corresponding to the processor utilization, the second weight corresponding to the network delay, and the third weight corresponding to the storage input and output delay, and determine the score value of the i-th object based on the first data, the second data, the third data, the first weight, the second weight, and the third weight.

[0075] The sum of the first weight, the second weight, and the third weight is 1.

[0076] In one example, a first product between the first data and the first weight, a second product between the second data and the second weight, and a third product between the third data and the third weight are calculated; and the sum of the first product, the second product, and the third product is determined as the score value of the i-th object.

[0077] It can be understood that if the score value of the i-th object is less than the preset score, the tasks of other objects can be migrated to the i-th object; if the score value of the i-th object is greater than or equal to the preset score, the task migration operation is performed on the i-th object to reduce the probability of failure of the i-th object.

[0078] In some optional implementations, the cloud platform determines the risk level corresponding to the posterior probability of historical abnormal events of the jth indicator based on the posterior probability of historical abnormal events of the jth indicator; and executes a solution strategy corresponding to the risk level based on the risk level.

[0079] Among them, risk levels include warning level, abnormal level and high-risk level.

[0080] An example, when the jth index of the historical abnormal event posterior probability is less than the first threshold value, the jth index of the historical abnormal event posterior probability corresponding to the risk level is determined as the warning level; according to the risk level, the solution strategy corresponding to the warning level is executed: automatically triggering the object to perform self-checking and synchronizing data to other objects.

[0081] An example, when the jth index of the historical abnormal event posterior probability is greater than or equal to the first threshold value and less than the second threshold value, the jth index of the historical abnormal event posterior probability corresponding to the risk level is determined as the abnormal level; according to the risk level, the solution strategy corresponding to the abnormal level is executed: automatically triggering the object to perform self-checking and synchronizing data to other objects.

[0082] An example, when the jth index of the historical abnormal event posterior probability is greater than the second threshold value, the jth index of the historical abnormal event posterior probability corresponding to the risk level is determined as the high-risk level; according to the risk level, the solution strategy corresponding to the high-risk level is executed: starting the redundant object to take over the task corresponding to the jth index, and generating a complete working video of the device through image fusion technology to assist manual intervention.

[0083] Among them, the first threshold value and the second threshold value can be set according to the actual situation. For example, the first threshold value is 0.2, and the second threshold value is 0.8.

[0084] It can be understood that, based on the method of Figure 2 and Figure 3 , through the unified interface protocol integration of the computing module, the network module, and the storage module of the system, multi-dimensional index data is collected in real time; a dynamic weight distribution algorithm is designed, and the priority of each object is adjusted according to the running index data of the object, so that all situations of all indexes are effectively tested. Compared with the 12%-15% missing rate in the related art, the missing rate of the present application is reduced to less than 4%; the abnormal response delay is reduced to 0.5-2 seconds.

[0085] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.

[0086] The embodiments of the present application also provide an object index abnormality monitoring device, applied to a cloud platform, the cloud platform is used for monitoring a plurality of running index data of each object in a plurality of objects, such as Figure 4 as shown, Figure 4 is a structural block diagram of an object index abnormality monitoring device provided by the embodiments of the present application; the object index abnormality monitoring device comprises:

[0087] The prior probability determination module 401 is used to determine the prior probability of historical abnormal events corresponding to each indicator based on the historical operating indicator data corresponding to each indicator of each object in the multiple objects and the abnormal data in the historical operating indicator data corresponding to each indicator.

[0088] The conditional probability determination module 402 is used to determine the abnormal conditional probability that the h-th indicator is abnormal and the j-th indicator is abnormal at the same time based on the prior probability of historical abnormal events corresponding to the j-th indicator of the i-th object, the historical operating indicator data corresponding to the h-th indicator other than the j-th indicator among multiple indicators, and the abnormal data in the historical operating indicator data, wherein the h-th indicator is an indicator other than the j-th indicator among multiple indicators.

[0089] The posterior probability determination module 403 is used to determine the posterior probability of historical abnormal events of the jth indicator based on the prior probability of historical abnormal events corresponding to each indicator and the abnormal conditional probabilities corresponding to the other indicators except the jth indicator among the multiple indicators after obtaining the abnormal conditional probabilities corresponding to the other indicators among the multiple indicators.

[0090] The processing module 404 is configured to determine that an abnormality exists in the jth indicator of the ith object when it is determined that the posterior probability of a historical abnormal event of the jth indicator is greater than a preset probability abnormality threshold corresponding to the jth indicator.

[0091] In some optional embodiments, when it is determined that the posterior probability of a historical abnormal event of the jth indicator is greater than a preset probability abnormality threshold corresponding to the jth indicator, after determining that the jth indicator of the ith object is abnormal, the processing module 404 is also used to determine the monitoring resource allocation weight of the qth object based on the historical operation indicator data corresponding to each indicator corresponding to the qth object, wherein the qth object is any object other than the ith object among the multiple objects; obtain multiple actual operation indicator data corresponding to the other objects other than the ith object among the multiple objects; after respectively obtaining the monitoring resource allocation weights and multiple actual operation indicator data corresponding to the other objects other than the ith object among the multiple objects, determine the compensation error rate of the ith object based on the historical operation indicator data corresponding to the jth indicator corresponding to the ith object, the monitoring resource allocation weights and multiple actual operation indicator data corresponding to the other objects other than the ith object among the multiple objects; and determine whether to perform a task migration operation on the task corresponding to the jth indicator in the ith object based on the compensation error rate.

[0092] In some optional implementations, when the indicators include processor utilization, network latency, and storage input / output latency, the monitoring resource allocation weight of the qth object is determined based on the historical operating indicator data corresponding to the indicators corresponding to the qth object, and is expressed by the following expression:

[0093] W q =(α*A q +β*B q ) / (γ*D q )

[0094] Among them, W q Assign a weight to the monitoring resource of the qth object; α is the first adjustment coefficient corresponding to the processor utilization; β is the second adjustment coefficient corresponding to the network delay; γ is the third adjustment coefficient corresponding to the storage input and output delay; A q is the historical operating indicator data of the processor utilization of the qth object; B q D is the historical operating indicator data of the network delay of the qth object; q The historical running indicator data of the storage input and output latency for the qth object.

[0095] In some optional implementations, after respectively obtaining the monitoring resource allocation weights and multiple actual operation indicator data corresponding to each of the multiple objects except the i-th object, the compensation error rate of the i-th object is determined according to the historical operation indicator data corresponding to the j-th indicator corresponding to the i-th object, the monitoring resource allocation weights and multiple actual operation indicator data corresponding to each of the multiple objects except the i-th object, and is expressed by the following expression:

[0096]

[0097] Among them, S i is the compensation error rate of the i-th object; W k Assign weights to the kth monitoring resources corresponding to each of the multiple objects except the i-th object; P k is the kth actual operating index data corresponding to each of the multiple objects except the i-th object; δ is the compensation coefficient; Q i The historical operating indicator data corresponding to the j-th indicator of the i-th object.

[0098] In some optional embodiments, the processing module 404 is specifically used to detect whether the compensation error rate is greater than the preset compensation error rate; if it is detected that the compensation error rate is greater than the preset compensation error rate, determine to perform a task migration operation on the task corresponding to the jth indicator in the i-th object.

[0099] In some optional embodiments, when it is determined that the posterior probability of a historical abnormal event for the jth indicator is greater than a preset probability abnormality threshold corresponding to the jth indicator, after determining that the jth indicator of the ith object is abnormal, the processing module 404 is further specifically used to determine the risk level corresponding to the posterior probability of a historical abnormal event for the jth indicator based on the posterior probability of a historical abnormal event for the jth indicator; and according to the risk level, execute a solution strategy corresponding to the risk level.

[0100] In some optional embodiments, before obtaining multiple actual operation indicator data corresponding to each object except the i-th object among the multiple objects, the processing module 404 is also used to generate a sorting order for the multiple objects according to the order of the monitoring resource allocation weights corresponding to each object except the i-th object among the multiple objects from large to small. The sorting order is used to indicate the order of obtaining the multiple actual operation indicator data corresponding to each object except the i-th object among the multiple objects.

[0101] For the description of the features in the embodiment corresponding to the object indicator abnormality monitoring device, please refer to the relevant description of the embodiment corresponding to the object indicator abnormality monitoring method, and no further details will be given here.

[0102] The embodiment of the present application also provides an electronic device, such as Figure 5 As shown, Figure 5 The hardware structure diagram of an electronic device provided in an embodiment of the present application is shown. The electronic device includes a processor 10 and a memory 20, wherein the memory 20 stores a computer program, and the processor 10 is configured to run the computer program to perform the steps of any of the above-mentioned object indicator abnormality monitoring method embodiments.

[0103] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned object indicator abnormality monitoring method embodiments when running.

[0104] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0105] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned object indicator abnormality monitoring method embodiments are implemented.

[0106] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned object indicator abnormality monitoring method embodiments.

[0107] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0108] The above is a detailed introduction to the object indicator abnormality monitoring method, device, electronic device and storage medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for monitoring abnormality of object indicators, characterized in that: Applied to a cloud platform, the cloud platform is used to monitor multiple operating indicator data of each of multiple objects, the method comprising: Determining a priori probability of a historical abnormal event corresponding to each of the indicators based on historical operating indicator data corresponding to each of the multiple objects and abnormal data in the historical operating indicator data corresponding to each of the indicators; Determine the abnormal conditional probability that the h-th indicator is abnormal and the j-th indicator is abnormal at the same time based on the prior probability of historical abnormal events corresponding to the j-th indicator of the i-th object, the historical operating indicator data corresponding to the h-th indicator other than the j-th indicator among the multiple indicators, and the abnormal data in the historical operating indicator data, wherein the h-th indicator is the indicator other than the j-th indicator among the multiple indicators; After obtaining the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among the multiple indicators, determining the posterior probability of the historical abnormal event of the j-th indicator according to the prior probability of the historical abnormal event corresponding to each indicator and the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among the multiple indicators; When it is determined that the posterior probability of a historical abnormal event of the jth indicator is greater than a preset probability abnormality threshold corresponding to the jth indicator, it is determined that the jth indicator of the i-th object is abnormal.

2. The method according to claim 1, characterized in that When it is determined that the posterior probability of a historical abnormal event of the jth indicator is greater than a preset probability abnormality threshold corresponding to the jth indicator, after determining that the jth indicator of the i-th object is abnormal, the method further includes: determining a monitoring resource allocation weight for the qth object according to historical operating indicator data corresponding to each indicator corresponding to the qth object, wherein the qth object is any object among the multiple objects except the i-th object; Obtaining a plurality of actual operation indicator data corresponding to the other objects except the i-th object among the plurality of objects; After respectively obtaining the monitoring resource allocation weights and the plurality of actual operation indicator data corresponding to the other objects among the plurality of objects except the i-th object, determining the compensation error rate of the i-th object according to the historical operation indicator data corresponding to the j-th indicator corresponding to the i-th object, the monitoring resource allocation weights and the plurality of actual operation indicator data corresponding to the other objects among the plurality of objects except the i-th object; Determine whether to perform a task migration operation on the task corresponding to the jth indicator in the i-th object according to the compensation error rate.

3. The method according to claim 2, characterized in that When the indicators include processor utilization, network latency, and storage input / output latency, the monitoring resource allocation weight of the qth object is determined based on the historical operating indicator data corresponding to each indicator corresponding to the qth object, and is expressed by the following expression: W q =(α*A q +β*B q ) / (γ*D q ) Among them, W q Assign a weight to the monitoring resource of the qth object; α is the first adjustment coefficient corresponding to the processor utilization; β is the second adjustment coefficient corresponding to the network delay; γ is the third adjustment coefficient corresponding to the storage input and output delay; A q B is the historical operating indicator data of the processor utilization of the qth object; q D is the historical operating indicator data of the network delay of the qth object; q The historical operating indicator data of the storage input and output delay of the qth object.

4. The method according to claim 2, characterized in that After respectively obtaining the monitoring resource allocation weights and multiple actual operation indicator data corresponding to the other objects except the i-th object among the multiple objects, the compensation error rate of the i-th object is determined according to the historical operation indicator data corresponding to the j-th indicator corresponding to the i-th object, the monitoring resource allocation weights and multiple actual operation indicator data corresponding to the other objects except the i-th object among the multiple objects, and is expressed by the following expression: Among them, S i is the compensation error rate of the i-th object; W k Assigning weights to the kth monitoring resources corresponding to the other objects except the i-th object among the plurality of objects; k is the kth actual operation index data corresponding to each of the objects except the i-th object among the plurality of objects; δ is the compensation coefficient; Q i The historical operating indicator data corresponding to the j-th indicator corresponding to the i-th object.

5. The method according to claim 4, characterized in that The determining, based on the compensation error rate, whether to perform a task migration operation on the task corresponding to the jth indicator in the i-th object includes: detecting whether the compensation error rate is greater than a preset compensation error rate; If it is detected that the compensation error rate is greater than the preset compensation error rate, it is determined to perform a task migration operation on the task corresponding to the jth indicator in the i-th object.

6. The method according to any one of claims 1 to 5, characterized in that: When it is determined that the posterior probability of a historical abnormal event of the j-th indicator is greater than a preset probability abnormality threshold corresponding to the j-th indicator, after determining that the j-th indicator of the i-th object is abnormal, the method further includes: Determining the risk level corresponding to the posterior probability of historical abnormal events for the j-th indicator based on the posterior probability of historical abnormal events for the j-th indicator; According to the risk level, a solution strategy corresponding to the risk level is executed.

7. The method according to claim 6, characterized in that Before obtaining a plurality of actual operation indicator data corresponding to the other objects except the i-th object among the plurality of objects, the method further includes: A sorting order of the multiple objects is generated according to the order from large to small of the monitoring resource allocation weights corresponding to each object except the i-th object among the multiple objects. The sorting order is used to indicate the order of obtaining the multiple actual operation indicator data corresponding to each object except the i-th object among the multiple objects.

8. A device for monitoring abnormality of an object index, characterized in that: Applied to a cloud platform, the cloud platform is used to monitor multiple operating indicator data of each of multiple objects, and the object indicator abnormality monitoring device includes: a priori probability determination module, configured to determine a priori probability of a historical abnormal event corresponding to each of the indicators based on historical operating indicator data corresponding to each of the multiple objects and abnormal data in the historical operating indicator data corresponding to each of the indicators; a conditional probability determination module, configured to determine, based on a priori probability of historical abnormal events corresponding to the jth indicator of the i-th object, historical operating indicator data corresponding to the hth indicator other than the jth indicator among the multiple indicators, and abnormal data in the historical operating indicator data, an abnormal conditional probability that the hth indicator is abnormal and the jth indicator is abnormal at the same time, wherein the hth indicator is an indicator other than the jth indicator among the multiple indicators; a posterior probability determination module, configured to, after obtaining the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among the multiple indicators, determine the posterior probability of the historical abnormal event of the j-th indicator according to the prior probabilities of the historical abnormal events corresponding to the respective indicators and the abnormal conditional probabilities corresponding to the other indicators except the j-th indicator among the multiple indicators; The processing module is used to determine that the jth indicator of the i-th object is abnormal when it is determined that the posterior probability of the historical abnormal event of the j-th indicator is greater than the preset probability abnormality threshold corresponding to the j-th indicator.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the object indicator abnormality monitoring method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the object indicator abnormality monitoring method according to any one of claims 1 to 7 are implemented.