Cloud pool equipment anomaly detection method and device

By constructing and analyzing the multi-scale feature matrix of cloud pool equipment and combining with the pre-trained anomaly detection model, the problems of poor abnormal detection accuracy and inability to locate the root cause of abnormality in the existing technology are solved, and the effect of high accuracy and accurate root cause positioning is achieved.

CN120066890APending Publication Date: 2025-05-30CHINA TELECOM CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510061129.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the cloud pool equipment abnormality detection, it is difficult to fully consider the relationship between the interrelationship and influence between various monitoring indicators, and the time information capture effect in the index data is poor, resulting in poor detection accuracy and the root cause of equipment abnormality cannot be located.

Method used

By obtaining the operating status data set of cloud pool equipment in each sub-time period of different lengths within the preset time period, a single-scale feature matrix is ​​constructed and spliced, the multi-scale feature matrix is ​​reconstructed and analyzed using the pre-trained anomaly detection model, the abnormal sub-time period is determined and the monitoring indicators that cause the abnormality are located.

Benefits of technology

It improves the accuracy of cloud pool equipment abnormal detection, can sensitively capture and accurately evaluate equipment operation abnormalities, and achieve accurate positioning of the root causes of abnormalities, reducing equipment maintenance difficulty and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066890A_ABST
    Figure CN120066890A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud pool equipment anomaly detection method and device. The method comprises the following steps: acquiring an operation state data set of cloud pool equipment in each sub-time period with different lengths in a preset time period; constructing a single-scale feature matrix corresponding to the sub-time period according to each running state data set, and splicing the single-scale feature matrixes to obtain a first multi-scale feature matrix; reconstructing the first multi-scale feature matrix by using an anomaly detection model, and analyzing the reconstructed second multi-scale feature matrix to obtain a first anomaly score of the cloud pool device in each sub-time period; and determining a target sub-time period when the cloud pool equipment is abnormal according to each first abnormal score, and positioning an abnormal monitoring index causing the cloud pool equipment to be abnormal. According to the method and the device, the technical problems that the analysis result is relatively poor in accuracy and the root cause of equipment abnormity cannot be positioned due to the fact that abnormity analysis is carried out on the cloud pool equipment through a single monitoring index in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and more particularly, to a method and device for detecting anomalies in cloud pool devices. Background Art

[0002] In the maintenance of cloud pool devices, the detection of anomalies in cloud pool devices is one of the important means to ensure the normal operation of cloud pool devices. At present, many enterprises still rely on automated inspections with fixed thresholds internally, but this method has weak adaptability to changes in business loads and cannot perform dynamic anomaly detection based on real-time data. Therefore, the research and development of AI-based anomaly detection play an important role in the maintenance of cloud pool devices.

[0003] Since the operating state of cloud pool devices is affected by multiple metrics, current AI-based anomaly detection methods are difficult to fully consider the mutual correlation and influence relationships among various metrics, and have poor capture effects on the time information in metric data, greatly reducing the accuracy of anomaly detection in cloud pool devices. In addition, existing anomaly detection methods can only determine whether a cloud pool device is in an abnormal state, but cannot evaluate the current anomaly level of the cloud pool device, let alone conduct corresponding root cause analysis to determine the main reasons for the anomaly.

[0004] To address the above problems, no effective solutions have been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a method and device for detecting anomalies in cloud pool devices, so as to at least solve the technical problem that the related art analyzes anomalies in cloud pool devices through a single monitoring metric, resulting in poor accuracy of the analysis results and inability to locate the root cause of device anomalies.

[0006] According to one aspect of the embodiments of the present application, a method for detecting anomalies in cloud pool devices is provided, including: obtaining an operation status data set of a target cloud pool device within sub-time periods of various different lengths within a preset time period, where the operation status data set includes operation status data sequences of each monitoring metric within the corresponding sub-time period among a plurality of monitoring metrics; constructing a single-scale feature matrix corresponding to each sub-time period based on the operation status data set within each sub-time period, and splicing the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix; using a pre-trained anomaly detection model to reconstruct the first multi-scale feature matrix to obtain a second multi-scale feature matrix, and analyzing the second multi-scale feature matrix to obtain a first anomaly score of the target cloud pool device within each sub-time period, where the anomaly detection model at least includes: a convolution module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, and a transposed convolution module; determining a target sub-time period in which the target cloud pool device has an anomaly based on the first anomaly scores of the target cloud pool device within each sub-time period, and positioning the anomaly monitoring metric that causes the target cloud pool device to have an anomaly within the target sub-time period based on the single-scale feature matrix and the second multi-scale feature matrix corresponding to the target sub-time period.

[0007] Optionally, constructing a single-scale feature matrix corresponding to each sub-time period based on the operation status data set within each sub-time period includes: for each sub-time period, calculating the inner product between every two operation status data sequences in the operation status data set of the sub-time period, and using the inner product as the correlation degree between the corresponding two operation status data sequences; taking each operation status data sequence within the operation status data set corresponding to the sub-time period as the rows and columns of a matrix respectively, and using the correlation degrees between the operation status data sequences as matrix elements to construct a single-scale feature matrix corresponding to the sub-time period.

[0008] Optionally, the anomaly detection model further includes: an output module, where using the pre-trained anomaly detection model to reconstruct the first multi-scale feature matrix to obtain a second multi-scale feature matrix, and analyzing the second multi-scale feature matrix to obtain a first anomaly score of the target cloud pool device within each sub-time period includes: using the convolution module in the anomaly detection model to extract features from the first multi-scale feature matrix to obtain a corresponding feature matrix; using the encoding module in the anomaly detection model to encode the feature matrix to obtain a corresponding feature encoding matrix; using the transposed convolution module in the anomaly detection model to decode the feature encoding matrix to obtain a second multi-scale feature matrix; using the output module in the anomaly detection model to analyze the second multi-scale feature matrix to obtain a first anomaly score of the target cloud pool device within each sub-time period.

[0009] Optionally, the convolutional module at least includes: a plurality of convolutional layers, and activation layers connected to each convolutional layer. Each convolutional layer is used to capture local features of a local area with the same size as the corresponding convolutional kernel in the first multi-scale feature matrix through a preset convolutional kernel, and form a corresponding feature matrix from the local features corresponding to each local area. The activation layer connected to the convolutional layer performs a non-linear transformation on the feature matrix output by the convolutional layer using a preset activation function. The encoding module at least includes: an attention layer and a feed-forward neural network. The attention layer analyzes the feature matrix output by the convolutional module using a first pooling unit based on the self-attention mechanism to capture the feature correlation between all time nodes in the feature matrix and obtain a first weight matrix. At the same time, the attention layer analyzes the feature matrix output by the convolutional module using a second pooling unit based on the Gaussian kernel function to capture the feature correlation between some time nodes near the last sub-time period in the feature matrix and obtain a second weight matrix. Then, the feature matrix output by the convolutional module is weighted and fused through the first weight matrix and the second weight matrix and output through a residual connection method. The feed-forward neural network maps the feature encoding matrix output by the attention layer and outputs it through a residual connection method to obtain a corresponding feature encoding matrix. The transposed convolutional module at least includes: a plurality of transposed convolutional layers. The transposed convolutional layer is used to adjust the length and pad zeros to the feature encoding matrix output by the encoding module to expand the size of the feature map until the size of the second multi-scale feature matrix output by the transposed convolutional module is equal to the size of the first multi-scale feature matrix. The output module at least includes: a linear layer. The linear layer is used to determine the reconstruction residual matrix between the second multi-scale feature matrix output by the transposed convolutional module and the single-scale feature matrix corresponding to each sub-time period, calculate the difference matrix between the first weight matrix and the second weight matrix using the symmetric divergence algorithm, and normalize the difference matrix. Determine the normalized divergence value of the last sub-time period from the normalized difference matrix, and multiply the normalized divergence value of the last sub-time period by the reconstruction residual matrix corresponding to each sub-time period to obtain the first anomaly score of the target cloud pool device in each sub-time period.

[0010] Optionally, determining the target sub-time period when the target cloud pool device has an anomaly based on the first anomaly score of the target cloud pool device in each sub-time period includes: determining whether the first anomaly score of the target cloud pool device in each sub-time period is greater than the anomaly threshold value corresponding to the sub-time period; in the case where the first anomaly score of the target cloud pool device in the target sub-time period is greater than the anomaly threshold value corresponding to the target sub-time period, determining that the target cloud pool device has an anomaly in the target sub-time period, where the first anomaly score is used to reflect the severity of the anomaly of the target cloud pool device in the target sub-time period.

[0011] Optionally, the method for determining the abnormal threshold value corresponding to each sub-time period includes: for each sub-time period of the same length, obtaining the second abnormal score of the target cloud pool device in multiple sub-time periods at the same scale obtained by analyzing multiple historical multi-scale feature matrices of the target cloud pool device by the abnormal detection model, where each historical multi-scale feature matrix is constructed at least from the historical operation state data set of at least one sub-time period with the same length traced back from the historical moment by the target cloud pool device, and the historical operation state data set includes the historical operation state data corresponding to at least one abnormal monitoring index; selecting the largest second target abnormal score from the multiple second abnormal scores, and using the result of multiplying the second target abnormal score by a preset empirical weight coefficient as the preset abnormal threshold value corresponding to each of the multiple sub-time periods at the same scale.

[0012] Optionally, locating the abnormal monitoring index that causes the target cloud pool device to be abnormal in the target sub-time period based on the single-scale feature matrix and the second multi-scale feature matrix corresponding to the target sub-time period includes: determining the target reconstruction residual matrix between the second multi-scale feature matrix and the single-scale feature matrix corresponding to the target sub-time period, where the matrix elements in the target reconstruction residual matrix are the reconstruction errors of the correlations between each monitoring index and other monitoring indexes; when the sum of the matrix elements in the target row in the target reconstruction residual matrix is greater than a preset threshold value, determining the monitoring index corresponding to the target row as the abnormal monitoring index that causes the target cloud pool device to be abnormal in the target sub-time period, and reconstructing.

[0013] According to another aspect of the embodiments of the present application, there is also provided a cloud pool device abnormal detection device, including: an acquisition module, configured to acquire the operation state data set of the target cloud pool device in sub-time periods of different lengths within a preset time period, where the operation state data set includes the operation state data sequence of each monitoring index in the corresponding sub-time period among multiple monitoring indexes; a feature matrix construction module, configured to construct a single-scale feature matrix corresponding to the sub-time period according to the operation state data set in each sub-time period, and splice the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix; an abnormal analysis module, configured to use a pre-trained abnormal detection model to reconstruct the first multi-scale feature matrix to obtain a second multi-scale feature matrix, and analyze the second multi-scale feature matrix to obtain the first abnormal score of the target cloud pool device in each sub-time period, where the abnormal detection model at least includes: a convolution module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, and a deconvolution module; an abnormal location module, configured to determine the target sub-time period in which the target cloud pool device is abnormal according to the first abnormal score of the target cloud pool device in each sub-time period, and locate the abnormal monitoring index that causes the target cloud pool device to be abnormal in the target sub-time period based on the single-scale feature matrix and the second multi-scale feature matrix corresponding to the target sub-time period.

[0014] According to another aspect of the embodiments of the present application, there is also provided a computer program product, which includes: a computer program, wherein when the computer program is executed by a processor, the above-mentioned cloud pool device anomaly detection method is implemented.

[0015] According to another aspect of the embodiments of the present application, there is also provided an electronic device, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned cloud pool device anomaly detection method through the computer program.

[0016] In the embodiments of the present application, the management system reconstructs and analyzes a multi-scale feature matrix reflecting the device state of the cloud pool device within a preset time period through a deep learning model, and obtains the anomaly scores of the cloud pool device in each sub-time period within the preset time period, realizing sensitive capture and accurate evaluation of the abnormal operation of the cloud pool device; then, according to the anomaly scores of the cloud pool device in each sub-time period, the target sub-time period in which the cloud pool device has an anomaly is determined, and the anomaly monitoring index that causes the cloud pool device to have an anomaly in the target sub-time period is located based on the single-scale feature matrix corresponding to the target sub-time period and the reconstructed multi-scale feature matrix, realizing accurate root cause location of the cloud pool device. Thus, the purpose of improving the accuracy of anomaly detection, reducing the device maintenance difficulty and labor cost, and ensuring the stable operation of the cloud pool device is achieved, and furthermore, the technical problem that in the related art, the anomaly analysis of the cloud pool device is carried out through a single monitoring index, resulting in poor accuracy of the analysis result and inability to locate the root cause of the device anomaly is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0018] Figure 1 is a schematic flowchart of an optional cloud pool device anomaly detection method according to the embodiments of the present application;

[0019] Figure 2 is a schematic architecture diagram of an optional anomaly detection model according to the embodiments of the present application;

[0020] Figure 3 is a schematic architecture diagram of an optional encoding module according to the embodiments of the present application;

[0021] Figure 4 is a schematic structural diagram of an optional cloud pool device anomaly detection device according to the embodiments of the present application;

[0022] Figure 5It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or cloud pool device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or cloud pool devices.

[0025] To better understand the embodiments of the present application, some nouns or terms that appear in the description process of the embodiments of the present application are translated and explained as follows:

[0026] The cloud security resource pool is a software-based integrated security tool, that is, it has unified management, unified monitoring, orchestration and automation, as well as compliance capabilities. The resource pool integrates various security tools of the manufacturer's own ecosystem and opens the integration of third-party security tools, providing elastic and on-demand security resources similar to cloud service resources. These security tools include: firewall (FW), web application and API protection (WAAP), vulnerability management (VM), cloud workload protection platform (CWPP), cloud security posture management (CSPM), and container and Kubernetes security tools, etc. These necessary core capabilities lay the foundation for the cloud security resource pool and the security tools supported by these capabilities.

[0027] Transformer: A deep learning model architecture based on self-attention mechanism, consisting of two main parts: an encoder and a decoder. It is mainly used to process sequence data and is naturally suitable for the data structure of time series, a type of sequence. It can capture long-term dependencies in time series and has good modeling ability for time series data with complex time-dependent patterns. It has achieved remarkable results in many fields such as natural language processing and computer vision.

[0028] Anomaly Detection, also known as Outlier Detection or novelty detection, in data science and machine learning, refers to the process of identifying data points in a dataset that are significantly different from other observations. These data points are usually considered "anomalous" because they do not conform to the general pattern or behavior of the data.

[0029] Root Cause Analysis (RCA) is a systematic problem-solving method aimed at identifying the root causes that lead to problems, rather than just addressing the surface symptoms. The goal of root cause analysis is to completely solve the problem and prevent it from happening again.

[0030] Embodiment 1

[0031] According to an embodiment of the present application, a method for detecting anomalies in cloud pool devices is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0032] Figure 1 It is a schematic flowchart of a method for detecting anomalies in cloud pool devices provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0033] Step S102, obtain the operating status data set of the target cloud pool device within sub-time periods of various different lengths within a preset time period.

[0034] In the technical solution provided in step S102 above, the target cloud pool device is a cloud pool device in the cloud security resource pool. Therefore, the target cloud pool device can be a physical cloud pool device, such as a server, a storage cloud pool device, a network cloud pool device, a switch cloud pool device, etc.; it can also be a virtualized security cloud pool device, such as a virtual firewall, a Web application firewall, a traffic processing platform, etc. In addition, the preset time period is a time period obtained by tracing back a preset duration from the current moment, and this time period can be divided into multiple sub-time periods of different lengths. Therefore, the operation status data set of each sub-time period can include: the operation status data sequence of each monitoring index in multiple monitoring indexes of the target cloud pool device within the corresponding sub-time period, where the monitoring indexes include at least one of the following: central processing unit (CPU) utilization rate, memory usage, network traffic, disk I / O read and write speed, sensor temperature, etc.

[0035] Step S104: Construct a single-scale feature matrix corresponding to each sub-time period based on the operation status data set of each sub-time period, and splice the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix.

[0036] In the technical solution provided in step S104 above, the management system takes into account that a cloud pool device has multiple monitoring indexes related to the device status. Therefore, when an abnormality occurs in the cloud pool device, it will inevitably lead to mutations in some monitoring indexes, and these abnormalities will definitely be reflected in the relationship between different monitoring indexes. Therefore, a single-scale feature matrix corresponding to representing the device status within a certain time period can be constructed first, and each element in this single-scale feature matrix represents the correlation degree between two monitoring indexes within the corresponding sub-time period. Then, splice the single-scale feature matrices corresponding to each sub-time period to form a first multi-scale feature matrix representing the device status within different sub-time periods.

[0037] Step S106: Use a pre-trained anomaly detection model to reconstruct the first multi-scale feature matrix to obtain a second multi-scale feature matrix, and analyze the second multi-scale feature matrix to obtain the first anomaly score of the target cloud pool device within each sub-time period.

[0038] In the technical solution provided in step S106 above, since the first multi-scale feature matrix may have a relatively low sensitivity to anomalies, especially for those anomalies with subtle changes but significant impacts on the device state, and the model mainly relies on data preprocessing and feature engineering to understand the device state, therefore, this method of directly using the original data may not be able to fully extract the complex patterns in the data. Especially for high-dimensional, multi-index time series data, it may be difficult for the model to distinguish normal fluctuations from anomalies. For this reason, the embodiments of this application propose to train an anomaly detection model that at least includes a convolutional module, an encoding module that combines a self-attention mechanism and a Gaussian kernel function, and a deconvolution module, and use the anomaly detection model to reconstruct the first multi-scale feature matrix, so that the model can learn the internal structure and patterns in the data. At the same time, the reconstruction process helps to amplify the difference between the anomaly signal and the normal signal, improving its visibility even when the anomaly signal is relatively hidden, and improving the accuracy, robustness, and sensitivity of anomaly detection.

[0039] Step S108: Determine the target sub-time period when the target cloud pool device has an anomaly based on the first anomaly scores of the target cloud pool device in each sub-time period, and locate the anomaly monitoring indicators that cause the target cloud pool device to have an anomaly in the target sub-time period according to the single-scale feature matrix and the second multi-scale feature matrix corresponding to the target sub-time period.

[0040] In the technical solution provided in step S108 above, the management system can first determine the target sub-time period when the target cloud pool device has an anomaly based on the first anomaly scores of the target cloud pool device in each sub-time period obtained in step S106 above, and then locate the anomaly monitoring indicators that cause the target cloud pool device to have an anomaly in the target sub-time period according to the single-scale feature matrix and the reconstructed second multi-scale feature matrix corresponding to the target sub-time period, so as to analyze the root cause of the anomaly of the cloud pool device.

[0041] Through the technical solutions provided in steps S102 - S108 above, it can be known that in the above embodiments, the management system reconstructs and analyzes the multi-scale feature matrix reflecting the device state of the cloud pool device in the preset time period through a deep learning model, and obtains the anomaly scores of the cloud pool device in each sub-time period in the preset time period, realizing sensitive capture and accurate evaluation of the abnormal operation of the cloud pool device; then, according to the anomaly scores of the cloud pool device in each sub-time period, determine the target sub-time period when the cloud pool device has an anomaly, and locate the anomaly monitoring indicators that cause the cloud pool device to have an anomaly in the target sub-time period according to the single-scale feature matrix and the reconstructed multi-scale feature matrix corresponding to the target sub-time period, realizing accurate root cause location of the cloud pool device. Thus, the purpose of improving the accuracy of anomaly detection, reducing the difficulty of device maintenance and labor costs, and ensuring the stable operation of the cloud pool device is achieved.

[0042] The following describes each step of the cloud pool device anomaly detection method in combination with a specific implementation process.

[0043] As an alternative implementation, in the technical solution provided in the above step S102, the management system can obtain the operation status data sets of each target cloud pool device in each sub-time period according to the following steps, including:

[0044] The first step: Set a preset time period as the analysis time window. And this preset time period can be divided into multiple sub-time periods, and the length of each sub-time period is different, that is, the lengths are different.

[0045] For example, if the preset time period is 1 day, it can be divided into multiple sub-time periods with lengths such as 1 hour, 3 hours, 6 hours, or 12 hours. It should be noted that the setting of the preset time period depends on the detection sensitivity requirements of the actual application scenario and the device specificity of the target cloud pool device.

[0046] The second step: According to the set sub-time periods, automatically or manually collect the initial operation status data sets of the target cloud pool device in each sub-time period from the preset monitoring system.

[0047] Among them, the initial operation status data set includes the operation status data sequences of each monitoring index in the corresponding sub-time period among multiple monitoring indexes. Therefore, each operation status data sequence includes the operation status data of the monitoring index of the target cloud pool device at each moment in the corresponding sub-time period. For example, the operation status data sequences of two monitoring indexes i and j in a sub-time period with a length of s can be respectively recorded as: Among them, t represents the starting moment of the sub-time period.

[0048] The third step: Preprocess the initial operation status data sets of the target cloud pool device in each sub-time period to obtain the operation status data sets of the target cloud pool device in each sub-time period within the preset time period.

[0049] Among them, the above preprocessing operations include but are not limited to: data cleaning, forward filling, or interpolation methods to process missing values, etc., to ensure the quality and applicability of the data.

[0050] In addition, due to the limitations of the data acquisition hardware, there are differences in the acquisition intervals of each monitoring index, which will cause missing values at certain moments or time periods. Therefore, the embodiment of the present application also adds time interval standardization in the preprocessing operation, that is, uses time interval standardization technology to make the acquisition time intervals of these operation status data on the same time scale.

[0051] In addition, since the initial operating state data set may contain a large amount of state data, in order to reduce the data dimension and enable subsequent analysis to focus on the data truly related to anomaly detection and root cause analysis, the embodiment of this application also adds key metric data extraction in the preprocessing operation, that is, using the principal component analysis method, the recursive feature elimination method, the screening algorithm based on correlation (such as Pearson correlation coefficient, Spearman rank correlation coefficient or other correlation measurement methods), etc., to screen out the key monitoring metrics that have the greatest influence on the operating state of the cloud pool device from multiple monitoring metrics.

[0052] As an alternative implementation manner, in the technical solution provided in the above step S104, the management system may construct a single-scale feature matrix corresponding to the target cloud pool device in each sub-time period according to the following steps, including:

[0053] For each sub-time period, calculate the inner product between each two operating state data sequences in the operating state data set of the sub-time period, and use the inner product as the correlation degree or similarity degree between the corresponding two operating state data sequences; respectively use each operating state data sequence in the operating state data set corresponding to the sub-time period as the rows and columns of the matrix, and use the correlation degree or similarity degree between each operating state data sequence as the matrix elements to construct the single-scale feature matrix corresponding to the sub-time period.

[0054] Therefore, through the above method, the single-scale feature matrix corresponding to each sub-time period can be determined, and the size of each single-scale feature matrix can be n*n, where n represents the number of monitoring metrics. In order to reflect the device state in different time periods, the single-scale feature matrices of different lengths of sub-time periods can be spliced to form the first multi-scale feature matrix representing the device state in different sub-time periods.

[0055] For example, respectively select the preset time periods corresponding to 40 time steps (i.e., 40 moments) traced back from the current moment, and these 40 moments can be divided into two sizes of sub-time periods corresponding to 10 moments and sub-time periods corresponding to 30 moments, and calculate the single-scale feature matrix corresponding to a single scale through the above correlation calculation, and splice the single-scale feature matrices of these two scales to obtain the first multi-scale feature matrix with a shape of 2×20×20.

[0056] As an alternative implementation manner, in the technical solution provided in the above step S106, the above anomaly detection model further includes: a convolution module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, a deconvolution module, and an output module, and its specific architecture is as Figure 2 shown. Among them:

[0057] (1) Convolution module.

[0058] The convolutional module can capture the spatial information of the first multi-scale feature matrix of the input. The specific operation is to construct multiple convolutional layers, and each convolutional layer is used to capture the local features of the local area of the first multi-scale feature matrix with the same size as the corresponding convolutional kernel through a preset convolutional kernel, and the local features corresponding to each local area form the corresponding feature matrix. Among them, since the convolutional kernel sizes of each convolutional layer are different, it can help the model learn the features of the data at different levels, from small-scale local patterns to large-scale complex associations, thereby improving the sensitivity and accuracy of the model for anomaly detection. For example, there are three convolutional layers in the convolutional module, and the convolutional kernel sizes of the first and second convolutional layers are 2×2, and the stride is 2, so as to extract more local spatial features from the input first multi-scale feature matrix, and at the same time reduce the size of the feature map by a stride of 2 to achieve a certain degree of dimensionality reduction and feature abstraction; while the convolutional kernel size of the third convolutional layer is 5×5 and the stride is 0, so as to capture more extensive context information and complex patterns through a larger convolutional kernel, and a stride of 0 means that it uses each input value at the same position for convolution operation, usually this refers to the use of padding technology, so that the size of the output feature map is the same as the input, which helps to retain more detailed information.

[0059] In addition, since the convolutional layer itself performs a linear transformation. However, a linear model cannot handle complex, non-linear data relationships. Therefore, in the embodiments of the present application, a preset activation function is further added after each convolutional layer as an activation layer to introduce a non-linear transformation, so that the model can learn and represent more complex data distributions and feature relationships. Among them, the above preset activation function can be the SELU activation function, because the SELU activation function has the self-normalization property, which means that it can automatically adjust the distribution of features during the training process of the network, making the mean of the features close to 0 and the standard deviation close to a fixed value, thereby helping to stabilize the training process of the model, avoiding the problems of gradient disappearance or gradient explosion, and ensuring that the model can still effectively learn and transmit information after multiple layers of convolution.

[0060] Therefore, assume that the convolutional kernel sizes of the first and second convolutional layers are 2×2 and the stride is 2; the size of the convolutional kernel of the third convolutional layer is 5×5 and the stride is 0. If the shape of the input first multi-scale feature matrix is 2×20×20, then after passing through this convolutional module and then through dimension conversion, a feature map with a shape of 1×128 can be formed.

[0061] (2) Encoding module.

[0062] The encoding module is designed based on the Transformer architecture, and its architecture diagram is as Figure 3As shown. The focus of this module lies in the construction of the attention head. The attention head of the encoding module in the embodiments of this application includes two branches: a first pooling unit based on the self-attention mechanism and a second pooling unit based on the Gaussian kernel function.

[0063] Specifically, the first pooling unit analyzes the feature matrix output by the convolutional module to capture the feature correlation between global time nodes within the feature matrix, obtaining a first weight matrix. Here, the global time nodes refer to k temporally correlated time steps (a preset time period) intercepted by a time window of a fixed time length k. Additionally, the above self-attention mechanism realizes the attention weight assignment to the value through the attention pooling of the query and the key (which means modeling the correlation between the query and the key to achieve pooling screening or weight assignment), generating the final output result. Therefore, based on the above knowledge, if the input sequence is (h represents the dimension of the feature matrix output by the convolutional module), its attention weight matrix (i.e., the first weight matrix) can be expressed as:

[0064]

[0065] where, Q = X·W q represents the word vector of the input sequence and the weight matrix which represents the query vector; K = X·W k represents the word vector of the input sequence and the weight matrix which represents the key vector; V = XW v represents the word vector of the input sequence and the weight matrix which represents the value vector; d K represents the dimension of the key vector. The Softmax function normalizes each row of the attention score matrix to obtain the attention weight matrix

[0066] Therefore, if the shape of the feature map output by the convolutional module is 1×128, then the first pooling unit can obtain the query matrix q 、W k 、 by three different learnable weight matrices W the key matrix the value matrix Then, through the query matrix Q, the key matrix K, and the value matrix V, the attention score matrix is obtained, and the first weight matrix obtained by normalizing each row of the attention score matrix through the Softmax function

[0067] The second convergence unit analyzes the feature matrix output by the convolution module to capture the feature correlation between some time nodes near the last sub-time period in the feature matrix. The partial time nodes near the last sub-time period here refer to the few time steps of the time window of a fixed time length k (preset time period) closest to the last sub-time period. In addition, the above-mentioned Gaussian kernel function, also known as the radial basis function, is a scalar function that is radially symmetric. It is usually defined as a monotonic function of the Euclidean distance between any point in space and a certain center point. Therefore, based on the above concept, if the input sequence is (h represents the dimension of the feature matrix output by the convolution module), its second weight matrix can be expressed as:

[0068] σ=X·W σ

[0069]

[0070] in, It represents a learnable weight matrix, and multiplying the input sequence X with it can obtain the scale matrix i and j represent the i-th time point and j-th time point in a time window of fixed length k, respectively. i represents the scale parameter corresponding to the i-th time point, Denotes the second weight matrix, and each row in the matrix represents a probability distribution.

[0071] Therefore, if the shape of the feature map output by the convolution module is 1×128, then the second aggregation unit can be obtained through a different learnable weight matrix Get the scale matrix Then, the Gaussian kernel is used to analyze the scale moment σ to obtain the second weight matrix Each row in this matrix represents a probability distribution.

[0072] In addition, the above-mentioned attention layer and the position-based feedforward neural network constitute two sublayers of the Transformer architecture, and each sublayer adopts a residual connection, that is, the output of the attention layer is added to the input, and the output of the feedforward neural network is also added to the input to form a residual connection. This connection can ensure the effective transmission of input information, allowing the model to capture deeper features. After the addition of the residual connection, the normalization layer is followed by normalization to finally form the encoding module. The output of the encoding module is the feature encoding map of the last sub-time period that combines the k outputs from the convolution module captured by the time window, and the last sub-time period combines the information of all sub-time periods within the preset time period.

[0073] (3) Deconvolution module.

[0074] The transposed convolution module decodes the feature encoding mapping output by the encoding module to obtain the reconstructed second multi-scale feature matrix. The specific operation is to construct multiple transposed convolution layers (deconvolution layers) and restore the input feature encoding matrix in the reverse steps of the convolution module (i.e., gradually restore and enlarge the size of the feature encoding mapping), and the output shape of the last transposed convolution layer of the transposed convolution module should be the same as the input shape of the first convolution layer of the convolution module. Finally, the output of the transposed convolution module is the reconstructed second multi-scale feature matrix.

[0075] Therefore, the transposed convolution layer enlarges the size of the feature encoding mapping by adjusting the stride and "zero-padding". For example, if the convolution layer operates with a stride of 2 and reduces the size of the input feature mapping, then the corresponding transposed convolution layer will also operate with a stride of 2, but it will increase the size of the output feature mapping.

[0076] In addition, the transposed convolution layer may also add additional "zero-padding" at the output edge to further enlarge the size of the feature mapping. Different from the input padding in the convolution operation, the transposed convolution layer usually uses output padding to ensure that the output size is enlarged as expected. Output padding is to add extra zeros at the edge of the feature mapping to restore or exceed the size of the original input.

[0077] It should be noted that if a pooling layer is used in the convolutional network to reduce the size of the feature mapping, then in the transposed convolutional network, a reverse pooling operation can be used to restore the original size of the feature mapping. The reverse pooling will restore the information lost in the pooling operation and provide a more detailed input for the transposed convolution.

[0078] (4) Output module.

[0079] The above output module obtains the anomaly scores of the target cloud pool device in each sub - time period by analyzing the first multi - scale feature matrix, the reconstructed second multi - scale feature matrix, the first weight matrix, and the second weight matrix. The specific operation is as follows: The second multi - scale feature matrix output by the de - convolution module is compared with the single - scale feature matrix corresponding to each sub - time period through a linear layer, and the square difference is calculated element - by - element to form a reconstructed residual matrix corresponding to each sub - time period. Each element in the reconstructed residual matrix represents the reconstruction error of each monitoring index in this sub - time period, and the error of abnormal points or abnormal segments will significantly increase. Then, the symmetric KL divergence is used to measure the difference between the two matrices output by the two attention branches (i.e., self - attention and Gaussian kernel - based attention) of the attention layer. Generally, for normal data, the difference between the two weight matrices is relatively large, while for abnormal data, the difference between the two weight matrices is relatively small. Therefore, the introduction of KL divergence helps to optimize the model during the training stage, making it better at distinguishing normal data and abnormal data. Finally, by combining the reconstruction error of the reconstructed residual matrix and the difference of the symmetric KL divergence, the anomaly monitoring score is obtained. The specific implementation process is as follows: The Softmax function is used to normalize the difference matrix to obtain the weight of each sub - time period. Since the output of the encoding module mainly focuses on the information of the last sub - time period, the KL divergence value of the last sub - time period is selected from the normalized difference matrix and multiplied by each squared error in the reconstructed residual matrix corresponding to each sub - time period and then summed to obtain the first anomaly detection score of the target cloud pool device in each sub - time period. Therefore, the expression of the above - mentioned first anomaly detection score can be written as:

[0080]

[0081] Among them, represents one of the single - scale feature matrices of m scales, represents the KL divergence of the last sub - time period among k sub - time periods after the normalization operation, represents the L2 norm.

[0082] Therefore, according to the above model structure, when performing end - to - end training on the anomaly detection model, the number of training iterations can be preset (for example, the number of iterations is 10), and by continuously adjusting the target loss function of the model, the performance of the model can reach the preset requirements.

[0083] Specifically, the loss function of this model includes the following two parts:

[0084] First: For the first branch in the attention head of the encoding module, that is, the branch constructed based on the self-attention mechanism, the calculated first weight matrix is obtained from the information of k time steps intercepted by a time window of fixed length k; for the second branch in the attention head of the encoding module, that is, the branch constructed based on the Gaussian kernel, the calculated second weight matrix is mainly obtained by capturing the information of adjacent time steps. For the feature matrix composed of normal data, the difference between the weight matrices obtained by these two branches is large, while for the abnormal feature matrix, the difference between these two weight matrices is relatively small. Therefore, the symmetric KL divergence is introduced to describe the difference between these two weight matrices, and it is made to continuously increase during the training phase to better meet the reconstruction of normal data.

[0085] Second: Another component of the loss function is the reconstruction error of the feature matrix, which is obtained by calculating the squared loss between the original feature matrix (i.e., the single-scale feature matrix corresponding to each sub-time period) and the reconstructed feature matrix (i.e., the second multi-scale feature matrix), and it is made to continuously decrease during the training process.

[0086] Therefore, the expression of the target loss function of the anomaly detection model in the embodiments of this application is:

[0087]

[0088] Among them, represents the input single-scale feature matrix, represents the second multi-scale feature matrix after I is reconstructed, S ij represents the element in the first weight matrix G ij represents the element in the second weight matrix i and j represent the i-th and j-th sub-time periods, and n represents the number of monitored indicators collected.

[0089] Based on the above anomaly detection model, the management system can obtain the first anomaly score of the target cloud pool device in each sub-time period according to the following steps, including:

[0090] Step S1061, use the convolution module in the anomaly detection model to extract features from the first multi-scale feature matrix to obtain the corresponding feature matrix;

[0091] Step S1062, use the encoding module in the anomaly detection model to encode the feature matrix to obtain the corresponding feature encoding matrix;

[0092] Step S1063, use the deconvolution module in the anomaly detection model to decode the feature encoding matrix to obtain the second multi-scale feature matrix;

[0093] Step S1064: Analyze the second multi-scale feature matrix using the output module in the anomaly detection model to obtain the first anomaly scores of the target cloud pool device in each sub-time period.

[0094] Furthermore, since a multi-scale feature matrix is constructed during the model analysis process, and each scale of the feature matrix represents the correlation between different indicators in different time periods, the anomaly scores obtained from the reconstructed residual matrices corresponding to different single-scale feature matrices represent the anomaly indices of different degrees of the target cloud pool device in the corresponding time periods. Among them, for the reconstructed residual matrix corresponding to the feature matrix of the small scale (i.e., the feature matrix composed of shorter time periods), because it constructs the feature matrix using the time series of shorter time periods, it is sensitive to anomalies and can detect anomalies regardless of whether they are long or short in duration; for the reconstructed residual matrix corresponding to the large-scale feature matrix, compared with the reconstructed residual matrix corresponding to the small-scale feature matrix, its sensitivity to anomalies is lower, and it can only detect anomalies with long durations. Therefore, during the anomaly detection stage, the anomaly scores of the device at different scales can be calculated separately to determine the severity of the anomaly.

[0095] As an optional implementation manner, after determining the first anomaly scores of the target cloud pool device in each sub-time period through the above Step S106, the management system can also determine the anomaly detection of the target cloud pool device and determine the corresponding severity of the anomaly according to the following method, including:

[0096] First, determine whether the first anomaly score of the target cloud pool device in each sub-time period is greater than the anomaly threshold value corresponding to the sub-time period.

[0097] In the case where the first anomaly score of the target cloud pool device in the target sub-time period is greater than the anomaly threshold value corresponding to the target sub-time period, it is determined that the target cloud pool device has an anomaly in the target sub-time period, where the first anomaly score is used to reflect the severity of the anomaly of the target cloud pool device in the target sub-time period.

[0098] Specifically, the anomaly threshold value corresponding to the above sub-time period is determined by the historical anomaly detection scores of multiple sub-time periods of the same scale. Therefore, for each sub-time period of the same scale, its corresponding anomaly detection threshold value can be determined according to the following method:

[0099] Obtain the second anomaly scores of the target cloud pool device in multiple sub-time periods at the same scale obtained by analyzing multiple historical multi-scale feature matrices of the target cloud pool device by the anomaly detection model. Each historical multi-scale feature matrix is constructed at least from the historical operation status data set of at least one sub-time period with the same length traced back from the historical moment by the target cloud pool device, and the historical operation status data set includes the historical operation status data corresponding to at least one anomaly monitoring metric; select the largest second target anomaly score from the multiple second anomaly scores, and use the result of multiplying the second target anomaly score by a preset empirical weight coefficient as the preset anomaly threshold value corresponding to each of the multiple sub-time periods at the same scale.

[0100] That is to say, according to the above anomaly detection model, the status data set containing anomaly monitoring data (that is, the cloud pool device is abnormal due to the anomaly of a certain monitoring metric) is imported to obtain the reconstructed historical multi-scale feature matrix, and the corresponding reconstructed residual matrix is determined based on the original historical multi-scale feature matrix and the reconstructed historical multi-scale feature matrix. The KL divergence calculated by the two weight matrices obtained by the encoding module is normalized by the Softmax function and combined into the reconstructed residual matrix to obtain the second anomaly scores of each sub-time period of the target cloud pool device at the same scale. Select the largest second target anomaly score from the multiple second anomaly scores, and use the result of multiplying the second target anomaly score by a preset empirical weight coefficient (α ∈ [1, 2]) as the preset anomaly threshold value corresponding to each of the multiple sub-time periods at the same scale, that is, T = α·max(Score).

[0101] Furthermore, after determining the target sub-time period in which the target cloud pool device has an anomaly, in order to further assist the relevant staff in solving the anomaly problem, the management system can also locate the root cause of the anomaly of the target cloud pool device according to the following method:

[0102] Determine the target reconstructed residual matrix between the second multi-scale feature matrix and the single-scale feature matrix corresponding to the target sub-time period, where the matrix elements in the target reconstructed residual matrix are the reconstructed errors of the correlations between each monitoring metric and other monitoring metrics;

[0103] In the case that the sum of the matrix elements in the target row in the target reconstructed residual matrix is greater than the preset threshold value, determine that the monitoring metric corresponding to the target row is the anomaly monitoring metric that causes the target cloud pool device to have an anomaly in the target sub-time period, where the sum of the reconstructed errors in the target row is the influence degree of the anomaly monitoring metric that causes the target cloud pool device to have an anomaly in the target time period.

[0104] Specifically, since the matrix shape of the single-scale feature matrix corresponding to the target sub-time period is n×n, where n is the number of monitored metrics collected, each row represents the correlation between a certain monitored metric and other monitored metrics. Therefore, for the target reconstruction residual matrix corresponding to the target sub-time period, each row represents the reconstruction error of the correlation between a certain metric and other monitored metrics. The larger the reconstruction error, the less ideal the reconstruction of the correlation between this monitored metric and other monitored metrics, that is, it represents that a problem with this monitored metric has caused an abnormality in the target cloud pool device. Therefore, the sum of the reconstruction errors of each row of the target reconstruction residual matrix represents the degree of influence of the abnormality caused by the monitored metric corresponding to this row.

[0105] Through the above steps, the management system innovatively integrates multiple monitored metrics that affect the device status to construct a feature matrix representing the device status. At the same time, the correlation between different monitored metrics is effectively encoded through convolution. Compared with the method of generally considering the influence of a single monitored metric on the device status alone, it is more complete and efficient. At the same time, the encoding module in the anomaly detection model can not only accurately capture the temporal correlation between the operation status data, but also cleverly utilize the difference in the reference to adjacent time information in the reconstruction of normal data and abnormal data, and combine it with the reconstruction error, significantly increasing the difference between the reconstruction of normal and abnormal data, thus greatly improving the accuracy of anomaly detection. In addition, compared with the method that can generally only detect device anomalies, the cloud pool device management method provided in the embodiments of the present application can not only provide the severity of the detected anomaly of the cloud pool device, facilitating the operator to take timely measures to solve the anomaly problem, but also can give the root cause of the anomaly, providing important reference for the operator.

[0106] Embodiment 2

[0107] According to the embodiments of the present application, there is also provided a cloud pool device anomaly detection device for implementing the cloud pool device anomaly detection method in Embodiment 1, as Figure 4 shown. The cloud pool device anomaly detection device at least includes: an acquisition module 42, a feature matrix construction module 44, an anomaly analysis module 46, and an anomaly location module 48, where:

[0108] The acquisition module 42 is configured to acquire the operation status data sets of the target cloud pool device within sub-time periods of different lengths within a preset time period. The operation status data sets include the operation status data sequences of each monitored metric within the corresponding sub-time period for each of the multiple monitored metrics;

[0109] The feature matrix construction module 44 is configured to construct a single-scale feature matrix corresponding to the sub-time period based on the operation status data set within each sub-time period, and splice the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix;

[0110] Anomaly analysis module 46 is configured to reconstruct the first multi-scale feature matrix by using a pre-trained anomaly detection model to obtain a second multi-scale feature matrix, and analyze the second multi-scale feature matrix to obtain the first anomaly scores of the target cloud pool device in each sub-time period. At least the following are included in the anomaly detection model: a convolution module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, and a transposed convolution module.

[0111] Anomaly localization module 48 is configured to determine the target sub-time period in which the target cloud pool device has an anomaly according to the first anomaly scores of the target cloud pool device in each sub-time period, and locate the anomaly monitoring metrics that cause the target cloud pool device to have an anomaly in the target sub-time period according to the single-scale feature matrix and the second multi-scale feature matrix corresponding to the target sub-time period.

[0112] It should be noted that each module in the cloud pool device anomaly detection device in the embodiments of the present application corresponds one by one to each implementation step of the cloud pool device anomaly detection method in Embodiment 1. Since Embodiment 1 has been described in detail, the details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.

[0113] Embodiment 3

[0114] According to an embodiment of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, the cloud pool device anomaly detection method in Embodiment 1 is implemented.

[0115] According to an embodiment of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. The cloud pool device where the non-volatile storage medium is located executes the cloud pool device anomaly detection method in Embodiment 1 by running the computer program.

[0116] According to an embodiment of the present application, there is also provided a processor, which is used to run a computer program. When the computer program runs, the cloud pool device anomaly detection method in Embodiment 1 is executed.

[0117] According to an embodiment of the present application, there is also provided an electronic cloud pool device, which includes: a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the cloud pool device anomaly detection method in Embodiment 1 through the computer program.

[0118] Specifically, when the computer program runs, it executes the following steps: obtaining a running state data set of the target cloud pool device in sub-time periods of different lengths within a preset time period, where the running state data set includes the running state data sequences of each monitoring metric in multiple monitoring metrics within the corresponding sub-time periods; constructing a single-scale feature matrix corresponding to the sub-time period based on the running state data set within each sub-time period, and splicing the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix; using a pre-trained anomaly detection model to reconstruct the first multi-scale feature matrix to obtain a second multi-scale feature matrix, and analyzing the second multi-scale feature matrix to obtain a first anomaly score of the target cloud pool device in each sub-time period, where the anomaly detection model at least includes: a convolutional module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, and a deconvolution module; determining a target sub-time period in which the target cloud pool device has an anomaly based on the first anomaly scores of the target cloud pool device in each sub-time period, and positioning the anomaly monitoring metric that causes the target cloud pool device to have an anomaly in the target sub-time period based on the single-scale feature matrix and the second multi-scale feature matrix corresponding to the target sub-time period.

[0119] As an optional implementation manner, the above-mentioned electronic cloud pool device may exist in the form of a mobile terminal, a computer terminal, or a similar computing device. Figure 5 The hardware structure block diagram of an electronic cloud pool device for implementing the cloud pool device anomaly detection method is shown. As Figure 5 shown, the electronic cloud pool device 50 may include one or more (shown as 502a, 502b,..., 502n in the figure) processors 502 (the processor 502 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 504 for storing data, and a transmission device 506 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 5 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic cloud pool device. For example, the electronic cloud pool device 50 may further include more or fewer components than Figure 5 shown, or have a different configuration from Figure 5 shown.

[0120] It should be noted that one or more of the above-mentioned processors 502 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other components in the electronic cloud pool device 50. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0121] The memory 504 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the cloud pool device anomaly detection method in the embodiments of the present application. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 504 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 504 can further include a memory remotely set relative to the processor 502, and these remote memories can be connected to the electronic cloud pool device 50 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0122] The transmission device 506 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the electronic cloud pool device 50. In one instance, the transmission device 506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network cloud pool devices through a base station and thus communicate with the Internet. In one instance, the transmission device 506 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0123] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables users to interact with the user interface of the electronic cloud pool device 50.

[0124] The above-mentioned embodiment numbers are only for description and do not represent the advantages or disadvantages of the embodiments.

[0125] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0126] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0127] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0128] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0129] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer cloud pool device (which can be a personal computer, a server, or a network cloud pool device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0130] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A cloud pool device anomaly detection method, characterized in that: include: Obtaining an operating status data set of a target cloud pool device in each sub-time period of different lengths within a preset time period, wherein the operating status data set includes an operating status data sequence of each monitoring indicator in a corresponding sub-time period in a plurality of monitoring indicators; Constructing a single-scale feature matrix corresponding to each sub-time period according to the running status data set in each sub-time period, and concatenating the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix; Reconstructing the first multi-scale feature matrix using a pre-trained anomaly detection model to obtain a second multi-scale feature matrix, and analyzing the second multi-scale feature matrix to obtain a first anomaly score of the target cloud pool device in each sub-time period, wherein the anomaly detection model at least includes: a convolution module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, and a deconvolution module; The target sub-time period in which the target cloud pool device has an abnormality is determined based on the first abnormality score of the target cloud pool device in each sub-time period, and the abnormal monitoring indicator causing the target cloud pool device to have an abnormality in the target sub-time period is located based on the single-scale feature matrix corresponding to the target sub-time period and the second multi-scale feature matrix.

2. The method according to claim 1, characterized in that Constructing a single-scale feature matrix corresponding to each sub-time period according to the running status data set in each sub-time period, including: For each of the sub-time periods, the inner product between each two operating status data sequences in the operating status data set of the sub-time period is calculated, and the inner product is used as the correlation between the corresponding two operating status data sequences; each operating status data sequence in the operating status data set corresponding to the sub-time period is used as the row and column of the matrix respectively, and the correlation between each of the operating status data sequences is used as the matrix element to construct a single-scale feature matrix corresponding to the sub-time period.

3. The method according to claim 1, characterized in that The anomaly detection model also includes: an output module, wherein the first multi-scale feature matrix is ​​reconstructed using a pre-trained anomaly detection model to obtain a second multi-scale feature matrix, and the second multi-scale feature matrix is ​​analyzed to obtain a first anomaly score of the target cloud pool device in each sub-time period, including: Using the convolution module in the anomaly detection model to perform feature extraction on the first multi-scale feature matrix to obtain a corresponding feature matrix; Encoding the feature matrix using the encoding module in the anomaly detection model to obtain a corresponding feature encoding matrix; Decoding the feature encoding matrix using a deconvolution module in the anomaly detection model to obtain the second multi-scale feature matrix; The second multi-scale feature matrix is ​​analyzed by using an output module in the anomaly detection model to obtain a first anomaly score of the target cloud pool device in each sub-time period.

4. The method according to claim 3, characterized in that: The convolution module at least includes: a plurality of convolution layers, and an activation layer connected to each of the convolution layers, wherein each of the convolution layers is used to capture the local features of the local area of ​​the first multi-scale feature matrix and the corresponding convolution kernel by a preset convolution kernel, and the local features corresponding to each local area form a corresponding feature matrix; the activation layer connected to the convolution layer is to perform a nonlinear transformation on the feature matrix output by the convolution layer using a preset activation function; The encoding module at least includes: an attention layer and a feedforward neural network, wherein the attention layer uses a first convergence unit based on a self-attention mechanism to analyze the feature matrix output by the convolution module to capture the feature correlation between all time nodes in the feature matrix to obtain a first weight matrix, and simultaneously uses a second convergence unit based on a Gaussian kernel function to analyze the feature matrix output by the convolution module to capture the feature correlation between some time nodes adjacent to the last sub-time period in the feature matrix to obtain a second weight matrix, and the feature matrix output by the convolution module is weightedly fused by the first weight matrix and the second weight matrix and then outputted by a residual connection method; the feedforward neural network maps the feature encoding matrix output by the attention layer and then outputs it by a residual connection method to obtain the corresponding feature encoding matrix; The deconvolution module at least includes: a plurality of transposed convolution layers, wherein the transposed convolution layers are used to adjust the length and perform zero padding operations on the feature encoding matrix output by the encoding module to expand the size of the feature map until the size of the second multi-scale feature matrix output by the deconvolution module is equal to the size of the first multi-scale feature matrix; The output module at least includes: a linear layer, wherein the linear layer is used to determine the reconstructed residual matrix between the second multi-scale feature matrix output by the deconvolution module and the single-scale feature matrix corresponding to each of the sub-time periods, and calculate the difference matrix between the first weight matrix and the second weight matrix using a symmetric divergence algorithm, and normalize the difference matrix; determine the normalized divergence value of the last sub-time period from the normalized difference matrix, and multiply the normalized divergence value of the last sub-time period by the reconstructed residual matrix corresponding to each of the sub-time periods to obtain the first anomaly score of the target cloud pool device in each sub-time period.

5. The method according to claim 1, characterized in that Determining a target sub-time period in which an abnormality occurs in the target cloud pool device according to the first abnormality score of the target cloud pool device in each sub-time period includes: Determine whether the first abnormality score of the target cloud pool device in each of the sub-time periods is greater than the abnormality threshold corresponding to the sub-time period; When the first abnormality score of the target cloud pool device within the target sub-time period is greater than the abnormality threshold value corresponding to the target sub-time period, it is determined that an abnormality has occurred in the target cloud pool device within the target sub-time period, wherein the first abnormality score is used to reflect the severity of the abnormality of the target cloud pool device within the target sub-time period.

6. The method according to claim 5, characterized in that The method for determining the abnormal threshold value corresponding to each of the sub-time periods includes: For each sub-time period of the same length, obtain the second anomaly score of the target cloud pool device in multiple sub-time periods at the same scale obtained by analyzing multiple historical multi-scale feature matrices of the target cloud pool device by the anomaly detection model, wherein each of the historical multi-scale feature matrices is constructed by at least a historical operating status data set of the target cloud pool device in at least one sub-time period of the same length from a historical moment back, and the historical operating status data set includes historical operating status data corresponding to at least one abnormal monitoring indicator; select the largest second target anomaly score from the multiple second anomaly scores, and use the result obtained by multiplying the second target anomaly score by a preset empirical weight coefficient as the preset anomaly threshold value corresponding to each of the multiple sub-time periods at the same scale.

7. The method according to claim 1, characterized in that The abnormal monitoring indicator causing the abnormality of the target cloud pool device in the target sub-time period is located according to the single-scale feature matrix corresponding to the target sub-time period and the second multi-scale feature matrix, including: Determine a target reconstruction residual matrix between the second multi-scale feature matrix and the single-scale feature matrix corresponding to the target sub-time period, wherein the matrix elements in the target reconstruction residual matrix are reconstruction errors of the correlation between each monitoring indicator and other monitoring indicators; When the sum of the matrix elements in the target row within the target reconstructed residual matrix is ​​greater than a preset threshold value, the monitoring indicator corresponding to the target row is determined to be an abnormal monitoring indicator that causes the target cloud pool device to have an abnormality within the target sub-time period, and the reconstruction.

8. A cloud pool device anomaly detection device, characterized in that: include: An acquisition module, used to acquire an operation status data set of a target cloud pool device in each sub-time period of different lengths within a preset time period, wherein the operation status data set includes an operation status data sequence of each monitoring indicator in a corresponding sub-time period in a plurality of monitoring indicators; A feature matrix construction module is used to construct a single-scale feature matrix corresponding to each sub-time period according to the running status data set in each sub-time period, and to splice the single-scale feature matrices corresponding to each sub-time period to obtain a first multi-scale feature matrix; An anomaly analysis module, used to reconstruct the first multi-scale feature matrix using a pre-trained anomaly detection model to obtain a second multi-scale feature matrix, and analyze the second multi-scale feature matrix to obtain a first anomaly score of the target cloud pool device in each sub-time period, wherein the anomaly detection model at least includes: a convolution module, an encoding module combining a self-attention mechanism and a Gaussian kernel function, and a deconvolution module; An anomaly locating module is used to determine the target sub-time period in which the target cloud pool device has an anomaly based on the first anomaly score of the target cloud pool device in each sub-time period, and locate the anomaly monitoring indicator that causes the target cloud pool device to have an anomaly in the target sub-time period based on the single-scale feature matrix corresponding to the target sub-time period and the second multi-scale feature matrix.

9. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the cloud pool device anomaly detection method according to any one of claims 1 to 7 is implemented.

10. An electronic cloud pool device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the cloud pool device anomaly detection method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Service system anomaly detection method and device, computer equipment and storage medium

    CN111880998A

  • Machine room anomaly detection method and device based on graph structure and abnormal attention mechanism

    CN115018021A

  • Abnormality detection and positioning method and device for multivariate time series data

    CN117076171A

  • Abnormality detection method and device, electronic equipment and storage medium

    CN117667587A

  • Time sequence multi-scale fusion prediction method based on TL-GAN

    CN118194236A