Fault component positioning method and device and nonvolatile storage medium

By acquiring multi-dimensional data from the cloud resource pool, determining weights, and constructing a knowledge graph, the problem of inaccurate fault location caused by neglecting the correlation of multi-dimensional data in existing technologies is solved, and accurate fault location and early warning of cloud resource pool are achieved.

CN121814554APending Publication Date: 2026-04-07CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, cloud resource pool fault location methods based on machine learning are unable to fully capture complex and ever-changing fault scenarios and ignore the correlation between multi-dimensional data, resulting in inaccurate fault location.

Method used

By acquiring multi-dimensional data from the cloud resource pool, determining the weight of each dimension, performing fusion processing, constructing a knowledge graph, and combining mutual information volume and dynamic time warping distance to build the relationship between service entities, the operational status of the cloud resource pool is analyzed, and faulty components are located.

Benefits of technology

It improves the accuracy of fault location in cloud resource pools by comprehensively considering the correlation between multi-dimensional data, accurately identifying faulty components, and achieving early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814554A_ABST
    Figure CN121814554A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for positioning a fault component and a nonvolatile storage medium. The method comprises the steps that multi-dimensional data of a cloud resource pool in a current time window are acquired, and the multi-dimensional data are used for describing the cloud resource pool from multiple dimensions; determining a weight of each dimension in a plurality of dimensions corresponding to the multi-dimensional data, and performing fusion processing on the multi-dimensional data according to the plurality of weights to obtain a fusion feature; determining the operation state of the cloud resource pool according to the fusion feature; under the condition that the operation state is an abnormal state, constructing a knowledge graph according to an association relationship between every two service entities in the cloud resource pool; and determining a fault component which causes the operation state of the cloud resource pool to be an abnormal state according to the knowledge graph. According to the method and the device, the technical problem that the fault component of the cloud resource pool is inaccurate in positioning due to the fact that mutual influence among different dimensions is not considered and the incidence relation among multi-dimensional data is ignored is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud resource pool fault location technology, and more specifically, to a method and apparatus for locating faulty components and a non-volatile storage medium. Background Technology

[0002] In related technologies, cloud resource pool fault location methods that are not based on machine learning mainly rely on a single data source or simple threshold judgment, which makes it difficult to cope with the complex and ever-changing fault scenarios in the cloud environment. While machine learning-based cloud resource pool fault location methods, those based on single-modal data analysis, struggle to comprehensively capture the complex states of the cloud resource pool and ignore the correlations between multi-dimensional data. Therefore, they suffer from inaccurate fault location.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method and apparatus for locating faulty components, as well as a non-volatile storage medium, to at least solve the technical problem of inaccurate location of faulty components in cloud resource pools caused by neglecting the mutual influence between different dimensions and ignoring the correlation between multi-dimensional data.

[0005] According to one aspect of the embodiments of this application, a method for locating faulty components is provided, comprising: acquiring multi-dimensional data of a cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; determining the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and performing fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features, wherein each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window, and the fusion features record the correlation relationships between multiple single dimensions; determining the operating state of the cloud resource pool according to the fusion features; in the case that the operating state is abnormal, constructing a knowledge graph according to the correlation relationships between every two service entities in the cloud resource pool; and determining the faulty component that causes the operating state of the cloud resource pool to be abnormal according to the knowledge graph.

[0006] Optionally, determining the weight of each dimension among the multiple dimensions corresponding to the multi-dimensional data includes: for each single dimension, performing feature extraction on the data belonging to that single dimension to obtain a latent state vector, wherein the latent state vector is used to record the latent state features of the single-dimensional data, the latent state features are used to describe the temporal change trend of the single-dimensional data, and the temporal change trend is used to indicate the change trend of the single-dimensional data with runtime; generating a global context query vector based on multiple latent state vectors, wherein the global context query vector is used to describe the running status of the cloud resource pool in the current time window; and determining the weights based on the latent state vectors and the global context query vector.

[0007] Optionally, the multi-dimensional data is fused according to multiple weights to obtain fused features, including: generating a tensor based on the multi-dimensional data, wherein the tensor has the same dimensions as the corresponding dimensions of the multi-dimensional data, the number of elements in the tensor is the same as the number of data generation times corresponding to the multi-dimensional data, each element corresponds to a data generation time, and each element is an array composed of multiple data with the same data generation time and belonging to the same dimension; decomposing the tensor to obtain multiple factor matrices, wherein the number of factor matrices is the same as the number of dimensions corresponding to the multi-dimensional data, and each factor matrix corresponds to a single dimension; and determining the fused features based on the multiple factor matrices and multiple weights.

[0008] Optionally, determining the operating status of the cloud resource pool based on fusion features includes: generating a target adjacency matrix based on multi-dimensional data, wherein each element in the target adjacency matrix indicates the association strength between any two service entities in the cloud resource pool within the current time window; performing convolution processing on the target adjacency matrix and fusion features in a spatial domain convolutional network layer to obtain spatial domain features of the fusion features, wherein the spatial domain features reflect the association relationship between any two service entities in the spatial dimension; and performing convolution processing on the fusion features in a temporal convolutional network layer to obtain temporal domain features of the fusion features, wherein the temporal domain features indicate the association relationship between two service entities in the temporal dimension; fusing the spatial domain features and temporal domain features to obtain spatiotemporal fusion features; determining the deviation between the spatiotemporal fusion features and reference features, and determining an operating status score based on the deviation, wherein the reference features are prediction results obtained by predicting the operating status of the cloud resource pool within the current time window based on historical operating data of the cloud resource pool; and determining the operating status of the cloud resource pool based on the status score interval to which the operating status score belongs.

[0009] Optionally, a knowledge graph is constructed based on the association between every two service entities in the cloud resource pool, including: determining the mutual information and dynamic time regularization distance between every two service entities, wherein the mutual information is used to indicate the correlation of the performance of the two service entities, and the dynamic time regularization distance is used to indicate the similarity of the performance change trends of the two service entities; determining each service entity as a node of the knowledge graph, and determining the straight line connecting two nodes as an edge of the knowledge graph; generating a knowledge graph based on multiple nodes and multiple edges, wherein each edge in the knowledge graph carries a weight coefficient, and the weight coefficient of each edge is determined based on the mutual information and dynamic time regularization distance between the two service entities corresponding to the two nodes associated with the edge.

[0010] Optionally, the mutual information and dynamic time warping distance between each pair of service entities are determined. The mutual information is determined using the following method: for each pair of service entities, the mutual information is determined based on the joint probability density and the marginal probability density of each service entity. The joint probability density indicates the probability that the performance metrics of both service entities change simultaneously within the current time window, and the marginal probability density indicates the probability that each service entity changes within the current time window. The dynamic time warping distance is determined by: [The text abruptly ends here, so the translation stops.] Multiple deviations at multiple time points are determined as elements of the deviation matrix, generating the deviation matrix. For each target element contained in the deviation matrix, the target element is replaced with the sum of the target element and its target neighbor elements to obtain the target matrix. Here, the target neighbor element is the smallest neighbor element in the target element's neighbor element set. The target element's neighbor element set consists of elements adjacent to the target element whose row identifier value is less than the target element's row identifier value, and / or whose column identifier value is less than the target element's column identifier value. The value of the last element in the target matrix is ​​determined as the dynamic time warping distance.

[0011] Optionally, the faulty component causing the cloud resource pool to be in an abnormal operating state is identified based on the knowledge graph, including: for each service entity, determining the quantification result of the impact of each service entity in the cloud resource pool on the operating state of the cloud resource pool when the service entity is a faulty component, based on the weight coefficient carried by each edge connected to the node corresponding to the service entity in the knowledge graph; determining the operating state score of the cloud resource pool based on the fusion features, and determining the final score of each service entity based on the operating state score and multiple quantification results corresponding to multiple service entities; and identifying the service entity corresponding to the final score with the largest value as the faulty component.

[0012] Optionally, after identifying the faulty component, the method further includes: identifying the root cause of the fault that causes the cloud resource pool to be in an abnormal operating state based on the faulty component, and identifying fault samples related to the root cause in the data related to the faulty component; generating adversarial samples based on the fault samples; and performing optimization training based on the fault samples and adversarial samples.

[0013] According to another aspect of the embodiments of this application, an apparatus for locating faulty components is also provided, comprising: an acquisition module for acquiring multi-dimensional data of a cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; a dimension weight determination module for determining the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and performing fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features, wherein each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window, and the fusion features record the correlation relationships between multiple single dimensions; a scoring module for determining the operating state of the cloud resource pool according to the fusion features; a knowledge graph construction module for constructing a knowledge graph according to the correlation relationships between every two service entities in the cloud resource pool when the operating state is abnormal; and a faulty component determination module for determining the faulty component that causes the operating state of the cloud resource pool to be abnormal according to the knowledge graph.

[0014] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, in which a computer program is stored, wherein the above-described method for locating faulty components is executed by running the computer program in the device where the non-volatile storage medium is located.

[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described method for locating faulty components through the computer program.

[0016] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the steps of the method for locating faulty components described above.

[0017] In this embodiment, multi-dimensional data of the cloud resource pool in the current time window is obtained, whereby the multi-dimensional data describes the cloud resource pool from multiple dimensions. The weight of each dimension in the multi-dimensional data is determined, and the multi-dimensional data is fused based on these weights to obtain a fused feature. Each weight quantifies the influence of a single dimension on the current state of the cloud resource pool in the current time window, and the fused feature records the relationships between multiple single dimensions. The operating state of the cloud resource pool is determined based on the fused feature. If the operating state is abnormal, a knowledge graph is constructed based on the relationships between every two service entities in the cloud resource pool. The faulty component causing the abnormal operating state of the cloud resource pool is determined based on the knowledge graph. By locating faults in the cloud resource pool, the relationships between the multi-dimensional data of the cloud resource pool are considered. Taking into account the time and spatial dependencies between multi-dimensional data of the cloud resource pool, this method determines whether the cloud resource pool is operating abnormally based on the integrated results of these dependencies, thus improving the accuracy of the operational status assessment. In the event of cloud resource pool malfunction, the method locates the component causing the failure by analyzing the time and spatial relationships between various components (i.e., service entities) within the cloud resource pool. This achieves the goal of comprehensively considering the relationships between various components in the cloud resource pool when locating faults, thereby improving the accuracy of cloud resource pool fault location. Furthermore, it solves the technical problem of inaccurate fault component location in cloud resource pools caused by neglecting the mutual influence between different dimensions and ignoring the relationships between multi-dimensional data. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0019] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing a method for locating faulty components according to an embodiment of this application;

[0020] Figure 2 This is a flowchart of the steps of a method for locating a faulty component according to an embodiment of this application;

[0021] Figure 3 This is a structural diagram of an apparatus for locating a fault component according to an embodiment of this application;

[0022] Figure 4 This is the experimental result of a comparative experiment based on an embodiment of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:

[0026] Mutual information: used to measure the degree of dependence between two variables. In the scheme provided in the embodiments of this application, mutual information is used to evaluate the correlation of performance changes of two service entities based on the performance indicators of the two service entities.

[0027] Dynamic time warping distance: used to measure the similarity between time series. In this embodiment, dynamic time warping distance is used to measure the mutual influence between different service entities in the cloud resource pool.

[0028] In related technologies, cloud resource pool fault location methods based on multi-source heterogeneous data often employ simple feature concatenation or weighted averaging when processing multi-source heterogeneous data. This method struggles to fully uncover the potential correlations between different data modalities, resulting in inaccurate fault location. To address this issue, this application provides a related solution, detailed below.

[0029] According to an embodiment of this application, a method embodiment for locating faulty components is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0030] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal for implementing a method for locating faulty components is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0031] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0032] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for locating faulty components in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned method for locating faulty components. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0033] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0034] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.

[0035] This application provides a method for locating faulty components suitable for the above-described operating environment. Figure 2 This is a flowchart of the method for locating faulty components according to embodiments of this application, such as... Figure 2 As shown, the method includes the following steps:

[0036] Step S202: Obtain multi-dimensional data of the cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions.

[0037] This application provides a cloud resource pool fault location method based on multi-dimensional data. By analyzing the correlation between multi-dimensional data, the accuracy of fault location is improved. In step S202, the device (or system) executing the method provided in this application uses distributed probes to collect multi-dimensional data from the cloud resource pool, including CPU utilization, memory usage, network latency, disk input / output operations per second (IOPS), and log events within the current time window. The multi-dimensional data collected by the distributed probes describes the cloud resource pool from different dimensions. Therefore, the acquired multi-dimensional data can comprehensively describe the operating status of the cloud resource pool. For example, the multi-dimensional data collected by the distributed probes can form a resource view, a performance view, and a log view. The resource view describes the cloud resource pool from the dimension of resource utilization, the performance view describes the cloud resource pool from the dimension of operating performance indicators, and the log view describes the cloud resource pool from the dimension of service records. The aforementioned resource view, performance view, and log view do not refer to representing this type of data in graphical form, but rather to classifying multi-dimensional data into three categories: describing the cloud resource pool from a resource perspective, describing the cloud resource pool from a performance perspective, and log data. The current time window refers to a period of time including the current moment. The current moment can be the starting point of this time interval represented by the current time window, or it can be a point within this time interval. The duration represented by the current time window is preset. In the method provided in this application embodiment, the current time window can be determined using a sliding time window method. The length of the sliding window can be adjusted according to specific application scenarios. For example, for rapidly changing network traffic data, a shorter window length (e.g., 5 minutes) can be selected; while for relatively stable hardware performance indicators, a longer window length (e.g., 1 hour) can be selected. This flexible window setting allows the system to adapt to the characteristics of different types of data, improving the accuracy of fault detection.

[0038] Step S204: Determine the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and perform fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features. Each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window. The fusion features record the correlation between multiple single dimensions.

[0039] In step S204, the multi-dimensional data obtained in step S202 is fused. Specifically, after determining the weight of each dimension based on the attention mechanism, the multi-dimensional data is fused based on the determined weight. In this embodiment, the weight determined based on the attention mechanism is a dynamic weight. Dynamic weight means that the weight of each dimension will change according to the degree of influence of the data under that dimension on the operation status of the cloud resource pool. Therefore, the weight of each dimension determined in the current time window can only indicate the degree of influence of the data of that dimension on the operation status of the cloud resource pool in the current time window. The greater the degree of influence, the greater the weight. For example, when the cloud resource pool encounters network-related failures, data related to network connection and traffic (such as network latency and packet loss rate) has a greater impact on the operation status of the cloud resource pool, so the weight of the resource dimension is greater. If the cloud resource pool encounters storage-related failures, then the weight of the performance dimension is greater. After determining the weight of each dimension, the multi-dimensional data is fused using these weights. A fused feature that integrates multiple dimensions such as resource dimension, performance dimension, and log dimension is obtained through weighted fusion. This fusion process not only captures information from a single dimension, but also captures the correlation between different dimensions. Therefore, the obtained fused feature can reflect the correlation between multiple single dimensions.

[0040] According to some optional embodiments of this application, determining the weight of each dimension among multiple dimensions corresponding to multi-dimensional data includes: for each single dimension, performing feature extraction on the data belonging to the single dimension to obtain a latent state vector, wherein the latent state vector is used to record the latent state features of the single-dimensional data, the latent state features are used to describe the temporal change trend of the single-dimensional data, and the temporal change trend is used to indicate the change trend of the single-dimensional data with runtime; generating a global context query vector based on multiple latent state vectors, wherein the global context query vector is used to describe the running status of the cloud resource pool in the current time window; and determining the weight based on the latent state vector and the global context query vector.

[0041] In this embodiment, determining the dynamic weight of each single dimension using a self-attention mechanism specifically includes the following steps: First, extract the latent state vector from the data under each single dimension. The extraction process of the latent state vector utilizes a Long Short-Term Memory (LSTM) network unit, which can capture long-term dependencies. Therefore, the extracted latent state vector can quantify the temporal change trend of the data contained in each dimension (i.e., the change trend of data as the cloud resource pool runs). For example, if a resource view, performance view, and log view are generated based on the multi-dimensional data obtained in step S202, feature extraction is performed on the data set corresponding to each view to obtain the latent state vector of the data in each dimension. Where i represents a single data dimension (resource dimension, performance dimension, etc.). Specifically, the hidden state vector of dimension i... The generation method can be described as a formula. In the formula, This represents an activation function (hyperbolic tangent function). and These are the parameters of the LSTM unit. It is the weight matrix of the LSTM unit. It is the bias vector of the LSTM unit; This is the factor matrix of dimension i at the k-th component. For each dimension, k is the index of a different component or pattern obtained in the tensor decomposition. The value of k affects the complexity of the tensor decomposition and can be preset. Next, a global context query vector q is generated based on the multiple hidden state vectors corresponding to multiple dimensions. The process of generating the global context query vector q can be described by the formula... In the formula, N represents the number of dimensions, and t represents each moment contained in the current time window. The hidden state vector is obtained by feature extraction of data belonging to a single dimension i at time t. According to the formula for generating the global context query vector q described above, the global context query vector can reflect the overall operating status of the cloud resource pool within the current time window, such as resource consumption level, performance efficiency, and the frequency of system log events. Finally, after determining the hidden state vector and the global context query vector, the dynamic weight of each single dimension is determined based on an attention mechanism. The dynamic weight aims to quantify the degree of influence of each dimension on the operating status of the cloud resource pool within the current time window. In this embodiment, the method for determining the dynamic weight can be described by the formula... ,in, Indicates the first Real-time weights of each view (i.e., the i-th dimension) at time t N represents the number of dimensions (the number of types of dimensions) corresponding to the multi-dimensional data. For each view (i.e., each individual dimension), there is a hidden state vector. yes transpose, For global context query vector, The activation function is sigmoid. The dynamic weight calculation method provided in this embodiment can effectively capture the complex interactions between different dimensions and improve the effect of feature fusion.

[0042] In the method provided in this application embodiment, before feature extraction from single-dimensional data, the acquired data undergoes the following preprocessing: specifically, log data is encoded to capture long-term dependencies in the log sequence; numerical features are standardized to eliminate the influence between different units; and a sliding window method is used to convert time-series data into fixed-length samples. These preprocessing steps improve the effectiveness and efficiency of subsequent analysis. The process of encoding log data can be described by the formula... ,in, This is the log event vector at the current moment. This represents the hidden state of the LSTM unit. In this way, the system can effectively capture long-term dependencies in log sequences and extract more meaningful feature representations. The standardization process for numerical features can be described by the formula: ,in, These are the original eigenvalues. and These are the mean and standard deviation of the feature, respectively. This is a small constant used to prevent the denominator from being zero. This standardization process allows characteristics of different dimensions to be compared and analyzed on the same scale.

[0043] Optionally, the multi-dimensional data is fused according to multiple weights to obtain fused features, including: generating a tensor based on the multi-dimensional data, wherein the tensor has the same dimensions as the corresponding dimensions of the multi-dimensional data, the number of elements in the tensor is the same as the number of data generation times corresponding to the multi-dimensional data, each element corresponds to a data generation time, and each element is an array composed of multiple data with the same data generation time and belonging to the same dimension; decomposing the tensor to obtain multiple factor matrices, wherein the number of factor matrices is the same as the number of dimensions corresponding to the multi-dimensional data, and each factor matrix corresponds to a single dimension; and determining the fused features based on the multiple factor matrices and multiple weights.

[0044] After determining the weight of each dimension according to the method provided in the previous embodiment, the multi-dimensional data is fused based on the multiple weights corresponding to multiple dimensions to obtain fused features. The process of fusing multi-dimensional data based on the multiple weights corresponding to multiple dimensions can be classified into the following steps: First, a tensor is constructed based on the multi-dimensional data. Specifically, the multi-dimensional data is converted into a unified representation form, and an N-dimensional tensor is constructed, where N represents the number of dimensions corresponding to the multi-dimensional data. The number of elements contained in this tensor is the same as the number of data generation times contained in the current time window, and each element is an array composed of multi-dimensional element data with the same data generation time. For example, the shape constructed based on the resource view, performance view, and log view data is as follows: 3D tensor ,in, yes , and Obtained by splicing along the feature dimension (D). It is a shape constructed based on resource view data (resource dimension data). The resource dimension tensor, It is a shape constructed based on performance view data (data with performance as the dimension). The performance dimension tensor, It is a shape constructed based on log view data (data categorized as logs). The tensor of the log data dimension, as described above, This indicates the number of resource nodes (i.e., the number of service entities contained in the cloud resource pool). This indicates the length of the current time window (i.e., the number of times data was generated). Representing the characteristic dimensions of resources, The value depends on the number of data categories classified as resource dimensions. For example, resource dimension data includes: CPU utilization, memory utilization, and network bandwidth utilization. The value of is 5; Feature dimensions representing performance The value depends on the number of data categories that are classified as performance dimensions; Representing the feature dimensions of the logs, The value depends on the number of data categories that are classified as log data dimensions; Next, we will examine the constructed N-dimensional tensor. The decomposition yields factor matrices for each dimension, where each factor matrix represents a tensor. In this embodiment, the tensor decomposition method for features in this dimension can be described by the following formula: In the formula, This is the concatenated tensor (i.e., a tensor generated from multi-dimensional data). For the first The weighting coefficients of each component (can be set according to the actual situation of the cloud resource pool). The value is preset based on the diversity of fault modes in the cloud resource pool. Representative dimension (Such as resource dimensions, performance dimensions, etc.) in the k-component factor matrix, then The number of dimensions is the same as the number of dimensions corresponding to multi-dimensional data. For example, if there are only three dimensions: resource dimension, performance dimension, and log data dimension, then... ,in, It is the factor matrix corresponding to the resource dimension r. Performance dimension The corresponding factor matrix, It is a log data dimension The corresponding factor matrix, Represents the tensor outer product. For the pre-defined residual terms, the tensor decomposition method provided in this embodiment can effectively capture the potential correlations between different views. Finally, based on the dynamic weights of each dimension and the corresponding factor matrix, a fusion process is performed to obtain the fused features. For example, the dynamic weights can be... The information from different dimensions is integrated by weighted averaging on the factor matrix corresponding to each dimension, thus obtaining a fusion feature F that integrates features from multiple dimensions.

[0045] Step S206: Determine the operating status of the cloud resource pool based on the fusion characteristics.

[0046] In step S206, the fused features obtained in step S204 are input into the Adaptive Spatiotemporal Graph Convolutional Network (AST-GCN). The fused features are processed by the Adaptive Spatiotemporal Graph Convolutional Network to determine the operating status score of the quantified cloud resource pool. By determining which operating status score of the cloud resource pool belongs to which score interval, the operating status of the cloud resource pool can be determined.

[0047] Optionally, determining the operating status of the cloud resource pool based on fusion features includes: generating a target adjacency matrix based on multi-dimensional data, wherein each element in the target adjacency matrix indicates the association strength between any two service entities in the cloud resource pool within the current time window; performing convolution processing on the target adjacency matrix and fusion features in a spatial domain convolutional network layer to obtain spatial domain features of the fusion features, wherein the spatial domain features reflect the association relationship between any two service entities in the spatial dimension; and performing convolution processing on the fusion features in a temporal convolutional network layer to obtain temporal domain features of the fusion features, wherein the temporal domain features indicate the association relationship between two service entities in the temporal dimension; fusing the spatial domain features and temporal domain features to obtain spatiotemporal fusion features; determining the deviation between the spatiotemporal fusion features and reference features, and determining an operating status score based on the deviation, wherein the reference features are prediction results obtained by predicting the operating status of the cloud resource pool within the current time window based on historical operating data of the cloud resource pool; and determining the operating status of the cloud resource pool based on the status score interval to which the operating status score belongs.

[0048] The method provided in this application embodiment for determining the operating status of a cloud resource pool based on fusion features can be divided into the following steps: generating a (target) adjacency matrix, extracting spatial domain features and temporal domain features, calculating an operating status score, and determining the operating status. The (target) adjacency matrix reflects the dynamic association strength between nodes (each node represents a service entity) in the cloud resource pool; that is, each element in the (target) adjacency matrix... This represents the relevance of the running states of service entity i and service entity j in the cloud resource pool within the current time window. The method for generating the dynamic adjacency matrix (i.e., the target adjacency matrix) can be described by the formula... In the formula, It is a function. , , These are the parameters of the neural network layer that processes fused features. and The first The query matrix and key matrix of the layer, Using the scaling factor dimension, where T represents transpose, the aforementioned dynamic adjacency matrix generation mechanism adaptively adjusts the association strength between nodes, better adapting to the complex and ever-changing topology of cloud resource pools. Each element of the dynamic adjacency matrix (i.e., the target adjacency matrix) The method for determining it can be described by the formula: In the formula, T represents transpose. It is the query vector of service entity i, and the feature representation of service entity i (generated based on the historical operation data of service entity i and the historical interaction information between service entity i and other service entities). It is the key vector of service entity j, and is the feature representation of service entity j (generated based on the historical operation information of service entity j). This represents the activation function (Sigmoid activation function), used to... and The dot product is mapped to the interval [0,1], and the mapping result represents the strength of the association between the two service entities.

[0049] The extraction of spatial and temporal features is accomplished using an Adaptive Spatiotemporal Graph Convolutional Network (AST-GCN). Specifically, the spatial domain convolutional network layer performs convolution processing on the target adjacency matrix and the fused features to obtain the spatial domain features in the fused features. The spatial domain convolutional network layer considers the spatial relationships between service entities (such as network topology, geographical location, or functional associations) during the convolution process. Therefore, the extracted spatial domain features can reflect the spatial relationship between two service entities. The temporal domain features are obtained by convolution processing the fused features using a temporal convolutional network layer. These temporal domain features indicate the temporal relationship between two service entities, such as the relationship between the trends of their operational states over time. After extracting the temporal and spatial domain features, they are fused to obtain the spatiotemporal fused features. The method for generating the spatiotemporal fused features can be described by the following formula: In the formula, Indicates the first Layer fusion feature representation, This is a dynamic adjacency matrix (target adjacency matrix). This represents a feature concatenation operation (which can be understood as a fusion process), where ReLU is the activation function. It is a spatial domain convolutional network layer. This represents the fused features of the input in the spatial domain convolutional network layer. ) and target adjacency matrix ( Convolution processing is performed to obtain spatial domain features. It is a temporal convolutional network layer. This represents the fused features of the input in the temporal convolutional network layer. Convolution processing is performed to obtain time-domain features. It is a fused feature obtained by fusing time domain features and spatial domain features (i.e., spatiotemporal fused feature).

[0050] As can be seen from the above, the next step after extracting time-domain and spatial-domain features is to calculate the operational status score. In this embodiment, by calculating the deviation between the fused features and the reference features, and applying a time decay function to weight the obtained deviation, an operational status score reflecting the health status of the cloud resource pool system can be obtained. The reference features mentioned above are the result of predicting the operational status of the current time window based on the historical operational data of the cloud resource pool. The formula for calculating the operational status score is as follows: In the formula, The operational status is scored. A higher operational status score indicates a greater deviation between the actual current operational status of the cloud resource pool and the current operational status predicted based on historical operational data, and a higher probability of abnormal operation. For time decay weight, This represents the reference feature of the cloud resource pool at the i-th time point in the current time window. This represents the spatiotemporal fusion characteristics of the cloud resource pool at the i-th time point in the current time window. Where e takes the value 2.71828, The timestamp of the current moment. The duration of the current time window. This is the time decay coefficient, adjusted by... The value can balance the focus on historical and recent operational data when calculating the operational status score; for example, when When calculating the operational status score, the data within the entire time window will be considered relatively evenly; while when When calculating the operational status score, the system pays more attention to data from the time closest to the current generation, making it more sensitive to sudden anomalies. Ultimately, the operational status score, along with pre-defined score intervals corresponding to different operational states, jointly determines the operational status of the cloud resource pool. For example, the score intervals can be divided into a first score interval corresponding to normal operation, a second score interval corresponding to fault warning status, and a third score interval corresponding to abnormal status. The operational status of the cloud resource pool is determined by which score interval its score falls into, thus quickly identifying the operational status of the cloud resource pool and enabling early warning of faults.

[0051] Step S208: When the running state is abnormal, construct a knowledge graph based on the relationship between every two service entities in the cloud resource pool.

[0052] If the cloud resource pool's operating status is determined to be abnormal according to the method provided in step S206, the root cause localization process is initiated in step S208. First, a causal relationship graph (i.e., a knowledge graph) between system components is constructed based on the association between every two service entities in the cloud resource pool. The aforementioned service entities are components in the cloud resource pool used to provide resources or services (such as servers, network devices, components providing software services, etc.). The association between every two service entities can be represented using mutual information and dynamic time warping distance. Mutual information is used to measure the correlation between the performance changes of two service entities, while dynamic time warping distance reflects the similarity of two service entities in a time series. In the method provided in this application embodiment, when constructing the knowledge graph, expert knowledge in the field is incorporated (such as the operation and maintenance experience of cloud resource pool operators, fault handling manuals and technical documents describing fault phenomena, common fault types, and system components and services that may be affected by the fault, etc.). Therefore, the knowledge graph constructed in step S208 is a knowledge graph that integrates data-driven and prior knowledge.

[0053] According to some optional embodiments of this application, a knowledge graph is constructed based on the association relationship between every two service entities in the cloud resource pool, including: determining the mutual information and dynamic time regularization distance between every two service entities, wherein the mutual information is used to indicate the correlation of the performance of the two service entities, and the dynamic time regularization distance is used to indicate the similarity of the performance change trends of the two service entities; determining each service entity as a node of the knowledge graph, and determining the straight line connecting two nodes as an edge of the knowledge graph; generating a knowledge graph based on multiple nodes and multiple edges, wherein each edge in the knowledge graph carries a weight coefficient, and the weight coefficient of each edge is determined based on the mutual information and dynamic time regularization distance between the two service entities corresponding to the two nodes associated with the edge.

[0054] In the method provided in this application embodiment, when the cloud resource pool is in an abnormal operating state, a knowledge graph is constructed based on the relationship between every two service entities in the cloud resource pool, and expert knowledge in the field is integrated into the constructed knowledge graph. In this embodiment, a causal relationship graph (i.e., a knowledge graph) between system components (i.e., service entities) is constructed based on mutual information (MI) and dynamic time warping distance (DTW). Mutual information reflects the statistical correlation between the performance metrics of two service entities in the cloud resource pool, while dynamic time warping distance measures the similarity of the performance change trends of the two service entities in the cloud resource pool. When constructing the knowledge graph, each service entity in the cloud resource pool is defined as a node in the knowledge graph, and the straight lines connecting nodes corresponding to related service entities are defined as edges in the knowledge graph. The aforementioned mutual information and dynamic time warping distance are used to determine the weight coefficient of each edge. That is, each edge in the knowledge graph constructed based on the relationship between two service entities carries its corresponding weight coefficient. The weight coefficient of each edge reflects the degree of association between the two service entities connected by this edge; a larger weight coefficient indicates a stronger association, and the more easily the operating states of the two service entities can influence each other. The formula for calculating the weight coefficient of the edge in the causal relationship graph (i.e., the knowledge graph) is as follows: in, Indicates service entity service entities mutual information content Indicates service entity service entities The dynamic time warp distance, To balance the coefficients, by adjusting The value of can strike a balance between the relevance and temporal similarity of statistical service entities; for example, when When the system is in a certain state, it focuses more on considering statistical correlation; while when... At that time, the system will give more consideration to temporal similarity.

[0055] Optionally, the mutual information and dynamic time warping distance between each pair of service entities are determined. The mutual information is determined using the following method: for each pair of service entities, the mutual information is determined based on the joint probability density and the marginal probability density of each service entity. The joint probability density indicates the probability that the performance metrics of both service entities change simultaneously within the current time window, and the marginal probability density indicates the probability that each service entity changes within the current time window. The dynamic time warping distance is determined by: [The text abruptly ends here, so the translation stops.] Multiple deviations at multiple time points are determined as elements of the deviation matrix, generating the deviation matrix. For each target element contained in the deviation matrix, the target element is replaced with the sum of the target element and its target neighbor elements to obtain the target matrix. Here, the target neighbor element is the smallest neighbor element in the target element's neighbor element set. The target element's neighbor element set consists of elements adjacent to the target element whose row identifier value is less than the target element's row identifier value, and / or whose column identifier value is less than the target element's column identifier value. The value of the last element in the target matrix is ​​determined as the dynamic time warping distance.

[0056] When constructing a knowledge graph, the following methods provided in the embodiments of this application can be used to determine the mutual information (MI) between two service entities. For example, for service entities... service entities According to the formula Identify service entities service entities mutual information In the formula, Representative service entity The performance indicators Representative service entity Performance metrics, for example It can be a service entity The specific value of CPU usage, It can be a service entity The specific value of CPU utilization; It is a service entity The performance indicators are The marginal probability density is used to represent the service entity when the influence of other service entities is ignored in the current time window. The performance indicators changed as The probability, similarly, It is a service entity The performance indicators are The marginal probability density is used to represent the service entity when the influence of other service entities is ignored in the current time window. The performance indicators changed as The probability of; It is a service entity service entities The joint probability density is used to represent the service entity within the current time window. Performance index value And service entities Performance index value The probability, The calculation takes into account the mutual influence of the two entities, reflecting the probability that the performance indicators of the two entities will simultaneously reach a specific value, when the service entity... service entities When the performance metrics are closely related (i.e., changes in one service entity strongly affect changes in another service entity), their joint probability... will be related to their respective marginal probabilities and The products are significantly different, leading to an increase in mutual information values, while when the service entity service entities When the performance metrics are independent of each other (i.e., changes in one service entity do not affect changes in another service entity). Therefore, when two service entities are independent of each other, the calculated mutual information tends to zero.

[0057] Service Entities service entities Dynamic time warping distance Used to evaluate service entities and The similarity of performance change trends within the current time window. The specific calculation process is as follows: First, based on the service entity... and Given performance metrics data at each point in the current time window, calculate the deviation of performance metrics between two service entities at each point in time, and generate a deviation matrix D representing the differences in performance metric trends based on these deviations. For example, for service entities... The performance metric data within the current time window are arranged into a data sequence X according to the time when the performance metric data was generated. (m represents the number of data points in data sequence X, which should be the same as the number of time points contained in the current time window), service entity The performance metrics data within the current time window are arranged as a data sequence Y according to the generation time of the performance metrics data. (where n represents the number of data points in data sequence Y, which should be the same as the number of time points contained in the current time window). Then, Euclidean distance can be used to calculate the deviation values ​​of different performance index data at the same time point in data sequences X and Y, and a deviation matrix is ​​generated based on these deviation values. Next, dynamic programming is used to find the normalized path with the minimum cumulative distance. Each point on the path has a predecessor point from above, left, or upper left. This distance accumulation method ensures the continuity and minimization of the path. Therefore, when applying dynamic programming to the deviation matrix to find the normalized path with the minimum cumulative distance, for each element in the deviation matrix D... (i.e., the target element), located in The target neighbor element with the smallest value is found in the set of adjacent elements consisting of the elements above, to the left, and to the upper left of the target element. Each target element is then replaced with the sum of the target element and its target neighbor elements. In this embodiment, the matrix obtained after performing the above replacement operation on all target elements in the deviation matrix is ​​called the target matrix. The last element in the target matrix is ​​the service entity. and Dynamic time warping distance The above is located in The elements above, to the left, and to the top left can be determined based on the row and column identifiers of each element; for example, the target element. The row identifier is m and the column identifier is n, located in the target element. The element to the left row identifier Smaller than target element The line identifier m is located in the target element. The element above column identifier Smaller than target element The column identifier n is located in the target element. The element in the upper left corner row identifier Smaller than target element The row identifier m, and the column identifier Smaller than target element The column identifier n.

[0058] Step S210: Identify the faulty components that cause the cloud resource pool to be in an abnormal operating state based on the knowledge graph.

[0059] After the knowledge graph is constructed in step S208, in step S210, the PageRank algorithm is used to score the importance of each node in the knowledge graph to determine the root cause component (i.e., the faulty component) that causes the cloud resource pool to be in an abnormal operating state. Each node in the knowledge graph represents a component (i.e., a service entity) in the cloud resource pool. The importance score of each node is used to reflect the probability that the node may be a faulty component that causes the cloud resource pool to be in an abnormal operating state. There is a positive proportional relationship between the importance score of a node and the probability that the component represented by the node is a faulty component. The faulty component can be determined by sorting the components represented by each node according to the importance score.

[0060] Optionally, the faulty component causing the cloud resource pool to be in an abnormal operating state is identified based on the knowledge graph, including: for each service entity, determining the quantification result of the impact of each service entity in the cloud resource pool on the operating state of the cloud resource pool when the service entity is a faulty component, based on the weight coefficient carried by each edge connected to the node corresponding to the service entity in the knowledge graph; determining the operating state score of the cloud resource pool based on the fusion features, and determining the final score of each service entity based on the operating state score and multiple quantification results corresponding to multiple service entities; and identifying the service entity corresponding to the final score with the largest value as the faulty component.

[0061] This application embodiment uses the PageRank algorithm to locate faulty components by applying it to the constructed knowledge graph. The specific process of applying the PageRank algorithm is as follows: For each service entity (or node) in the knowledge graph... According to the formula The described method identifies each service entity (or node). Final score Among them, the final score Used to reflect service entities The root cause (i.e., a faulty component) that could lead to the abnormal operating state of the cloud resource pool is, in this embodiment, The service entity with the highest value is identified as the faulty component. The above calculation results in the final score. In the formula, n is the total number of service entities contained in the cloud resource pool. The operating status score of the cloud resource pool is determined by the fusion features obtained from the fusion processing of multi-dimensional data of the cloud resource pool (see the above embodiment for the method of determining the operating status score). Indicates rating Regarding service entities The partial derivative reflects For all other nodes The contribution of abnormal scores; Representatives, taking into account service entities Under these conditions, service entities The PageRank value reflects When considered a potential root cause, The degree of impact on the operational status of the cloud resource pool, that is to say, yes In the case of being regarded as a faulty component, The quantitative results of the impact on the operational status of the cloud resource pool. Formulas can be used Calculations are performed, in which, Represents the node To the node The weight of the edges (i.e., the weight coefficient carried by each edge). It is a set of nodes Any node in, It is a node The set of adjacent nodes, Includes all of the above. Directly connected nodes; From node To the node The weight of each edge (i.e., the weight coefficient carried by each edge); d represents the weight coefficient passed from neighboring nodes to the node. The probability is such that the value of d is preset, for example, it can be set to 0.85.

[0062] According to some alternative embodiments of this application, after determining the faulty component, the method further includes: determining the root cause of the fault that causes the cloud resource pool to be in an abnormal operating state based on the faulty component, and determining the fault sample related to the root cause in the data related to the faulty component; generating adversarial samples based on the fault samples; and performing optimization training based on the fault samples and adversarial samples.

[0063] After identifying the faulty component using the method provided in this application, a thorough analysis of the faulty component can determine the root cause of the fault. Following the determination of the root cause, the system extracts fault samples directly related to the root cause from log events, performance metrics, and resource usage statistics associated with the faulty component. These samples typically contain multimodal data before and after the fault occurs, reflecting the specific impact of the root cause on the system state. To enhance the robustness and adaptability of the system, in the method provided in this application, after identifying the fault samples using the above method, perturbation processing of the fault samples can generate adversarial samples with similar fault patterns but with minor perturbations. For example, a fault generator can be trained based on a Generative Adversarial Network (Wasserstein GAN) framework. The generator accepts noise input and conditional labels, and outputs adversarial samples similar to the fault samples but with specific perturbations. The generated adversarial samples can be used in conjunction with the fault samples to optimize and train the model executing the method provided in this application. During the optimization training process, an Adaptive Moment Estimator (AdamW) weight decay optimizer is used, and the optimization is performed according to the formula... Adjust the learning rate, where, The adjusted learning rate, The pre-set base learning rate, This represents the current number of training steps. This refers to the number of steps in the warm-up phase. This learning rate adjustment strategy maintains a small learning rate in the early stages of training to avoid drastic changes in model parameters, then gradually increases the learning rate to accelerate convergence, and finally gradually decreases the learning rate for fine-tuning. For example, the base learning rate can be... Set to 0.001, the number of steps in the preheating phase. Set to 10% of the total training steps.

[0064] The process of optimizing training can be achieved using formulas. To describe, in the formula, The training objective of the generator (G) is to minimize the loss, while The training objective of the discriminator (D) is to maximize the loss. Let D(x) represent the expected value of the real data sample, and D(x) be the score given by the discriminator to the real data sample (x). A negative number representing the expected value of the generated data sample. Let G(z) be the probability distribution of the noise (z) input to the generator, G(z) be the data sample generated by the generator, and (D(G(z)) be the score given by the discriminator to the generated data sample. To balance the parameters, GP is the gradient penalty term. ,in, It is data obtained through linear interpolation between real data samples and generated data samples. This indicates that the discriminator evaluates the sample. The output of the discriminator, that is, the discriminator's response to the random interpolated sample. Authenticity assessment This indicates finding the derivative.

[0065] When optimizing and training a neural network model (a model including temporal convolutional network layers and spatial convolutional network layers) that implements the methods provided in the embodiments of this application, the neural network model can be loaded into memory. For example, the raw data of the neural network model can be loaded from non-volatile memory into volatile memory so that the processor can run the neural network model. The raw data of the neural network model refers to unprocessed data, which typically includes the parameters and structural data of the neural network model. The structural data can be the computational relationships based on the parameters, such as the forward propagation computational relationships between intermediate layers and between neurons. Specifically, the structural data can include the structure-related code of the neural network model, such as code used to perform related calculations between intermediate layers and between neurons.

[0066] In one implementation, a region can be partitioned in memory for loading the neural network model, which may include a structure data storage area and a parameter storage area. The structure data storage area stores structure-related code, and the parameters referenced by it can be accessed via pointers pointing to the addresses of specific parameters in the parameter storage area. During the training of the neural network model, frequent parameter updates may be required; in this case, updating the parameter values ​​in the parameter storage area is sufficient.

[0067] Through the above steps, multi-dimensional data fusion processing and adaptive weight adjustment of the cloud resource pool can be achieved. The operating status of the cloud resource pool can be determined based on the fused multi-dimensional data, improving the accuracy of operating status analysis. In the case of abnormal operating status, a knowledge graph is constructed based on the multi-dimensional associations of multiple service entities. The faulty component can be quickly located through the knowledge graph, shortening the fault response time, improving the accuracy of fault location, and solving the problem of inaccurate fault location caused by isolated information and insufficient extraction of spatiotemporal features.

[0068] Figure 3 This is a structural diagram of a device for locating fault components according to an embodiment of this application, such as... Figure 3As shown, the device for locating faulty components includes: an acquisition module 30, which acquires multi-dimensional data of the cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; a dimension weight determination module 32, which determines the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and performs fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features, wherein each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window, and the fusion features record the correlation between multiple single dimensions; a scoring module 34, which determines the operating state of the cloud resource pool according to the fusion features; a knowledge graph construction module 36, which constructs a knowledge graph according to the correlation between every two service entities in the cloud resource pool when the operating state is abnormal; and a faulty component determination module 38, which determines the faulty component that causes the operating state of the cloud resource pool to be abnormal according to the knowledge graph.

[0069] It should be noted that, Figure 3 Preferred embodiments of the shown examples can be found in [reference needed]. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.

[0070] Figure 4 This is the result of a comparative experiment. To verify the effectiveness of the method provided in this application embodiment, simulation experiments were conducted in the same simulation environment using the method provided in this application embodiment and a traditional Long Short-Term Memory (LSTM) network model. During the simulation, three months of normal operation data from the same cloud resource pool were collected as the training set, containing 80% fault-free data and 20% faulty data. An additional month's worth of data was prepared as the test set. During model training, the model was first trained using normal data, and then faulty data was gradually introduced. An optimizer was used, with an initial learning rate set to 0.001, using the aforementioned learning rate adjustment strategy. The number of training epochs was set to 100, and the batch size was 64. During the test, 1-3 faults were randomly injected every hour. The system's fault detection results, root cause localization results, and response time were recorded. Finally, when evaluating the performance of different methods, the accuracy, recall, and F1 score of fault detection were calculated. The accuracy of root cause localization was evaluated. The average time from fault injection to the system providing a diagnostic result was statistically analyzed. The experimental results are as follows: Figure 4 As shown, Figure 4 As shown, the solution provided in this application embodiment has higher performance in fault location under different fault scenarios (CPU overload, memory leak, network packet loss, disk I / O storm, service deadlock).

[0071] This application also provides a non-volatile storage medium storing a computer program, wherein the device containing the non-volatile storage medium executes the above-described method for locating faulty components by running the computer program.

[0072] The aforementioned non-volatile storage medium is used to store programs that perform the following functions: acquiring multi-dimensional data of the cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; determining the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and performing fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features, wherein each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window, and the fusion features record the correlation between multiple single dimensions; determining the operating state of the cloud resource pool according to the fusion features; in the case of an abnormal operating state, constructing a knowledge graph according to the correlation between every two service entities in the cloud resource pool; and identifying the faulty components that cause the operating state of the cloud resource pool to be abnormal according to the knowledge graph.

[0073] This application also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to execute the above-described method for locating faulty components through the computer program.

[0074] The processor in the aforementioned electronic device is used to run a program that performs the following functions: acquiring multi-dimensional data of the cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; determining the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and performing fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features, wherein each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window, and the fusion features record the correlation between multiple single dimensions; determining the operating state of the cloud resource pool according to the fusion features; in the case of an abnormal operating state, constructing a knowledge graph according to the correlation between every two service entities in the cloud resource pool; and determining the faulty component that caused the operating state of the cloud resource pool to be abnormal according to the knowledge graph.

[0075] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-described method for locating faulty components.

[0076] It should be noted that each module in the above-mentioned fault location component device can be a program module (e.g., a set of program instructions to implement a certain function) or a hardware module. For the latter, it can be manifested in the following forms, but is not limited to them: each of the above modules is manifested as a processor, or the functions of each of the above modules are implemented by a processor.

[0077] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0078] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0079] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0080] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0081] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0082] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0083] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for locating a faulty component, characterized in that, include: Obtain multi-dimensional data of the cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; The weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data is determined, and the multi-dimensional data is fused according to the multiple weights to obtain a fusion feature. Each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window. The fusion feature records the correlation between multiple types of single dimensions. The operating status of the cloud resource pool is determined based on the fusion characteristics; In the event that the operating state is abnormal, a knowledge graph is constructed based on the relationship between every two service entities in the cloud resource pool. The faulty component that caused the cloud resource pool to be in the abnormal state was identified based on the knowledge graph.

2. The method according to claim 1, characterized in that, Determining the weight of each dimension among the multiple dimensions corresponding to the multi-dimensional data includes: For each of the single dimensions, feature extraction is performed on the data belonging to the single dimension to obtain a latent state vector. The latent state vector is used to record the latent state features of the single-dimensional data, the latent state features are used to describe the temporal change trend of the single-dimensional data, and the temporal change trend is used to indicate the change trend of the single-dimensional data with the duration of runtime. A global context query vector is generated based on multiple hidden state vectors, wherein the global context query vector is used to describe the operating status of the cloud resource pool in the current time window; The weights are determined based on the hidden state vector and the global context query vector.

3. The method according to claim 1, characterized in that, The multi-dimensional data is fused based on multiple weights to obtain fused features, including: A tensor is generated based on the multi-dimensional data, wherein the dimensions of the tensor are the same as the dimensions corresponding to the multi-dimensional data, the number of elements contained in the tensor is the same as the number of data generation times corresponding to the multi-dimensional data, each element corresponds to one data generation time, and each element is an array composed of multiple data with the same data generation time and belonging to the same dimension. The tensor is decomposed to obtain multiple factor matrices, wherein the number of factor matrices is the same as the number of dimensions corresponding to the multi-dimensional data, and each factor matrix corresponds to a single dimension. The fusion feature is determined based on the multiple factor matrices and the multiple weights.

4. The method according to claim 1, characterized in that, Determining the operating status of the cloud resource pool based on the fusion characteristics includes: A target adjacency matrix is ​​generated based on the multi-dimensional data, wherein each element in the target adjacency matrix is ​​used to indicate the association strength between any two service entities in the cloud resource pool in the current time window; The target adjacency matrix and the fusion feature are convolved in the spatial domain convolutional network layer to obtain the spatial domain feature of the fusion feature, wherein the spatial domain feature is used to reflect the spatial relationship between any two service entities. Furthermore, the fused features are convolutionally processed in a temporal convolutional network layer to obtain temporal features of the fused features, wherein the temporal features are used to indicate the relationship between the two service entities in the temporal dimension; The spatial domain features and the temporal domain features are fused to obtain spatiotemporal fused features; The deviation between the spatiotemporal fusion feature and the reference feature is determined, and the operating status score is determined based on the deviation. The reference feature is a prediction result obtained by predicting the operating status of the cloud resource pool in the current time window based on the historical operating data of the cloud resource pool. The operating status of the cloud resource pool is determined based on the status score range to which the operating status score belongs.

5. The method according to claim 1, characterized in that, A knowledge graph is constructed based on the relationships between every two service entities in the cloud resource pool, including: Determine the mutual information and dynamic time regularization distance between each pair of service entities, wherein the mutual information is used to indicate the correlation of the performance of the two service entities, and the dynamic time regularization distance is used to indicate the similarity of the performance change trends of the two service entities. Each of the service entities is identified as a node in the knowledge graph, and the straight line connecting two nodes is identified as an edge in the knowledge graph. The knowledge graph is generated based on the multiple nodes and the multiple edges, wherein each edge in the knowledge graph carries a weight coefficient, and the weight coefficient of each edge is determined based on the mutual information and dynamic time warping distance of the two service entities corresponding to the two nodes associated with the edge.

6. The method according to claim 5, characterized in that, Determine the mutual information and dynamic time-warped distance between every two service entities, wherein, The mutual information is determined by the following method: for every two service entities, the mutual information of the two service entities is determined based on the joint probability density of the two service entities and the marginal probability density of each service entity, wherein the joint probability density is used to indicate the probability that the performance indicators of the two service entities change simultaneously in the current time window, and the marginal probability density is used to indicate the probability that each service entity changes in the current time window. The dynamic time warping distance is determined by the following method: Multiple deviations in the performance index data of the two service entities at multiple moments within the current time window are identified as elements of a deviation matrix, generating a deviation matrix; for each target element contained in the deviation matrix, the target element is replaced with the sum of the target element and its target neighbor elements to obtain a target matrix, wherein the target neighbor element is the smallest neighbor element in the set of neighbor elements of the target element, and the set of neighbor elements of the target element consists of elements adjacent to the target element whose row identifier value is less than the row identifier value of the target element, and / or whose column identifier value is less than the column identifier value of the target element; the value of the last element in the target matrix is ​​determined as the dynamic time warping distance.

7. The method according to claim 5, characterized in that, The faulty components that cause the cloud resource pool to enter the abnormal state, as determined by the knowledge graph, include: For each service entity, the degree of influence of each service entity in the cloud resource pool on the operating status of the cloud resource pool is determined based on the weight coefficient carried by each edge connected to the node corresponding to the service entity in the knowledge graph. In the case that the service entity is the faulty component, the quantitative result of the influence of each service entity in the cloud resource pool on the operating status of the cloud resource pool is determined. The operational status score of the cloud resource pool is determined based on the fusion characteristics, and the final score of each service entity is determined based on the operational status score and the multiple quantification results corresponding to the multiple service entities. The service entity corresponding to the final score with the highest value is identified as the faulty component.

8. The method according to claim 1, characterized in that, After identifying the faulty component, the method further includes: The root cause of the failure that caused the cloud resource pool to be in the abnormal state is determined based on the faulty component, and the fault sample related to the root cause is determined in the relevant data of the faulty component. Adversarial samples are generated based on the fault samples; optimization training is performed based on the fault samples and the adversarial samples.

9. A device for locating a faulty component, characterized in that, include: The acquisition module acquires multi-dimensional data of the cloud resource pool in the current time window, wherein the multi-dimensional data is used to describe the cloud resource pool from multiple dimensions; The dimension weight determination module is used to determine the weight of each dimension in the multiple dimensions corresponding to the multi-dimensional data, and to perform fusion processing on the multi-dimensional data according to the multiple weights to obtain fusion features. Each weight is used to quantify the degree of influence of a single dimension on the current state of the cloud resource pool in the current time window. The fusion features record the correlation between multiple types of single dimensions. The scoring module is used to determine the operating status of the cloud resource pool based on the fusion characteristics; The knowledge graph construction module is used to construct a knowledge graph based on the relationship between every two service entities in the cloud resource pool when the running state is in an abnormal state. The fault component determination module is used to determine, based on the knowledge graph, the fault components that cause the cloud resource pool to be in the abnormal state.

10. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a computer program, wherein the device containing the non-volatile storage medium executes the method for locating a faulty component as described in any one of claims 1 to 8 by running the computer program.

11. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method for locating a faulty component as described in any one of claims 1 to 8 through the computer program.

12. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method for locating a faulty component as described in any one of claims 1 to 8.