Fault positioning method and device, electronic equipment and storage medium

CN115809161BActive Publication Date: 2026-09-29DEBON SECURITIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211515026.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-09-29
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

但是这些大量的系统特征跟系统故障之间不是一对一的关系,可能一种故障的发生会带动多个特征的变化,或是某个特征的变化可能是由多种故障引起的

Benefits of technology

[0049]本申请实施例提供一种故障定位方法、装置、电子设备及存储介质,通过在监控到数据中心的节点发生故障时,确定故障对应的目标故障类型;确定所述目标故障类型对应的第一目标特征子集;计算所述数据中心中各个节点的特征与所述第一目标特征子集中各个特征之间的相似度,并基于相似度将各个节点的特征进行排列,得到特征集合;基于特征选择算法从所述特征集合中确定第二目标特征子集;基于第二目标特征子集中各个特征在各个节点的实际位置,确定故障的位置,能够实现对故障节点进行准确定位。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115809161B_ABST
    Figure CN115809161B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a fault positioning method and device, electronic equipment and storage medium. When a fault of a node in a data center is monitored, a target fault type corresponding to the fault is determined. A first target feature subset corresponding to the target fault type is determined. Similarities between features of each node in the data center and each feature in the first target feature subset are calculated, and the features of each node are arranged based on the similarities to obtain a feature set. A second target feature subset is determined from the feature set based on a feature selection algorithm. The position of the fault is determined based on the actual positions of each feature in the second target feature subset in each node, and the fault node can be accurately positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data center technology, and in particular to a fault location method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of cloud computing and big data technologies, modern data centers are evolving towards greater complexity, precision, and efficiency. The numerous network nodes within these data center networks generate vast amounts of data during operation. This real-time acquisition of information, which sensitively reflects changes in the data center system's state, is called features and can be used for system status monitoring and fault diagnosis. However, the relationship between these numerous system features and system faults is not one-to-one. A single fault may cause changes in multiple features, or a change in a particular feature may be caused by multiple faults. Due to the complex correlation between features and faults, it is difficult for human intervention to pinpoint the true cause and location of faults within the features themselves. Summary of the Invention

[0003] To address the problems in the aforementioned related technologies, this application provides a fault location method, apparatus, electronic device, and storage medium.

[0004] This application provides a fault location method, including:

[0005] When a node failure is detected in the data center, determine the target failure type corresponding to the failure.

[0006] Determine the first target feature subset corresponding to the target fault type;

[0007] Calculate the similarity between the features of each node in the data center and the features in the first target feature subset, and arrange the features of each node based on the similarity to obtain a feature set;

[0008] A second target feature subset is determined from the feature set based on a feature selection algorithm;

[0009] The location of the fault is determined based on the actual position of each feature in the second target feature subset at each node.

[0010] In some embodiments, determining the target fault type corresponding to the fault when a fault is detected in a node of the data center includes:

[0011] Monitor the characteristics of each node in the data center;

[0012] If the data change corresponding to a feature is detected to exceed a preset threshold, the feature exceeding the preset threshold is identified as an abnormal feature to obtain an abnormal feature set.

[0013] Compare the abnormal feature set with the first target feature subset corresponding to each fault type;

[0014] The target fault type corresponding to the fault is determined by the fault type corresponding to the first target feature subset that is the same as the abnormal feature set.

[0015] In some embodiments, the method further includes:

[0016] Obtain feature data for each feature;

[0017] The correlation between each feature and each fault type is calculated based on the feature data, and each feature is assigned a corresponding weight based on the correlation.

[0018] The initial feature subset corresponding to each fault type is determined based on the weights corresponding to each feature.

[0019] Based on the mutual information between each feature and the fault type in the initial feature subset and the mutual information between each feature, the first target feature subset corresponding to each fault type is determined from the initial feature subset.

[0020] In some embodiments, determining the first target feature subset corresponding to each fault type from the initial feature subset based on the mutual information between each feature and the fault type and the mutual information between the features includes:

[0021] Calculate the mutual information between each feature in the initial feature subset and the fault type, and calculate the mutual information between each feature in the initial feature subset;

[0022] The correlation measure between each feature and the fault type in the initial feature subset is calculated based on the mutual information between each feature and the fault type, and the correlation measure between features in the initial feature subset is calculated based on the mutual information between each feature in the feature subset.

[0023] Based on the correlation measurement between each feature and the fault type and the correlation measurement between features in the initial feature subset, a feature selection algorithm is used to determine the first target feature subset corresponding to each fault type from the initial feature subset.

[0024] In some embodiments, determining the second target feature subset from the feature set based on the feature selection algorithm includes:

[0025] Determine the first data for each feature in the feature set;

[0026] And determine the second and third data that are most adjacent to the first data;

[0027] Calculate the first difference between the first data and the second data, and calculate the second difference between the first data and the third data;

[0028] The weights of each feature in the feature set are determined based on the first difference and the second difference.

[0029] Based on the weights of each feature, a subset of the second target features is determined from the feature set.

[0030] In some embodiments, the method further includes:

[0031] Obtain the original data of the feature set;

[0032] The original data is subjected to structured and standardized processing to obtain data for each feature, wherein there is no coupling relationship between the features.

[0033] In some embodiments, the method further includes:

[0034] Obtain the current actual status data and target status data of the data center;

[0035] The target state data is converted into principal component set values;

[0036] The actual state data was dimensionality reduced using independent component analysis.

[0037] The current principal component value is determined based on the actual state data after dimensionality reduction.

[0038] The error value is determined based on the principal element setting value and the current principal element value;

[0039] Based on the error value, determine the change information of the feature;

[0040] The change information of the operational characteristics is sent to the data center so that the data center can perform fault-tolerant control based on the change information of the characteristics.

[0041] This application provides a fault location device, including:

[0042] The monitoring module is used to determine the target fault type when a fault is detected in a node of the data center.

[0043] The first determining module is used to determine the first target feature subset corresponding to the target fault type;

[0044] The calculation module is used to calculate the similarity between the features of each node in the data center and each feature in the first target feature subset, and to arrange the features of each node based on the similarity to obtain a feature set;

[0045] The second determining module is used to determine a second target feature subset from the feature set based on a feature selection algorithm;

[0046] The location module is used to determine the location of the fault based on the actual location of each feature in the second target feature subset at each node.

[0047] This application provides an electronic device, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, performs the fault location method as described above.

[0048] This application provides a storage medium storing a computer program that can be executed by one or more processors and can be used to implement the fault location method described above.

[0049] This application provides a fault location method, apparatus, electronic device, and storage medium. When a fault is detected in a node of a data center, the method determines the target fault type; determines a first target feature subset corresponding to the target fault type; calculates the similarity between the features of each node in the data center and the features in the first target feature subset, and arranges the features of each node based on the similarity to obtain a feature set; determines a second target feature subset from the feature set based on a feature selection algorithm; and determines the location of the fault based on the actual location of each feature in the second target feature subset at each node, thereby enabling accurate location of the faulty node. Attached Figure Description

[0050] The present application will be described in more detail below based on embodiments and with reference to the accompanying drawings.

[0051] Figure 1 A schematic diagram illustrating the implementation process of a fault location method provided in an embodiment of this application;

[0052] Figure 2 This is a schematic diagram of the structure of a fault location device provided in an embodiment of this application;

[0053] Figure 3 This is a schematic diagram of the structure of a fault location device provided in an embodiment of this application;

[0054] Figure 4 This is a schematic diagram of the structure of a feature selection module provided in an embodiment of this application;

[0055] Figure 5 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application.

[0056] In the accompanying drawings, the same parts are referred to by the same reference numerals, and the drawings are not drawn to scale. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0058] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0059] If the application documents contain similar descriptions such as "first, second, third", the following explanation shall be added: In the following description, the terms "first, second, third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0061] Based on the problems existing in related technologies, this application provides a fault location method. The method is applied to electronic devices, which may specifically be mobile phones, tablets, laptops, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, servers, etc. This application does not impose any limitations on these devices. The functions implemented by the fault location method provided in this application can be achieved by the processor of the electronic device calling program code, wherein the program code can be stored in a computer storage medium. Figure 1 This is a schematic diagram illustrating the implementation process of a fault location method provided in an embodiment of this application, as shown below. Figure 1 As shown, it includes:

[0062] Step S101: When a fault is detected in a node of the data center, determine the target fault type corresponding to the fault.

[0063] In this embodiment, the data center can be connected to multiple nodes, each of which can be considered a server. Each node in the data center can be in a normal state or a fault state. The normal state refers to the data transmission process when the data center system is operating smoothly and normally. The fault state refers to the abnormal state caused by a fault in an internal network node during the operation of the data center system. Fault types include abnormal states that may occur during system operation, such as server crashes, large data transmission delays, and reception failures, which constitute a set of states that prevent the data center network system from operating normally.

[0064] In this embodiment of the application, the electronic device can monitor the data center and thus detect faults in each node in the data center. The electronic device can monitor whether certain variables output by the data center deviate from the expected range, or whether the process state and certain characteristics suddenly change beyond the preset value, and then determine that a node in the data center has failed.

[0065] In this embodiment of the application, the electronic device can determine the target fault type corresponding to the fault.

[0066] In this embodiment of the application, step S101 can be implemented through the following steps:

[0067] Step S1011: Monitor the characteristics of each node in the data center;

[0068] Step S1012: If the data change corresponding to a feature is monitored to exceed a preset threshold, the feature exceeding the preset threshold is identified as an abnormal feature to obtain an abnormal feature set.

[0069] In this embodiment, features can also be considered variables. The real-time monitoring system monitors all variables in the system and can detect which variable is abnormal. When a feature is abnormal, the feature is set to 1, resulting in the abnormal feature combination X. fault .

[0070] Step S1013: Compare the abnormal feature set with the first target feature subset corresponding to each fault type.

[0071] In this embodiment of the application, a correspondence between each fault type and a first target feature subset can be established in advance, and the first target feature subset may include multiple features.

[0072] For example, for fault type A (A can take different values, representing different abnormal state types), the first target feature subset corresponding to A is X. SelectA X SelectA It is a set of abnormal features.

[0073] Step S1014: Determine the target fault type corresponding to the fault by identifying the fault type corresponding to the first target feature subset that is the same as the abnormal feature set.

[0074] Following the example above, we can compare X. fault With X SelectA If the elements are completely identical, it is considered that an exception of this type has occurred, thus determining the target fault type corresponding to the fault.

[0075] The feature refers to feature data, which is the end-to-end transmission process data sampled by each network node server during the operation of the data center. The data of each node server is used as the system feature, and the structured feature matrix is ​​used, with each feature corresponding to each node.

[0076] Step S102: Determine the first target feature subset corresponding to the target fault type.

[0077] In this embodiment, a first target feature subset corresponding to each fault type can be pre-established. After determining the target fault type, the first target feature subset corresponding to the target fault type can be determined.

[0078] Prior to step S102, the method further includes:

[0079] Step S1021: Obtain the feature data of each feature.

[0080] In this embodiment, the feature data of each feature can be represented as S, the number of iterations for each feature is set to m, and the number of nearest neighbor samples is k. The feature data can be simulated data or historical data.

[0081] Step S1022: Calculate the correlation between each feature and each fault type based on the feature data, and assign corresponding weights to each feature based on the correlation.

[0082] In this embodiment, the ReliefF algorithm can be used to calculate the correlation between each feature in the sample set and each fault type, and the feature can be assigned a corresponding weight matrix ω.

[0083] Step S1023: Determine the initial feature subset corresponding to each fault type based on the weights corresponding to each feature.

[0084] In this embodiment of the application, the feature with the smallest weight can be removed to obtain the initial feature subset S. n and the corresponding weight matrix ω n .

[0085] Step S1024: Based on the mutual information between each feature and the fault type in the initial feature subset and the mutual information between each feature, determine the first target feature subset corresponding to each fault type from the initial feature subset.

[0086] In this embodiment of the application, step S1024 can be implemented through the following steps:

[0087] Step S1: Calculate the mutual information between each feature in the initial feature subset and the fault type, and calculate the mutual information between each feature in the initial feature subset.

[0088] In this embodiment of the application, mutual information is defined as follows: Given two random variables x and y, whose probability density functions are p(x), p(y), and p(x,y), the mutual information between them is defined as follows:

[0089] The mutual information between each feature and the fault type in the initial feature subset can be calculated using the above formula, and the mutual information between each feature in the initial feature subset can also be calculated.

[0090] Step S2: Calculate the correlation measure between each feature and the fault type in the initial feature subset based on the mutual information between each feature and the fault type, and calculate the correlation measure between features in the initial feature subset based on the mutual information between each feature in the feature subset.

[0091] In the embodiments of this application, the formula can be adopted. Calculate the correlation measure between each feature and the fault type, where c is the fault type.

[0092] In the embodiments of this application, the formula can be used. Calculate the correlation measure between the various features.

[0093] Step S3: Based on the correlation measurement between each feature and the fault type and the correlation measurement between features in the initial feature subset, a feature selection algorithm is used to determine the first target feature subset corresponding to each fault type from the initial feature subset.

[0094] In this embodiment of the application, the target dimension of the first target feature subset can be set to d, and the solution formula can be used to solve the problem. To feature subset S d The first feature most relevant to the fault type is added. Then, based on the feature selection algorithm, the solution is obtained. Select the features that meet the requirements in sequence and add them to the first target feature subset S. d The algorithm ends when d features are selected, thus obtaining the first target feature subset corresponding to each fault type.

[0095] Step S103: Calculate the similarity between the features of each node in the data center and the features in the first target feature subset, and arrange the features of each node based on the similarity to obtain a feature set.

[0096] For example, the first target feature subset is x1; the similarity between the features of each node and x1 can be calculated. The similarity measure is the Pearson correlation coefficient or information entropy.

[0097] In this embodiment of the application, x1…x can be arranged in descending order of similarity to x1. nv We obtain X = {x1, x2, ..., x} nv That is, the feature set.

[0098] Step S104: Determine a second target feature subset from the feature set based on the feature selection algorithm.

[0099] In this embodiment of the application, X can be processed using an optimized feature selection algorithm to obtain a subset of X with redundant features removed, and then a second target feature subset is determined from the feature set.

[0100] In this embodiment of the application, step S104 can be implemented through the following steps:

[0101] Step S1041: Determine the first data of each feature in the feature set.

[0102] In this embodiment, the data for each feature can be represented by D, the number of sample samplings m, the threshold for feature weights δ, and X. select The initial value is an empty set, W A The weights for the A-th feature are initially set to 0. In this embodiment, the first data can be any sample R.

[0103] Step S1042: Determine the second and third data that are most adjacent to the first data.

[0104] In this embodiment of the application, the nearest neighboring second data H and third data M can be determined from the first data and different sample sets, respectively.

[0105] Step S1043: Calculate the first difference between the first data and the second data, and calculate the second difference between the first data and the third data.

[0106] In this embodiment, diff(A,R,H) represents the difference between sample R and sample H under the A-th feature, and diff(A,R,M) represents the difference between sample R and sample M under the A-th feature.

[0107] Step S1044: Determine the weight of each feature in the feature set based on the first difference and the second difference.

[0108] In this embodiment, W can be calculated by iterating through all data under each feature. A '=W A -diff(A,R,H) / m+diff(A,R,M) / m, thus obtaining the weights of each feature.

[0109] Step S1045: Determine the second target feature subset from the feature set based on the weights of each feature.

[0110] In this embodiment, all calculated feature weights are constructed into a feature weight matrix ω. n The feature set S here n Let X be the input ω, and let X be the input ω using a feature selection algorithm. n With S n The calculation ultimately yields the second target feature subset S. d It contains d important features. The feature subset S obtained by the feature selection algorithm d d important features are added to X select In the selected X select If the features in the first target feature set are the features most relevant to the fault type, then the second target feature subset is the feature most relevant to the fault type.

[0111] In some embodiments, prior to step S1041, the method further includes:

[0112] Step S11: Obtain the original data of the feature set.

[0113] In this embodiment of the application, the original data of the feature set can form a matrix X. ns×nv Where ns is the number of samples (sampling time is t) s ), where nv is the number of process variables or features.

[0114] Step S12: The original data is subjected to structured and standardized processing to obtain data for each feature, wherein there is no coupling relationship between the features.

[0115] In this embodiment of the application, the characteristics of the data matrix in the data center operation process are described according to the conditions of the process, that is...

[0116] To incorporate more system information, it is necessary to expand the features and identify statistical attributes that better describe each process state. These expanded features will be added to the attributes of each process sample. The resulting new dataset has more features than the original dataset X, but the number of samples remains the same, and the new dataset consists of structured features.

[0117] When the data center system can be completely decoupled, the aforementioned state space Then, by decomposing the eigenvalues ​​of matrix A, the expression can be transformed into: in: intermediate state Where Z = TX, and the transformation matrix is: T = [P1P2...P...]. n ], λ i P i -AP i =0, the eigenvalues ​​λi (i = 1, 2, ..., n) of the state characteristic equation are usually obtained by the following formula: / λI - A / = 0, the nth order system equation has n eigenvalues. When A can be completely expressed as eigenvalues ​​along the diagonal, it indicates that there is no coupling between the system variables, and complete decoupling can be achieved through linear transformation. However, this situation is very rare in reality because there are many coupling relationships between the system's variables.

[0118] When a data center system is not fully decoupled, the eigenvalue decomposition of some matrices can only be converted to Jordan canonical form, where Λ is:

[0119] If Λ is in Jordan canonical form, when the state coefficient matrix A of a linear time-invariant system has repeated roots, the eigenvalue decomposition of the matrix can be transformed into a more general Jordan canonical form through singular value transformation. It is evident from the formula that variable groups corresponding to the same eigenvalue cannot be completely decoupled, meaning there is a strong coupling relationship between them. However, different groups can have their coupling relationship completely decoupled through transformation, thereby achieving structured data processing.

[0120] In this embodiment of the application, data standardization processing can be performed by writing the original data into a state matrix form in the following way: X m = (X1, X2... X n ); X m Normalization is performed, where the normalization formula is: in: Represents X m The mean of X, σ represents the mean. m The standard deviation is used to standardize the original data.

[0121] Step S105: Determine the location of the fault based on the actual location of each feature in the second target feature subset at each node.

[0122] In this embodiment of the application, since each feature in the second target feature subset corresponds to each node, the location of the fault can be determined based on the actual location of each feature at each node. The location of the fault may include: physical location and logical location.

[0123] This application provides a fault location method. When a fault occurs in a node of a data center, the method determines the target fault type; determines a first target feature subset corresponding to the target fault type; calculates the similarity between the features of each node in the data center and each feature in the first target feature subset, and arranges the features of each node based on the similarity to obtain a feature set; determines a second target feature subset from the feature set based on a feature selection algorithm; and determines the location of the fault based on the actual location of each feature in the second target feature subset at each node, thereby enabling accurate location of the faulty node.

[0124] In some embodiments, the method further includes:

[0125] Step S106: Obtain the current actual status data and target status data of the data center.

[0126] In this embodiment of the application, the actual state data is collected in real time during the operation of the data center, and the target state data is the set ideal data.

[0127] In this embodiment, both the actual state data and the target state data have undergone data standardization processing. For example, taking the actual state data as an example, the system operating state collected for this period is: X1=(χ1,χ2...χ...). n ) T Write the states over m periods in state matrix form: X m = (X1, X2... X n ); X m Normalization yields: in: Represents X m The mean of X, σ represents the mean. m The standard deviation.

[0128] Step S107: Convert the target state data into master element set values.

[0129] In this embodiment of the application, the target state data can be the target output data, and the target output data can be set as X. q Using principal component regression, a product quality model X is established.q =Tθ T +F, where θ is the principal component regression model coefficient and F is the model error. Then, the target output data is transformed into principal component setpoints t. sp =χ qsp (θ T ) f , where: t sp It is the principal setpoint, x qsp This is the data quality setting, (θ) T ) f It is θ T The generalized inverse, t = χP, where x is the process variable and operation variable at the current time step, χ = [X q / u].

[0130] Step S108: Dimensionality reduction of the actual state data is performed using independent component analysis.

[0131] In this embodiment of the application, the actual state data can be decomposed into eigenvalues. Continuing from the example above, eigenvalue decomposition is performed using the covariance matrix: in:

[0132] Take the first k principal components of Λ as the analysis elements, and take the vector P = (u1, u2, ..., u3) of the corresponding first k U matrices. k ), thus obtaining the dimensionality reduction form of X. Where: T = XP; data variable matrix X m Operate on variable u, and store the two values ​​in a matrix X = [X / u].

[0133] Step S109: Determine the current principal component value based on the actual state data after dimensionality reduction.

[0134] In this embodiment of the application, principal component analysis can be performed on X to obtain the principal component model X = TP. T +E, where: T is the principal component score, P is the principal component load, and E is the model error. By performing principal component analysis on X, the changes in X can be summarized using a low-dimensional principal component space. Changes in the data matrix X are caused by changes in the manipulated variables and process disturbances. The principal component load indicates the direction of change in the process variables and manipulated variables. Thus, the current principal component values ​​of the actual state data can be obtained, and these current principal component values ​​are the principal component scores.

[0135] Step S110: Determine the error value based on the principal component setting value and the current principal component value.

[0136] Following the example above, Δt = t sp -t represents the error between the primary variable's set value and the current primary variable's value.

[0137] Step S111: Determine the change information of the feature based on the error value.

[0138] In this embodiment of the application, the principal component error can be mapped to the X space ΔX = ΔtP. T That is, [ΔX / Δu]=ΔtP T When the principal component model is correct, Δu in the above equation represents the change of the operation variable, i.e., the control flow.

[0139] The control flow includes: single or continuous control commands.

[0140] Step S112: Send the change information of the operation feature to the data center so that the data center can perform fault-tolerant control based on the change information of the feature.

[0141] In this embodiment, the control flow can be sent to the data center, thereby enabling the data center to perform fault-tolerant control and form a closed loop of information flow.

[0142] Based on the foregoing embodiments, this application provides a fault location device. The various modules and units included in the device can be implemented by a processor in a computer device; of course, they can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0143] This application provides a fault location device. Figure 2 This is a schematic diagram of the structure of a fault location device provided in an embodiment of this application, as shown below. Figure 2 As shown, the fault location device 200 includes:

[0144] The monitoring module 201 is used to determine the target fault type when a fault is detected in a node of the data center.

[0145] The first determining module 202 is used to determine the first target feature subset corresponding to the target fault type;

[0146] The calculation module 203 is used to calculate the similarity between the features of each node in the data center and each feature in the first target feature subset, and to arrange the features of each node based on the similarity to obtain a feature set;

[0147] The second determining module 204 is used to determine a second target feature subset from the feature set based on a feature selection algorithm;

[0148] The positioning module 205 is used to determine the location of the fault based on the actual location of each feature in the second target feature subset at each node.

[0149] In some embodiments, determining the target fault type corresponding to the fault when a fault is detected in a node of the data center includes:

[0150] Monitor the characteristics of each node in the data center;

[0151] If the data change corresponding to a feature is detected to exceed a preset threshold, the feature exceeding the preset threshold is identified as an abnormal feature to obtain an abnormal feature set.

[0152] Compare the abnormal feature set with the first target feature subset corresponding to each fault type;

[0153] The target fault type corresponding to the fault is determined by the fault type corresponding to the first target feature subset that is the same as the abnormal feature set.

[0154] In some embodiments, the fault location device 200 is further configured to:

[0155] Obtain feature data for each feature;

[0156] The correlation between each feature and each fault type is calculated based on the feature data, and each feature is assigned a corresponding weight based on the correlation.

[0157] The initial feature subset corresponding to each fault type is determined based on the weights corresponding to each feature.

[0158] Based on the mutual information between each feature and the fault type in the initial feature subset and the mutual information between each feature, the first target feature subset corresponding to each fault type is determined from the initial feature subset.

[0159] In some embodiments, determining the first target feature subset corresponding to each fault type from the initial feature subset based on the mutual information between each feature and the fault type and the mutual information between the features includes:

[0160] Calculate the mutual information between each feature in the initial feature subset and the fault type, and calculate the mutual information between each feature in the initial feature subset;

[0161] The correlation measure between each feature and the fault type in the initial feature subset is calculated based on the mutual information between each feature and the fault type, and the correlation measure between features in the initial feature subset is calculated based on the mutual information between each feature in the feature subset.

[0162] Based on the correlation measurement between each feature and the fault type and the correlation measurement between features in the initial feature subset, a feature selection algorithm is used to determine the first target feature subset corresponding to each fault type from the initial feature subset.

[0163] In some embodiments, determining the second target feature subset from the feature set based on the feature selection algorithm includes:

[0164] Determine the first data for each feature in the feature set;

[0165] And determine the second and third data that are most adjacent to the first data;

[0166] Calculate the first difference between the first data and the second data, and calculate the second difference between the first data and the third data;

[0167] The weights of each feature in the feature set are determined based on the first difference and the second difference.

[0168] Based on the weights of each feature, a subset of the second target features is determined from the feature set.

[0169] In some embodiments, the fault location device 200 is further configured to:

[0170] Obtain the original data of the feature set;

[0171] The original data is subjected to structured and standardized processing to obtain data for each feature, wherein there is no coupling relationship between the features.

[0172] In some embodiments, the fault location device 200 is further configured to:

[0173] Obtain the current actual status data and target status data of the data center;

[0174] The target state data is converted into principal component set values;

[0175] The actual state data was dimensionality reduced using independent component analysis.

[0176] The current principal component value is determined based on the actual state data after dimensionality reduction.

[0177] The error value is determined based on the principal element setting value and the current principal element value;

[0178] Based on the error value, determine the change information of the feature;

[0179] The change information of the operational characteristics is sent to the data center so that the data center can perform fault-tolerant control based on the change information of the characteristics.

[0180] Based on the foregoing embodiments, this application further provides a fault location device. This device performs node fault location and fault-tolerant control using an optimized ReliefF feature selection algorithm. Employing data mining analysis, it does not require a full understanding of the system's mechanistic characteristics. Instead, it analyzes the characteristics and information contained in the system node data to select the variable features that have the greatest impact on the fault type. Then, it accurately locates the fault node based on the actual location of these variable features. For the control system, independent component analysis is used to reduce the dimensionality of the system's feature information. The reduced-dimensional data automatically tracks the set ideal data output. Fault diagnosis is performed by evaluating the residual between the actual system output value and the estimated value of the model's ideal output. The residual value is used to determine the control flow that needs to be increased, and this control flow is fed back to the system to control the fault node in an ideal state.

[0181] In some embodiments, Figure 3 This is a schematic diagram of the structure of a fault location device provided in an embodiment of this application, as shown below. Figure 3 As shown, the fault location device may include: a feature selection module, a fault node location module, a real-time monitoring module, and a fault-tolerant controller. Specifically: the feature selection module filters the node data features collected from the data center network, selecting a set of feature subspaces that have the greatest impact on the identification of fault and normal states, i.e., important features (similar to the first target feature subset in the above embodiment); the fault node location module accurately locates the fault state based on the location of the important features within the data center system; the real-time monitoring module determines the fault type based on the type of important features; and the fault-tolerant controller performs inverse principal component calculation based on preset ideal output data (similar to the target state data in the above embodiment) to obtain ideal principal components, and extracts feature principal components based on real-time system data (similar to the actual state data in the above embodiment) to obtain the current principal component. By calculating the difference between the two, the control flow that needs to be added is obtained, thereby realizing a signal feedback mechanism. Feature selection and fault node location complete the fault location function; the real-time monitoring module implements the function of detecting whether a system fault has occurred; and the fault-tolerant controller implements fault-tolerant control in the event of a fault in the data center system.

[0182] Figure 4 This is a schematic diagram of the structure of a feature selection module provided in an embodiment of this application, as shown below. Figure 4As shown, the feature selection module includes: a historical database, a feature selection unit, and a feature location unit. The historical database is connected to the data center system and transmits historical data information of the system. The feature selection algorithm is connected to the historical database and transmits important features of the historical data of the system. The historical database stores network node data. The feature selection unit is used to select a subset of data. The feature location unit is used for fault location.

[0183] The real-time monitoring module includes a real-time monitoring unit and an important feature unit, wherein the real-time monitoring unit is connected to the important feature unit and transmits the health status of the data center system.

[0184] The fault-tolerant controller includes: a current principal unit, an ideal principal unit, and a controller, wherein: the current principal unit is connected to the complete feature space and transmits the selected current system information principals; the ideal principal unit is connected to the ideal output data and outputs ideal principal information; and the controller is connected to the current principal unit and the ideal principal unit and outputs control flow.

[0185] The aforementioned method of determining fault type based on important features refers to the real-time monitoring module selecting important features based on different fault states, monitoring important feature samples, and identifying an anomaly in the combination of important features. If an anomaly occurs in a combination of important features, then the corresponding anomaly type is considered to have occurred.

[0186] In building a data-driven diagnostic system for data center network fault nodes, the first step is data acquisition. This data can be obtained from historical data from each network node server or through virtual simulation. All system variables include end-to-end data from the node server transmission process, such as packet loss rate, latency, and throughput parameters. The data center system has a variable set X = {x1, x2, ..., xn}, and the relationships between these system variables can be described using a state space: Where X includes state variables and control variables, it is not necessary to list the input state U separately. Since the state space is not uniquely represented and can be expressed as a combination of many X variables, the form of the state space can be changed through linear transformations.

[0187] The data features mentioned refer to the end-to-end transmission process data sampled from each network node server during the operation of the data center. The data from each node server is used as system features, and the structured feature matrix is ​​used, with each feature corresponding to each node.

[0188] The feature selection is defined as follows: From a feature set X containing D features, select a subset containing d features, where d is known and d ≤ D, such that, according to a certain evaluation criterion, this subset is the optimal subset among all subsets of the original set containing d elements, and contains the most state information, thus completely finding the subset of these anomalous state features. Under certain conditions, this can achieve the localization of the root cause of anomalous states. Where: X Select It is a subset of X, with an indefinite number of features, but the number of features is less than or equal to the number of features in X. Feature selection is an important technique in data mining, pattern recognition, and machine learning, playing a crucial role in data preparation and preprocessing, and is also widely used in fault diagnosis.

[0189] The basic idea of ​​the ReliefF feature selection algorithm is to assign different weights to each feature based on the difference in relevance between each feature and the sample class. If the weight is less than a set threshold, the feature will be removed; therefore, it belongs to the feature weighting algorithm. The Relief algorithm determines the relevance between features and classes based on the feature's ability to distinguish sample classes.

[0190] The optimized ReliefF feature selection algorithm is the mRMR-ReliefF algorithm. As mentioned above, the ReliefF algorithm calculates the weight of each feature based on its correlation with the sample class, selecting the subset of features with the highest correlation to the sample class. However, it ignores the correlation between features, thus failing to select the optimal subset. The mRMR (Maximum Relevance Minimum Redundancy) algorithm, on the other hand, can select the optimal subset of features that are correlated with the sample class and independent of each other. However, the features selected by this algorithm all have the same effect on sample classification. Therefore, the idea of ​​the mRMR algorithm is used to optimize the ReliefF algorithm.

[0191] The fault location device provided in this application embodiment utilizes a data information processing method. It does not require a full understanding of the mechanism characteristics of the data center system. It can accurately locate the fault state of the system, identify the fault state type, and control the ideal state by analyzing the data characteristics of the data center system.

[0192] It should be noted that, in the embodiments of this application, if the above-described fault location method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0193] Accordingly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the fault location method provided in the above embodiments.

[0194] This application provides an electronic device; Figure 5 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application, such as... Figure 5 As shown, the electronic device 900 includes: a processor 901, at least one communication bus 902, a user interface 903, at least one external communication interface 904, and a memory 905. The communication bus 902 is configured to enable communication between these components. The user interface 903 may include a display screen, and the external communication interface 904 may include standard wired and wireless interfaces. The processor 901 is configured to execute a fault location method program stored in the memory to implement the steps of the fault location method provided in the above embodiment.

[0195] The descriptions of the above embodiments of the electronic devices and storage media are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the embodiments of the computer devices and storage media of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0196] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0197] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0198] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0199] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0200] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0201] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0202] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a controller to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0203] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A fault location method, characterized in that, include: When a node failure is detected in the data center, determine the target failure type corresponding to the failure. Determining the first subset of target features corresponding to the target fault type includes: Obtain feature data for each feature; The correlation between each feature and each fault type is calculated based on the feature data, and each feature is assigned a corresponding weight based on the correlation. The initial feature subset corresponding to each fault type is determined based on the weights corresponding to each feature. Based on the mutual information between each feature and the fault type in the initial feature subset and the mutual information between the features themselves, a first target feature subset corresponding to each fault type is determined from the initial feature subset, including: Calculate the mutual information between each feature in the initial feature subset and the fault type, and calculate the mutual information between each feature in the initial feature subset; The correlation measure between each feature and the fault type in the initial feature subset is calculated based on the mutual information between each feature and the fault type, and the correlation measure between features in the initial feature subset is calculated based on the mutual information between each feature in the feature subset. Based on the correlation measurement between each feature and the fault type and the correlation measurement between features in the initial feature subset, a feature selection algorithm is used to determine the first target feature subset corresponding to each fault type from the initial feature subset; Calculate the similarity between the features of each node in the data center and the features in the first target feature subset, and arrange the features of each node based on the similarity to obtain a feature set; A second target feature subset is determined from the feature set based on a feature selection algorithm; The location of the fault is determined based on the actual position of each feature in the second target feature subset at each node.

2. The method according to claim 1, characterized in that, When a node failure is detected in the data center, determining the target failure type includes: Monitor the characteristics of each node in the data center; If the data change corresponding to a feature is detected to exceed a preset threshold, the feature exceeding the preset threshold is identified as an abnormal feature to obtain an abnormal feature set. Compare the abnormal feature set with the first target feature subset corresponding to each fault type; The target fault type corresponding to the fault is determined by the fault type corresponding to the first target feature subset that is the same as the abnormal feature set.

3. The method according to claim 1, characterized in that, The feature selection algorithm determines the second target feature subset from the feature set, including: Determine the first data for each feature in the feature set; And determine the second and third data that are most adjacent to the first data; Calculate the first difference between the first data and the second data, and calculate the second difference between the first data and the third data; The weights of each feature in the feature set are determined based on the first difference and the second difference. Based on the weights of each feature, a subset of the second target features is determined from the feature set.

4. The method according to claim 3, characterized in that, The method further includes: Obtain the original data of the feature set; The original data is subjected to structured and standardized processing to obtain data for each feature, wherein there is no coupling relationship between the features.

5. The method according to claim 1, characterized in that, The method further includes: Obtain the current actual status data and target status data of the data center; The target state data is converted into principal component set values; The actual state data was dimensionality reduced using independent component analysis. The current principal component value is determined based on the actual state data after dimensionality reduction. The error value is determined based on the principal element setting value and the current principal element value; Based on the error value, determine the change information of the feature; The change information of the feature is sent to the data center so that the data center can perform fault-tolerant control based on the change information of the feature.

6. A fault location device, characterized in that, include: The monitoring module is used to determine the target fault type when a fault is detected in a node of the data center. The first determining module is used to determine a first target feature subset corresponding to the target fault type, including: Obtain feature data for each feature; The correlation between each feature and each fault type is calculated based on the feature data, and each feature is assigned a corresponding weight based on the correlation. The initial feature subset corresponding to each fault type is determined based on the weights corresponding to each feature. Based on the mutual information between each feature and the fault type in the initial feature subset and the mutual information between the features themselves, a first target feature subset corresponding to each fault type is determined from the initial feature subset, including: Calculate the mutual information between each feature in the initial feature subset and the fault type, and calculate the mutual information between each feature in the initial feature subset; The correlation measure between each feature and the fault type in the initial feature subset is calculated based on the mutual information between each feature and the fault type, and the correlation measure between features in the initial feature subset is calculated based on the mutual information between each feature in the feature subset. Based on the correlation measurement between each feature and the fault type and the correlation measurement between features in the initial feature subset, a feature selection algorithm is used to determine the first target feature subset corresponding to each fault type from the initial feature subset; The calculation module is used to calculate the similarity between the features of each node in the data center and each feature in the first target feature subset, and to arrange the features of each node based on the similarity to obtain a feature set; The second determining module is used to determine a second target feature subset from the feature set based on a feature selection algorithm; The location module is used to determine the location of the fault based on the actual location of each feature in the second target feature subset at each node.

7. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores a computer program that, when executed by the processor, performs the fault location method as described in any one of claims 1 to 5.

8. A storage medium, characterized in that, The storage medium stores a computer program, which is executed by a computer using the fault location method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Abnormal state monitoring and control system based on feature selection and principal component control

    CN109932904A

  • Data monitoring system, device and method

    CN113839827A