A method and system for cleaning operation and maintenance data of an intelligent platform
By training the correlation between fault types and operation and maintenance parameters, the optimal fault type is located and data correction is performed, which solves the problem of difficult fault identification in the operation and maintenance of data center equipment, and realizes efficient cleaning of operation and maintenance data and stable equipment operation.
Patent Information
- Application Number
- CN202511182927.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing technologies make it difficult to quickly identify the root cause of faults in data center equipment operation and maintenance, resulting in low efficiency of operation and maintenance data cleaning, inability to identify faults in a timely manner and perform effective cleaning, and affecting the normal operation of equipment.
By training the correlation between fault types and operation and maintenance parameters, the optimal fault type is located and data correction is performed. An intelligent platform operation and maintenance data cleaning method and system are established, including a training correlation module, an abnormal data judgment module, a fault location module and a data cleaning module. The correlation coefficient and correction model are used to clean the operation and maintenance data.
It enables accurate location and cleaning of operation and maintenance data, improves the accuracy of fault identification and the accuracy of dynamic acquisition of operation and maintenance data, and ensures the stable operation of data center equipment.
Smart Images

Figure CN120744323B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of operation and maintenance data processing technology, and relates to an intelligent platform operation and maintenance data cleaning method and system. Background Technology
[0002] Operations and maintenance data cleaning is a crucial step in ensuring the data quality of production and operations systems, and it is especially important in the era of intelligent manufacturing and the Industrial Internet. With the widespread application of AI technology in the industrial field, the volume of operations and maintenance data is growing exponentially, and data quality directly affects the accuracy of key business processes such as parameter optimization, quality analysis, security protection, and predictive maintenance.
[0003] In a data center environment, data cleaning refers to the process of identifying, correcting, and deleting raw data collected from data center equipment. Data center equipment is the core support for the efficient operation of a data center. Data center equipment mainly includes uninterruptible power supplies (UPS), servers, network equipment, storage devices, environmental monitoring equipment, and fire protection equipment. UPS primarily provides power backup for servers, network equipment, storage devices, environmental monitoring equipment, and fire protection equipment within the data center, ensuring that all equipment can continue to operate normally in the event of power failure. Servers, storage devices, and network communication equipment together constitute the infrastructure of the data center, ensuring that data processing, storage, and transmission can be carried out quickly and stably. The performance of these devices directly affects the accuracy of the data and is a crucial cornerstone of enterprise IT infrastructure development.
[0004] The prior art CN117575559A discloses an equipment operation and maintenance management system, including: an operation and maintenance personnel database for recording and updating operation and maintenance personnel information; a triggering module for parsing the user terminal information to obtain semantic information and matching the operation and maintenance command corresponding to the semantic information, and triggering a new work task according to the operation and maintenance command; a work order generation module for generating a corresponding work order according to the work task; an allocation module for allocating the work order to a suitable operation and maintenance personnel according to the work task and the skill information and work arrangement information of the operation and maintenance personnel in the operation and maintenance personnel database; a recording module for recording the execution process of the work task; a confirmation module for confirming the completion of the work task; an evaluation module for evaluating the completion effect of the work task and generating an operation and maintenance score; and an ending module for ending the work task.
[0005] If manual maintenance is used in the existing technology, it is difficult to identify the root cause of the fault from several abnormal operation and maintenance parameters. It also requires a lot of experience, which increases the difficulty and time of fault identification. It is impossible to identify faults in a timely manner based on the operation and maintenance parameters of each data center equipment, and to clean abnormal operation and maintenance data based on the relationship between faults and various operation and maintenance parameters. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent platform operation and maintenance data cleaning method and system, which solves the problems existing in the prior art.
[0007] The objective of this invention can be achieved through the following technical solutions:
[0008] A method for cleaning operation and maintenance data of an intelligent platform, which extracts operation and maintenance data corresponding to each operation and maintenance parameter under each fault type;
[0009] Analyze the changes in the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type to train the correlation between each fault type and each operation and maintenance parameter. The correlation is determined by the correlation coefficient. The correlation coefficient is determined based on the changes in the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type or the ratio between the changes in the operation and maintenance data corresponding to each operation and maintenance parameter and the operation and maintenance data of the operating parameter under normal operating conditions.
[0010] Extract the operation and maintenance data corresponding to each current operation and maintenance parameter and the operation and maintenance data under normal operation status, and analyze the abnormal operation and maintenance parameters;
[0011] Based on the correlation coefficient between different fault types and various operation and maintenance parameters, the optimal fault type corresponding to the current abnormal operation and maintenance parameter is located.
[0012] Based on the correlation coefficient between the optimal fault type and each operation and maintenance parameter, the operation and maintenance data corresponding to the operation and maintenance parameters that are correlated with the optimal fault type are corrected in order to clean the detected operation and maintenance data and obtain the corrected operation and maintenance data of each operation and maintenance parameter.
[0013] Preferably, the correlation coefficient between training each fault type and each operation and maintenance parameter is determined by judging the change of each operation and maintenance parameter under each fault type, comparing it with a set screening threshold, and / or comparing it with the operation and maintenance data of the corresponding operation and maintenance parameter under normal operation.
[0014] Preferably, the method for constructing a preliminary fault type combination based on the fault type corresponding to the current abnormal operation and maintenance parameters includes:
[0015] A1. Extract the fault types associated with the operation and maintenance parameters of each anomaly as the initial fault types, wherein the number of fault types associated with the operation and maintenance parameters of the anomaly is at least one.
[0016] A2. Compare the initial fault types associated with any two abnormal operation and maintenance parameters, and obtain at least two fault types associated with the operation and maintenance parameters of the two abnormal parameters as the initial screening fault types.
[0017] A3. Analyze the number of maintenance parameters that have a correlation coefficient greater than the set threshold between each initial screening fault type and all current abnormal maintenance parameters. Determine whether the number of maintenance parameter types with a correlation coefficient greater than the set threshold is less than the number of current abnormal maintenance data types. If it is less, proceed to the next initial screening fault type until the abnormal maintenance parameters corresponding to the selected initial screening fault type cover all current abnormal maintenance parameters or all initial screening fault types have been screened.
[0018] A4. Construct at least one set of initial screening fault type combinations, wherein the operation and maintenance parameter types corresponding to the initial screening fault type combinations cover all current abnormal operation and maintenance parameter types.
[0019] Preferably, when the number of initial screening fault type combinations is greater than or equal to 1, the method for determining the optimal initial screening fault type combination includes:
[0020] Analyze the variation coefficients corresponding to the abnormal operation and maintenance parameters, and filter out the operation and maintenance parameters corresponding to the largest variation coefficients;
[0021] The number of times the fault type associated with the operation and maintenance parameter with the largest coefficient of variation appeared in all the initial screening fault type combinations was statistically analyzed.
[0022] Based on the relationship between the frequency of occurrence and the frequency threshold, the fault type corresponding to at least one maintenance parameter with the largest variation coefficient is selected from the initial screening fault type combination and used as the optimal initial screening fault type combination.
[0023] Preferably, when the number of initial screening fault type combinations is greater than or equal to 1, reference abnormal operation and maintenance data corresponding to the operation and maintenance parameters that are related to the initial screening fault types are obtained;
[0024] In the initial screening of fault type combinations, the weight corresponding to the correlation coefficient between the operation and maintenance parameters of the same anomaly and each initial screening fault type is determined based on the correlation coefficient between the operation and maintenance parameters of the anomaly and each initial screening fault type.
[0025] Based on the reference abnormal operation and maintenance data corresponding to the operation and maintenance parameters of each abnormality, analyze all fault types in each initial screening fault type combination in turn, and predict the abnormal operation and maintenance data corresponding to the operation and maintenance parameters of each current abnormality.
[0026] For each initial screening fault type combination, calculate the difference between the predicted abnormal operation and maintenance data corresponding to the operation and maintenance parameters of each current abnormality and the operation and maintenance data corresponding to the operation and maintenance parameters of each current abnormality. Sum the differences and select the initial screening fault type combination with the smallest summed difference as the optimal initial screening fault type combination.
[0027] Preferably, the predicted abnormal operation and maintenance data corresponding to each abnormal operation and maintenance parameter is obtained by multiplying the weight of each initial screening fault type in the initial screening fault type combination with the reference abnormal operation and maintenance data of the abnormal operation and maintenance parameters under the initial screening fault type, and then summing them up.
[0028] Preferably, the operation and maintenance data corresponding to the operation and maintenance parameters that are correlated with the optimal fault type are corrected according to the correlation coefficient between the optimal fault type and each operation and maintenance parameter. This includes: obtaining the operation and maintenance data affected by the fault type at any time, and correcting the operation and maintenance data corresponding to the operation and maintenance parameters at the current time according to the number of optimal fault types and the correlation coefficient between each optimal fault type and each operation and maintenance parameter.
[0029] Preferably, the calibration model using operation and maintenance parameters is used to calibrate the operation and maintenance data corresponding to the operation and maintenance parameters that are correlated with the optimal fault type. This includes: obtaining the operation and maintenance data under the influence of the fault type at any given time and the operation and maintenance data under normal operating conditions; and calibrating the operation and maintenance data corresponding to the operation and maintenance parameters at the current time based on the duration of the fault type and the correlation coefficient between the optimal fault type and each operation and maintenance parameter.
[0030] An intelligent platform operation and maintenance data cleaning system includes:
[0031] The training correlation module is used to extract the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, analyze the change of the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, so as to train the correlation between each fault type and each operation and maintenance parameter. The correlation is determined by the correlation coefficient.
[0032] The abnormal data determination module is used to extract the operation and maintenance data corresponding to each current operation and maintenance parameter and the operation and maintenance data under normal operation status, and analyze the abnormal operation and maintenance parameters.
[0033] The fault location module determines the optimal combination of initial screening fault types based on the correlation coefficient between different fault types and various operation and maintenance parameters, and locates the optimal fault type corresponding to the current abnormal operation and maintenance parameter from the optimal combination of initial screening fault types.
[0034] The data cleaning module uses a calibration model for operation and maintenance parameters to calibrate the operation and maintenance data corresponding to the operation and maintenance parameters that are related to the optimal fault type, thereby cleaning the detected operation and maintenance data and obtaining the calibrated operation and maintenance data.
[0035] The beneficial effects of this invention are:
[0036] This invention analyzes the operational data corresponding to each current operational parameter to identify abnormal operational parameters. By combining the correlation coefficients between different fault types and each operational parameter, it selects the optimal initial fault type combination. Furthermore, it filters the optimal fault type corresponding to the current abnormal operational parameter from this optimal initial fault type combination. By locating the root cause of the abnormal operational parameter, it can accurately identify the optimal fault type. Based on the correlation coefficient between the optimal fault type and the current operational parameter, it cleans the operational data of the detected operational parameters correlated with the optimal fault type to eliminate the influence of fault types on the operational parameters. This cleaned operational data eliminates the deviation of operational data corresponding to operational parameters from normal operating conditions caused by faults in the data center equipment, thereby obtaining the actual operating data of the data center equipment.
[0037] This invention analyzes the changes in operation and maintenance data corresponding to the operation and maintenance parameters associated with each fault type, and trains the changes in operation and maintenance data corresponding to different fault types and operation and maintenance parameters to establish the correlation between different fault types and operation and maintenance parameters. This allows for the subsequent screening of fault types that are correlated with the operation and maintenance parameters based on the changes in operation and maintenance data corresponding to the currently detected operation and maintenance parameters.
[0038] This invention filters out the fault types associated with the currently abnormal operation and maintenance parameters, compares the initial screening fault types associated with any two abnormal operation and maintenance parameters, and filters out the initial screening fault types. This ensures that the operation and maintenance parameter types corresponding to all initial screening fault types in each combination cover the types of the currently abnormal operation and maintenance parameters, thereby obtaining all combinations of initial screening fault types. The invention then filters out the operation and maintenance parameter with the largest variation coefficient among the currently abnormal operation and maintenance parameters and determines the frequency of occurrence of the fault type associated with the operation and maintenance parameter with the largest variation coefficient in the initial screening fault type combinations. Based on the frequency of occurrence, the optimal initial screening fault type combination is selected from several groups of initial screening fault type combinations. This improves the screening effect and the accuracy of locating the optimal fault type.
[0039] This invention establishes reference abnormal operation and maintenance data for operation and maintenance parameters under abnormal conditions, compares it with the current abnormal operation and maintenance data, or uses the correlation coefficient of each initial screening fault type to the same operation and maintenance parameter to determine the weight of each initial screening fault type combination for a certain operation and maintenance parameter. This allows for the selection of the optimal initial screening fault type combination from several groups of initial screening fault type combinations. Furthermore, by combining multiple methods, the optimal fault type is determined from the optimal initial screening fault types. This enables the localization of the fault type causing the current abnormal operation and maintenance parameter. By employing dual constraints, the initial screening fault type combination is limited from several groups of initial screening fault type combinations, thereby determining the optimal fault type. This can accurately locate the root cause of the fault type causing the abnormal operation and maintenance data, improving the accuracy of fault type determination.
[0040] This invention corrects the operation and maintenance data corresponding to the operation and maintenance parameters that are correlated with the optimal fault type by using the correlation coefficient between the optimal fault type and each operation and maintenance parameter. This can eliminate the dynamic changes in the operation and maintenance data corresponding to the operation and maintenance parameters when a fault type exists. It can predict or correct the operation and maintenance data corresponding to each operation and maintenance parameter, improve the accuracy of dynamic acquisition of operation and maintenance data, and analyze the stability of each data center device based on the operation and maintenance parameters corresponding to the current operation and maintenance parameters, so as to control or maintain the data center device according to the predicted stability status of the data center device. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a schematic diagram of the operation and maintenance data cleaning method in this invention;
[0043] Figure 2 This is a schematic diagram illustrating how fault types cause abnormal operation and maintenance parameters in this invention.
[0044] Figure 3 This is a schematic diagram illustrating the fault types associated with operation and maintenance parameters in this invention;
[0045] Figure 4 This is a preliminary screening of fault type combinations for abnormal operation and maintenance parameters in this invention;
[0046] Figure 5 This is a schematic diagram of the method for locating the optimal fault type in this invention;
[0047] Figure 6 This is a schematic diagram of the method for analyzing the stability of computer room equipment according to the present invention. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] Uninterruptible power supplies (UPS) are used to ensure power quality and provide uninterrupted power supply. Through voltage stabilization and filtering functions, UPSs prevent equipment failures in the data center caused by mains power fluctuations and harmonics. Furthermore, in the event of a mains power outage, the UPS continues to supply power through batteries and inverters, ensuring the normal operation of other equipment in the data center. Data cleaning of data center equipment is crucial for ensuring the quality of data center equipment operation. If, during operation, data acquisition, transmission, or equipment malfunctions cause the monitored operating parameters of the data center equipment to deviate from normal operating conditions, it becomes impossible to analyze the actual operating status of the data center equipment and thus cannot provide reliable operational decisions.
[0050] The main equipment in a computer room includes uninterruptible power supplies (UPS), servers, network equipment, storage devices, environmental monitoring equipment, and fire-fighting equipment. UPS mainly provides power protection for the equipment in the computer room, such as servers, network equipment, storage devices, environmental monitoring equipment, and fire-fighting equipment, to ensure that each device can still operate normally in the event of a power outage.
[0051] When equipment in the data center malfunctions, it can also affect the UPS. For example, when servers, network devices, and storage devices are overloaded, the UPS output power may exceed the set range or internal components may overheat, accelerating equipment aging. When environmental monitoring equipment malfunctions, it affects the UPS's charging and discharging efficiency. When network devices or storage devices detect abnormal power parameters (unstable current or voltage), the UPS may be unable to adjust its output in a timely manner. When servers or network devices experience short circuits, overloads, or overheating, the UPS will activate its protection mechanism, cutting off power to the load and preventing it from continuing to provide the necessary power for the data center's equipment. When the UPS malfunctions, it can lead to problems such as server data loss, network interruptions, data loss or damage to storage devices, delayed data uploads from environmental monitoring devices, and failure to upload environmental data from fire protection equipment or to detect temperature or fires in a timely manner.
[0052] Therefore, if the detected operation and maintenance data is abnormal, it may cause the uninterruptible power supply (UPS) to make incorrect judgments and handle the situation, affecting the normal operation of the UPS and the normal operation of the equipment in the data center. It is necessary to identify the detected abnormal operation and maintenance parameters and process the operation and maintenance data corresponding to the abnormal operation and maintenance parameters based on the fault type.
[0053] like Figure 1 As shown, a method for cleaning operation and maintenance data of an intelligent platform includes the following steps:
[0054] Step 1: Train the change in operation and maintenance data corresponding to the operation and maintenance parameters associated with each fault type;
[0055] The operation and maintenance parameters of the uninterruptible power supply (UPS) in the computer room are monitored in real time. When different types of faults occur in the UPS, the change in the operation and maintenance data corresponding to each operation and maintenance parameter is detected. The change is relative to the operation and maintenance data under normal operation. Once a fault occurs, the operation and maintenance data corresponding to the operation and maintenance parameters associated with the fault type will change accordingly, deviating from the operation and maintenance data under normal operation, thus affecting the normal operation of the uninterruptible power supply (UPS).
[0056] The types of UPS (Uninterruptible Power Supply) faults include abnormal voltage, overload, overheating, battery malfunction, and automatic shutdown. Each fault type is affected by various maintenance data. Abnormal voltage may be caused by unstable mains voltage or a fault in the UPS's internal voltage regulation circuit. Overload may be caused by the load power exceeding the UPS's rated power, or by a short circuit or open circuit in the load itself. Overheating may be caused by a faulty cooling fan, high ambient temperature, or overload. Battery malfunction may be caused by battery aging or poor wiring. Automatic shutdown may be caused by overheating, overload, or a short circuit. These are common UPS faults encountered in actual operation and are issues that frequently require monitoring during daily management.
[0057] The training program tracks the changes in operational data corresponding to various operational parameters in the data center equipment under different fault types, establishes the correlation between different fault types and various operational parameters, and obtains the changes in operational data corresponding to operational parameters under different fault types.
[0058] The operation and maintenance parameters include mains voltage, UPS internal temperature, ambient temperature, and the power, current, voltage, and load temperature of the UPS and each load that has a power demand from the UPS (servers, network devices, storage devices, environmental monitoring equipment, fire-fighting equipment, etc.). Each operation and maintenance parameter is associated with at least one type of fault. Based on the operation and maintenance data corresponding to each operation and maintenance parameter, it is possible to directly or indirectly reflect whether the uninterruptible power supply has a fault or is about to have a fault.
[0059] Collect and extract the operation and maintenance data corresponding to previously detected operation and maintenance parameters, and analyze the impact of each fault type on the operation and maintenance data corresponding to the same operation and maintenance parameter. Analyze the operation and maintenance data of a certain operation and maintenance parameter under different fault types, and use a linear regression model to analyze the correlation between different fault types and the same operation and maintenance data. The correlation can be determined by the change in operation and maintenance data or the ratio of operation and maintenance data.
[0060] To ensure the accuracy of the above data analysis, it is necessary to ensure that the identified fault types and the detected operation and maintenance data are synchronized and their timestamps are aligned. Based on the duration of each fault type and the corresponding operation and maintenance data for each parameter during the fault's duration, a correlation between fault types and operation and maintenance parameters can be established. This correlation may change with the duration of the fault.
[0061] During the correlation analysis between fault types and operation and maintenance parameters, the operation and maintenance data corresponding to abnormal operation and maintenance parameters are cleaned and corrected to eliminate the impact of fault types on operation and maintenance parameters, thereby reducing the impact of abnormal or missing operation and maintenance data on the performance judgment of data center equipment.
[0062] Furthermore, based on the magnitude of the changes in operational data corresponding to each operational parameter under a certain fault type, operational parameters that are correlated with each fault type can be identified. The changes in operational data reflect the degree to which the operational parameter is affected by the fault type. If, when a certain fault type occurs, the operational data of a certain operational parameter remains relatively stable compared to the operational data under normal operating conditions, it indicates that the operational parameter is not affected by that fault type.
[0063] Alternatively, the ratio between the change in the operation and maintenance data corresponding to each operation and maintenance parameter under a certain fault type and the operation and maintenance data under the corresponding normal operation state can be used to determine the operation and maintenance data that are related to each fault type.
[0064] If we select the magnitude of the change in the operation and maintenance data corresponding to each operation and maintenance parameter under a certain fault type for correlation judgment, we can determine the screening threshold for the magnitude of the change based on experience or numerical values. That is, when the change is greater than the set screening threshold, it is considered to be correlated, and the correlation coefficient is set to 1; otherwise, there is no correlation, and the correlation coefficient is set to 0. The screening threshold can be set uniformly for each operation and maintenance data, or it can be set according to the change in the operation and maintenance data corresponding to each operation and maintenance parameter under different fault types in the historical data of previous detections.
[0065] If the correlation is determined by comparing the change in the maintenance data corresponding to each maintenance parameter under a certain fault type with the maintenance data of the same parameter under normal operating conditions, the ratio threshold can be determined based on the change in the maintenance data corresponding to each maintenance parameter under different fault types in the historical data of previous detections. This threshold can be adjusted by personnel based on experience or numerical values. When the ratio is greater than the set ratio threshold, a correlation is considered to exist, and the correlation coefficient is set to 1. Conversely, no correlation is considered to exist, and the correlation coefficient is set to 0. Similarly, the ratio threshold can be set according to actual needs or standards, which will not be explained in detail here.
[0066] In addition, the correlation coefficient can also be determined based on the ratio between the change in the operation and maintenance data corresponding to the operation and maintenance parameters and the maximum allowable change in the operation and maintenance parameters, thus determining the correlation coefficient between each fault type and each operation and maintenance parameter.
[0067] For the same operation and maintenance parameter, it is associated with at least one fault type. If the operation and maintenance data corresponding to one of the operation and maintenance parameters is abnormal, it will cause the operation and maintenance data of other operation and maintenance parameters to be abnormal, thereby indirectly causing multiple fault types, or directly affecting other fault types.
[0068] For example, if the load is too high, it will cause the UPS to overheat and shut down automatically, which will affect the UPS's ability to provide power to other equipment and prevent them from operating normally.
[0069] Step 2: Obtain the current operation and maintenance parameters, analyze the operation and maintenance data, identify abnormal operation and maintenance data, and locate the optimal fault type corresponding to the current abnormal operation and maintenance parameters based on the correlation coefficient of operation and maintenance parameters under different fault types.
[0070] Since at least one fault type is associated with the same operation and maintenance parameter, when an anomaly is detected in the operation and maintenance data of at least one operation and maintenance parameter, several possible fault types are filtered out based on the correlation between the operation and maintenance parameter and the fault type, and the optimal fault type is located from several fault types. Compared with directly determining the existing fault type based on the fault type associated with the abnormal operation and maintenance parameter, this greatly improves the accuracy of fault identification and reduces misjudgment of fault type determination due to interference between faults.
[0071] like Figure 2 and 3 As shown, where, Figure 3 It reflects the abnormal operation and maintenance parameters and the fault types associated with the abnormal operation and maintenance parameters.
[0072] Specifically, at least one fault type will cause an anomaly in one maintenance data. Relying solely on the correlation between each maintenance data and the fault type is insufficient to accurately pinpoint the fault type. Alternatively, the existence of a fault type may cause an anomaly in the maintenance data of a certain maintenance parameter 'a' (in this case, the maintenance data deviates from the maintenance data under normal operating conditions, resulting in data deviation). After an anomaly occurs in maintenance parameter 'a', it will cause an anomaly in the maintenance data of another maintenance parameter 'b'. If only the correlation between each maintenance parameter and the fault type is used, the fault types corresponding to maintenance parameters 'a' and 'b' can be located, resulting in at least one fault type. This makes it impossible to accurately pinpoint the actual fault type, i.e., it is impossible to locate the source of the fault from the fault types corresponding to maintenance parameter 'a'.
[0073] When there are multiple operational data anomalies corresponding to multiple operational parameters, there are relatively many fault types that are related to each operational parameter. It is difficult to locate the best fault type that caused the abnormal operational parameter from multiple fault types. However, the following method can be used to filter out the optimal fault type from the fault types corresponding to multiple abnormal operational parameters.
[0074] like Figure 5 As shown, based on the fault types corresponding to the current abnormal operation and maintenance parameters, a preliminary combination of fault types is constructed, and then the optimal fault type is located. The method includes the following steps:
[0075] Step 21: Filter out the operation and maintenance data corresponding to the current abnormal operation and maintenance parameters, and extract the fault types associated with each abnormal operation and maintenance parameter as the initial fault types; the number of fault types associated with the abnormal operation and maintenance parameters is at least one.
[0076] Step 22: Compare the initial fault types associated with the operation and maintenance parameters of any two anomalies, and obtain the fault types associated with the operation and maintenance parameters of at least two anomalies as the initial screening fault types.
[0077] Step 23: Analyze the number of maintenance parameter types whose correlation coefficient with each initially screened fault type is greater than a set threshold (the set threshold is less than the first correlation coefficient s2 and greater than the second correlation coefficient s1) with all current abnormal maintenance parameters. Determine whether the number of maintenance parameter types with a correlation coefficient greater than the set threshold is less than the number of current abnormal maintenance parameter types. If it is less, proceed to the next initially screened fault type until the abnormal maintenance parameter types corresponding to the selected initially screened fault types cover all current abnormal maintenance parameter types or all initially screened fault types have been screened.
[0078] The threshold is used to determine the correlation coefficients in the calculation. It is determined based on the correlation coefficients between each fault type and the changes in the corresponding operation and maintenance data of each operation and maintenance parameter. The purpose is to classify the correlation coefficients.
[0079] The correlation coefficient can be determined by the correlation between the fault type and each operation and maintenance data. That is, whether the operation and maintenance data of one operation and maintenance parameter deviates from the operation and maintenance data of the operation and maintenance parameter under normal operation will trigger at least one fault type, or the change of the operation and maintenance data corresponding to each operation and maintenance parameter when the fault occurs. When a certain fault type occurs, if the change in the operation and maintenance data corresponding to the operation and maintenance parameter is less than a set first numerical threshold, or the ratio between the change in the operation and maintenance data of the operation and maintenance parameter under normal operation is less than a first ratio threshold, then the correlation coefficient is considered small, and the correlation coefficient is set to s1. Conversely, if the change in the operation and maintenance data corresponding to the operation and maintenance parameter is not less than the set first numerical threshold, or the ratio between the change in the operation and maintenance data of the operation and maintenance parameter under normal operation is not less than the first ratio threshold, then the correlation coefficient is considered large, and the correlation coefficient is set to s2.
[0080] Among them, s1 and s2 are used to reflect the degree of correlation between fault type and operation and maintenance parameters. Specific values can be set according to actual needs. s2 is greater than s1, and s2 can be set to be close to or equal to the upper limit value of the correlation coefficient. s1 is close to or equal to the lower limit value of the correlation coefficient, so that operation and maintenance parameters and each fault type can be distinguished based on the correlation coefficient.
[0081] Alternatively, the correlation coefficient between the maintenance parameters and the fault type can be determined by the ratio between the change in maintenance data corresponding to each fault type and the maintenance data of the maintenance parameters under normal operating conditions.
[0082] The explanation of the abnormal operation and maintenance parameter types corresponding to the initially screened fault types covering all current abnormal operation and maintenance parameter types is as follows: This involves sequentially screening operation and maintenance parameters for a specific initially screened fault type that have a correlation coefficient greater than a set threshold with all current abnormal operation and maintenance parameters. The types of operation and maintenance parameters with a correlation coefficient greater than the set threshold are then counted. These counted operation and maintenance parameter types include the types of abnormal operation and maintenance parameters. For example, if there are n types of current abnormal operation and maintenance parameters, and m types of operation and maintenance parameters with a correlation coefficient greater than the set threshold for the initially screened fault type selected in step 23 (m ≥ n), and all types of current abnormal operation and maintenance parameters fall within the range of the selected operation and maintenance parameter types with a correlation coefficient greater than the set threshold for the initially screened fault type.
[0083] Step 24: Construct a set of initial screening fault types that cover all currently abnormal operation and maintenance parameter types, and select the optimal fault type from the set of initial screening fault types based on the operation and maintenance data corresponding to the abnormal operation and maintenance parameters.
[0084] like Figure 4 As shown, Figure 4 against Figure 3 The initial screening fault type combination is formed by the fault types associated with the abnormal operation and maintenance parameters. In a set of initial screening fault type combinations, the combination of operation and maintenance parameters corresponding to each initial screening fault type includes the operation and maintenance parameters of the current operation and maintenance data that are abnormal. That is, the types of operation and maintenance parameters of the current operation and maintenance data that are abnormal belong to a subset of the types of operation and maintenance parameters corresponding to all initial screening fault types in the set of initial screening fault type combinations.
[0085] There is at least one combination of initial screening fault types, and in each combination of initial screening fault types, there is at least one different fault type or a different number of fault types.
[0086] When the number of initial screening fault type combinations is greater than 1, the optimal initial screening fault type combination is further selected from the initial screening fault type combinations, and then the optimal fault type is selected from the optimal initial screening fault type combination, so as to accurately locate the fault type that causes each abnormal operation and maintenance data.
[0087] Based on the fault types with the highest correlation, the combination of fault types with the highest correlation is selected. Specifically, when the number of fault type combinations is greater than 1, the operation and maintenance data corresponding to the current abnormal operation and maintenance parameters is compared with the operation and maintenance data of the operation and maintenance parameters under normal operation and maintenance conditions. The variation coefficient of each operation and maintenance parameter is calculated, and the operation and maintenance parameter with the largest variation coefficient is extracted.
[0088] Wherein, the coefficient of variation is equal to the absolute value of the difference between the maintenance data of the abnormal maintenance parameter and the maintenance data of the maintenance parameter under normal operation, and the ratio between the maintenance data of the maintenance parameter under normal operation.
[0089] Determine the number of times the fault type associated with the maintenance parameter with the largest variation coefficient appears in the initial screening of fault type combinations;
[0090] If the number of occurrences equals the number of occurrences threshold, then the initial screening fault type combination corresponding to the fault type associated with the maintenance parameter with the largest variation coefficient in the initial screening fault type combination is taken as the optimal initial screening fault type combination. The number of occurrences threshold is set to 1, which limits the number of fault types that occur in the initial screening fault type combination. Here, the preferred value is 1.
[0091] If the number of occurrences exceeds the threshold, the fault type corresponding to the next maintenance parameter with the largest variation coefficient is extracted (excluding maintenance parameters with the largest variation coefficient). The fault types corresponding to the two maintenance parameters with the largest variation coefficients are used to filter the initial screening fault type combinations to determine the optimal initial screening fault type combination.
[0092] When the number of occurrences exceeds the threshold, the optimal combination of initial screening fault types is selected, including the fault types corresponding to the maintenance parameters with the largest and second largest variation coefficients.
[0093] By filtering the magnitude of the variation coefficients of operation and maintenance parameters, the fault type associated with the operation and maintenance parameter with the largest variation coefficient is selected. Based on this fault type, several initial screening fault type combinations are further filtered. Combining the number of initial screening fault type combinations that meet the conditions, the screening conditions are further determined so as to select the best initial screening fault type combination from several initial screening fault type combinations, in order to prepare for selecting the optimal fault type from the optimal initial screening fault type combination.
[0094] In addition, when the number of initial screening fault type combinations is greater than 1, this invention discloses another method for screening initial screening fault type combinations. This method involves extracting abnormal operation and maintenance parameters and their correlation coefficients with each fault type, extracting operation and maintenance parameters with correlation coefficients greater than a set threshold, obtaining reference abnormal operation and maintenance data for these parameters under abnormal conditions, comparing the reference abnormal operation and maintenance data with the current abnormal operation and maintenance data, and selecting the operation and maintenance parameter with the smallest deviation rate between the current abnormal operation and maintenance data and the reference abnormal operation and maintenance data (the ratio of the difference between the reference abnormal operation and maintenance data and the current abnormal operation and maintenance data to the reference abnormal operation and maintenance data is used as the deviation rate). Then, it extracts fault types with correlation coefficients greater than a set threshold with the operation and maintenance parameter with the smallest deviation rate, and based on this fault type, selects the initial screening fault type combination with the highest correlation from the initial screening fault type combinations.
[0095] This invention establishes reference abnormal operation and maintenance data for operation and maintenance parameters under abnormal conditions, and compares it with the current abnormal operation and maintenance data. It then filters out the fault types where the correlation coefficient between operation and maintenance parameters with the smallest deviation rate is greater than a set threshold. By using multiple filtering criteria such as correlation coefficient and deviation rate, it can effectively eliminate the filtering interference of operation and maintenance parameters with a lower than the set threshold and the interference of operation and maintenance parameters with a large deviation rate. While simplifying the amount of data, it can quickly and relatively accurately filter out the expected combination of initial fault types.
[0096] In addition, when the number of initial screening fault type combinations is greater than 1, this invention discloses another method for screening fault type combinations, which can compare abnormal operation and maintenance parameters with the operation and maintenance parameters corresponding to each initial screening fault type combination to determine the initial screening fault type combination with the highest matching degree.
[0097] In the initial screening fault type combination, if any initial screening fault type is correlated with at least one abnormal operation and maintenance parameter, then the reference abnormal operation and maintenance data corresponding to the operation and maintenance parameter correlated with the initial screening fault type (the reference abnormal operation and maintenance data is the average value of the operation and maintenance data corresponding to the abnormal operation and maintenance parameter under a certain fault type) is obtained as the predicted abnormal operation and maintenance data corresponding to the abnormal operation and maintenance parameter. If at least two initial screening fault types are correlated with the same abnormal operation and maintenance parameter, the weight of the correlation coefficient is determined. According to the initial screening fault type, the reference abnormal operation and maintenance data corresponding to the operation and maintenance parameter with the correlation coefficient greater than a set threshold is analyzed. The product of the weight of each initial screening fault type in the initial screening fault type combination and the reference abnormal operation and maintenance data of the abnormal operation and maintenance parameter under the initial screening fault type is used and accumulated to obtain the result. Calculate the predicted abnormal operation and maintenance data for each fault type in each initial screening fault type combination for the current abnormal operation and maintenance parameters, and calculate the sum of the differences between the predicted abnormal operation and maintenance data of each abnormal operation and maintenance parameter and the corresponding abnormal operation and maintenance data. Select the initial screening fault type combination with the smallest sum of differences as the initial screening fault type combination with the highest matching degree, i.e., the optimal initial screening fault type combination.
[0098] The weight of the correlation coefficient ,in, , where aij represents the correlation coefficient between the j-th initial screening fault type and the i-th abnormal operation and maintenance parameter, and p represents the number of abnormal operation and maintenance parameters.
[0099] Predict abnormal operation and maintenance data Ms, t=1,2,...,p, Ms=ui1*Reference abnormal operation and maintenance data corresponding to the i-th operation and maintenance parameter under the first initial screening fault type +ui2*Reference abnormal operation and maintenance data corresponding to the i-th operation and maintenance parameter under the second initial screening fault type +......+uis*Reference abnormal operation and maintenance data corresponding to the i-th operation and maintenance parameter under the s-th initial screening fault type.
[0100] This method takes into account the impact of various fault types on the operation and maintenance data of each operation and maintenance parameter. It uses the correlation coefficient of each initially screened fault type on the same operation and maintenance parameter to determine the weight of each initially screened fault type combination on a certain operation and maintenance parameter. It also comprehensively analyzes the joint impact of all fault types in each initially screened fault type combination on the operation and maintenance data of each abnormal operation and maintenance parameter. This enables a comprehensive comparison of the operation and maintenance data of each operation and maintenance parameter under the influence of all fault types in the initially screened fault type combination, so as to screen out the initially screened fault type combination with the highest matching degree, laying the foundation for the subsequent selection of the optimal fault type, thereby reducing data processing and improving screening efficiency.
[0101] The method of selecting the most matching combination of initial screening fault types from multiple initial screening fault type combinations is not limited to the three methods mentioned above.
[0102] When there is only one initial screening fault type combination, the current initial screening fault type combination, as the initial screening fault type combination with the highest matching degree, needs to be selected as the optimal fault type from the initial screening fault type combination.
[0103] Methods for selecting the optimal fault type from the initial screening fault type combinations with the highest matching degree include:
[0104] Step A1: Extract the variation coefficients corresponding to the deviations of the operation and maintenance parameters of each anomaly from the operation and maintenance data under normal operating conditions;
[0105] Step A2: Determine whether the difference between the variation coefficient corresponding to each abnormal operation and maintenance parameter and the set variation coefficient threshold corresponding to the corresponding operation and maintenance parameter is within the allowable first error range. If it is within the allowable first error range, extract the operation and maintenance parameters within the allowable error range as the first operation and maintenance parameter. The allowable first error range is used to limit the degree to which the variation coefficient corresponding to each abnormal operation and maintenance parameter deviates from the set variation coefficient threshold corresponding to the corresponding operation and maintenance parameter. It can be set by R&D or testing personnel according to actual needs.
[0106] Among them, the threshold of the variation coefficient corresponding to the set operation and maintenance parameters is: the fault type with a correlation coefficient between the operation and maintenance parameters and each fault type that is greater than the set threshold. When a fault occurs, the difference between the reference abnormal operation and maintenance data of the operation and maintenance parameters and the operation and maintenance data of the operation and maintenance parameters under normal operation status, and the ratio between the operation and maintenance data of the operation and maintenance parameters under normal operation status.
[0107] Step A3: Determine if the number of the first maintenance parameters is less than 1. If yes, filter out the maintenance parameters with the smallest difference between the variation coefficient corresponding to each abnormal maintenance parameter and the set variation coefficient threshold, and proceed to step A5. If no, determine if the number of the first maintenance parameters is equal to 1. If yes, proceed to step A5. Otherwise, proceed to step A4, indicating that the current abnormal maintenance data is caused by at least two fault types.
[0108] Step A4: Select the fault type with the largest sum of correlation coefficients with each of the first maintenance parameters from the fault type combination with the highest matching degree, and use it as the optimal fault type. Remove the first maintenance parameters that are related to the optimal fault type from all the first maintenance parameters, and use them as the second maintenance parameters. Select the fault type with the largest correlation coefficient with the second maintenance parameter from the fault type combination with the highest matching degree, and use it as the optimal fault type.
[0109] Step A5: Select the fault type with the highest correlation coefficient with the first maintenance parameter from the fault type combination with the highest matching degree, and use it as the optimal fault type.
[0110] By comparing the variation coefficient of abnormal operation and maintenance parameters with the set variation coefficient threshold of the corresponding operation and maintenance parameters, the most influential first operation and maintenance parameter can be selected from several abnormal operation and maintenance parameters. Combined with the number of first operation and maintenance parameters, the corresponding optimal fault type screening strategy can be selected. This allows for the analysis of at least one optimal fault type caused by multiple fault types from at least one abnormal operation and maintenance parameter. In this way, the root cause of the fault type that causes abnormal operation and maintenance data can be accurately located, improving the accuracy of fault type determination.
[0111] When multiple abnormal operation and maintenance parameters fail to identify the optimal fault type—that is, the actual fault causing the abnormal operation and maintenance layer—it is easy to misidentify other faults and miss the actual fault type, affecting the accuracy of fault type localization and making it difficult to obtain the root cause fault from several operation and maintenance parameters. The method described above can analyze the associated fault type combinations from multiple abnormal operation and maintenance parameters, and further determine the optimal fault type from these combinations, achieving precise fault type localization.
[0112] Among them, multiple abnormal operation and maintenance parameters may be caused by a fault type, which may lead to an abnormal operation and maintenance parameter, and then the abnormal operation and maintenance parameter may lead to other abnormal operation and maintenance parameters. In other words, multiple abnormal operation and maintenance parameters may be caused by the same fault type.
[0113] This invention discloses another method for selecting the optimal fault type from a preliminary fault type combination. The method for locating the optimal fault type (at least one abnormal operation and maintenance parameter exists) includes the following steps:
[0114] Step B1: Filter out the operation and maintenance parameters that are abnormal in the current operation and maintenance data, and extract the fault type associated with each abnormal operation and maintenance parameter.
[0115] Step B2: Extract the variation coefficients corresponding to all current abnormal operation and maintenance parameters. The variation coefficient is equal to the ratio of the absolute value of the difference between the operation and maintenance data of the abnormal operation and maintenance parameter and the operation and maintenance data of the operation and maintenance parameter under normal operation to the operation and maintenance data of the operation and maintenance parameter under normal operation.
[0116] Step B3: Filter out the operation and maintenance parameters with the largest variation coefficient, and based on the operation and maintenance parameters, filter out the initial screening fault types that have a correlation coefficient with the operation and maintenance parameters greater than a set threshold, as the first initial screening fault types; the number of the first initial screening fault types is greater than or equal to 1. This method can filter out the fault types that are relatively likely to cause abnormal operation and maintenance parameters from several initial screening fault types, so as to facilitate the next step of precise screening.
[0117] Step B4: Extract the first initial screening fault type with the largest correlation coefficient with the operation and maintenance parameters, and extract the correlation coefficient between the first initial screening fault type and other abnormal operation and maintenance parameters. Based on the correlation coefficient, analyze the operation and maintenance data of other abnormal operation and maintenance parameters under the interference of the first initial screening fault type.
[0118] Among them, other abnormal operation and maintenance parameters are affected by the first initial screening fault type, that is, the product between the first initial screening fault type and the operation and maintenance data corresponding to the operation and maintenance parameters under normal operation status.
[0119] Step B5: Determine whether the difference between the operation and maintenance data corresponding to other abnormal operation and maintenance parameters under the first initial screening fault type and the operation and maintenance data corresponding to the current abnormal operation and maintenance parameter is within the allowable second error range. If it is within the allowable second error range, the first initial screening fault type is taken as the optimal fault type. If it exceeds the allowable second error range, the next initial screening fault type with a correlation coefficient greater than the set threshold with the operation and maintenance parameter is selected. Steps B4 and B5 are executed until the optimal fault type is determined.
[0120] If no optimal fault type that meets the above-mentioned allowable error range requirement is found according to steps B1-B5 above, the second error range can be adjusted so that only one initially screened fault type meets the requirement.
[0121] In addition, when two or more fault types exist simultaneously, it is impossible to calculate an optimal fault type using the above method. If necessary, one initial fault type can be selected first. Then, based on the difference between the operation and maintenance data corresponding to the abnormal operation and maintenance parameters caused by the selected initial fault type and the current abnormal operation and maintenance parameters, the next initial fault type can be selected based on the above steps B1-B5, and so on. Generally, the possibility of more than two fault types existing simultaneously is small. In the actual operation and maintenance data monitoring process, one or two fault types often occur. When one fault type exists and is not resolved in time, it may lead to more than two fault types.
[0122] The present invention also discloses another method: if the number of current abnormal operation and maintenance parameters is 1, extract the fault type associated with the abnormal operation and maintenance parameter, compare the operation and maintenance data under the influence of the fault type on the operation and maintenance parameter with the current abnormal operation and maintenance data, calculate the difference between the current abnormal operation and maintenance data and the operation and maintenance data corresponding to the operation and maintenance parameter under each associated fault type, and extract the fault type corresponding to the smallest difference as the optimal fault type.
[0123] The above method can filter out the fault types associated with the abnormal operation and maintenance parameters from the initial screening of fault type combinations when the operation and maintenance parameters are abnormal, thereby identifying the fault source of the abnormal operation and maintenance parameters and making it easier to filter out the fault source that caused the abnormal data and achieve accurate fault type location.
[0124] The above analysis of abnormal operation and maintenance parameters, based on the relationship between each initial screening fault type combination and each abnormal operation and maintenance parameter, selects the initial screening fault type combination with the highest matching degree. Further analysis of the abnormal operation and maintenance parameters identifies the optimal fault type from the initial screening fault type combination with the highest matching degree. By using the initial screening fault type combination as a constraint to determine the optimal fault type, the situation where the optimal fault type cannot be accurately located due to the correlation between fault types or the correlation between the same fault type and multiple operation and maintenance parameters can be avoided, greatly improving the accuracy of fault location.
[0125] Step 3: Based on the optimal fault type, the operation and maintenance data of the operation and maintenance parameters associated with the optimal fault type are corrected to clean the detected operation and maintenance data and obtain corrected operation and maintenance data, which improves the accuracy of operation and maintenance data detection and reduces the interference of fault type on each operation and maintenance data.
[0126] Correction strategies for operation and maintenance data, k = 1, 2, ..., q, where q represents the number of operational parameters and m represents the number of optimal fault types. If there is at least one optimal fault type in the system, interference exists between the optimal fault types. This is represented by the correlation coefficient between the nth optimal fault type and the kth maintenance parameter. The value can be greater than 0, less than 0, or equal to 0. This represents the maintenance data for the k-th maintenance parameter under the influence of the optimal fault type at time t. The operation and maintenance data is represented as the corrected k-th operation and maintenance parameter. If the correlation coefficient between the optimal fault type and the k-th operation and maintenance parameter is equal to 0, the operation and maintenance parameter is not affected by the optimal fault type.
[0127] Since the operation and maintenance data of the operation and maintenance parameters are dynamic values, the impact of the optimal fault type on the operation and maintenance parameters can be eliminated or reduced through the dynamic correction of the operation and maintenance data mentioned above. This allows us to obtain the operation and maintenance data of the corrected dynamic operation and maintenance parameters and ensures that the operation and maintenance data corresponding to the corrected dynamic operation and maintenance parameters are close to the real operation and maintenance data.
[0128] In addition, the detected maintenance data can be corrected based on the maintenance data corresponding to each maintenance parameter under normal operating conditions, and the maintenance data corresponding to each maintenance parameter under the optimal fault type. The correction strategy is as follows: , This represents the maintenance data corresponding to the k-th maintenance parameter under the influence of the optimal fault type at time t. This represents the average maintenance data for the k-th maintenance parameter under normal operating conditions. Represented as the first The maintenance data after correction of the k-th maintenance parameter under the influence of fault type at a given time, preferably, when This represents the maintenance data corresponding to the k-th maintenance parameter at time t after the occurrence of the optimal fault type. Synchronization is represented as the period after the occurrence of the optimal fault type. The maintenance data after the k-th maintenance parameter is corrected under the influence of the fault type at any given time.
[0129] The second method for correcting operational parameters is based on the rate of change of operational data for each operational parameter under the duration of a fault type. It corrects the operational data corresponding to each operational parameter at a certain moment, making the corrected operational data of each operational parameter closer to the actual operational data. Compared with the first method, the second method is better because it integrates the operational data of each operational parameter under the combined influence of various optimal fault types.
[0130] When a fault type exists, the operation and maintenance data of the operation and maintenance parameters will change as the fault type persists. The correction of the operation and maintenance data corresponding to the above-mentioned operation and maintenance parameters can eliminate the dynamic changes of the operation and maintenance data corresponding to the operation and maintenance parameters when a fault type exists. It can predict or correct the operation and maintenance data corresponding to the operation and maintenance parameters at the next moment, thereby improving the accuracy of dynamic acquisition of operation and maintenance data.
[0131] In cases where no fault occurs or no fault type exists, if the maintenance data of the maintenance parameters becomes abnormal or is lost due to transmission anomalies, the abnormal maintenance data is discarded. The discarded abnormal maintenance data is then filled in with maintenance data from before and after the maintenance parameter detection, thus cleaning the maintenance data of the abnormal maintenance parameters.
[0132] This method also includes data model analysis of the corrected operation and maintenance parameters to assess the current stability of the data center equipment. Based on this, the data center equipment is optimized and controlled to achieve intelligent operation and control of the data center equipment, thereby avoiding damage to the required equipment (servers, network equipment, and storage devices in the data center) or affecting the normal operation of the uninterruptible power supply (UPS) or other data center equipment due to damage to the data center equipment.
[0133] To account for the use of space and computing resources, a trigger time is set to analyze the stability of the operation and maintenance parameters of the data center equipment for a fixed duration, thereby achieving timed monitoring.
[0134] Trigger conditions can also be set to trigger a stability analysis of the operation and maintenance parameters of the data center equipment when abnormalities are detected, thereby enabling monitoring of the stability status that meets the trigger conditions.
[0135] Assess the stability of the data center equipment (UPS and other equipment requiring power) under the current operational data. Abnormal operational parameters of the equipment requiring power may affect the operation of the UPS. For example, when abnormal operational parameters are detected, the UPS will trigger power demand control for that equipment. For instance, if the temperature of a piece of equipment rises, continued power supply will cause the temperature to rise further, potentially damaging the equipment and causing the UPS to adjust its power demand, which may affect the normal operation of the UPS. Furthermore, abnormal operational data in the UPS system can lead to damage to the UPS itself, directly impacting the equipment requiring power from the UPS.
[0136] like Figure 6 As shown, based on the operational data corresponding to the current abnormal operational parameters, the stability of each data center device is analyzed. The analysis method includes the following steps:
[0137] W1. Train the interference coefficient between the operation and maintenance parameters of each data center device, that is, the change in the operation and maintenance data corresponding to the operation and maintenance parameters of a certain data center device affects the change in the operation and maintenance data corresponding to other operation and maintenance parameters of the same data center device or other data center devices. When the operation and maintenance data corresponding to a certain operation and maintenance parameter changes, does the operation and maintenance data of other operation and maintenance parameters change?
[0138] When the operation and maintenance data corresponding to the operation and maintenance parameter 'a' of a certain data center device becomes abnormal (the operation and maintenance data of the operation and maintenance parameter changes), train the change amount of other operation and maintenance data of this device and the operation and maintenance data corresponding to the operation and maintenance parameters of other devices.
[0139] △ai=λ (ai,bj) *△bj; △bj represents the change in the operation and maintenance data of the j-th operation and maintenance parameter, that is, the difference between the current operation and maintenance data of the j-th operation and maintenance parameter and the operation and maintenance parameter under normal operation; △ai represents the change in the data of the i-th operation and maintenance parameter. Similarly, i and j belong to any one of the operation and maintenance parameters, and the case where i equals j is not considered. The range of values for i and j is 1-n.
[0140] λ (ai,bj)Let λ represent the interference coefficient of the i-th operation and maintenance parameter on the j-th operation and maintenance parameter. Since the change in the operation and maintenance data corresponding to the operation and maintenance parameter is affected by the change in the operation and maintenance data corresponding to other operation and maintenance parameters, the impact of the change in the operation and maintenance data of the i-th parameter on the change in the data of other operation and maintenance parameters will vary depending on the change in the operation and maintenance data of the i-th parameter. There are cases where the interference coefficient between two operation and maintenance parameters approaches a fixed value, and there are cases where the interference coefficient between two operation and maintenance parameters changes with different dependent variables. Data fitting is used to establish a variable interference function λ. (ai,bj) =f(△ai).
[0141] W2. Extract the current abnormal operation and maintenance parameters and analyze the interference coefficient between the current abnormal operation and maintenance parameters and the operation and maintenance parameters of each data center equipment.
[0142] Based on the operational data corresponding to the current abnormal operational parameters under normal operating conditions, determine the amount of data change of the current abnormal operational parameters. Based on the variable interference function, determine the interference coefficient between the current abnormal operational parameters and other operational parameters. Then, determine the interference coefficient between each abnormal operational parameter and the operational parameters that are related to each data center equipment.
[0143] W3. Calculate the comprehensive interference coefficient of a maintenance parameter under the same equipment in the same computer room;
[0144] Since various operation and maintenance parameters influence each other, the same operation and maintenance parameter under the same data center equipment is affected by at least one other operation and maintenance parameter. In order to improve the accuracy of the operation and maintenance parameter analysis of the data center equipment to be evaluated and stabilized, the comprehensive interference degree of each abnormal operation and maintenance parameter on each operation and maintenance parameter in the data center equipment is analyzed.
[0145] Calculate the impact of the current abnormal change in the operation and maintenance parameters on the operation and maintenance parameters of each data center device, and calculate the comprehensive interference coefficient of the j-th operation and maintenance parameter of data center device T. The comprehensive interference coefficient has dimensions and reflects the degree of interference. △ai represents the change in the operation and maintenance data of the i-th operation and maintenance parameter. When △ai is greater than 0, it indicates that △ai is an abnormal operation and maintenance parameter.
[0146] W4. Based on the stability assessment model, analyze the stability coefficient of each data center device under the current operation and maintenance parameters. Analyze whether the stability coefficient of each data center device is less than the first threshold. If it is less than the first threshold, it indicates that the current abnormal operation and maintenance parameters have no impact on the normal operation of the data center device. If it is greater than the second threshold, it indicates that the current abnormal operation and maintenance parameters directly affect the normal operation of the data center device. Conversely, it indicates that the current abnormal operation and maintenance parameters may affect the normal operation of the data center device. Further monitoring of the operation and maintenance parameters of the data center device is required in the future. Based on the monitoring results, the second threshold should be adjusted to improve the accuracy of judging the stability of each data center device during operation.
[0147] The stability assessment model is as follows: ; Let be the stability coefficient of the T-th computer room device. Let x be the correlation coefficient between the j-th maintenance parameter and the x-th fault type. This represents the weight of the x-th fault type in the process of causing a fault in the data center equipment. The weight is based on the ratio of the number of times the x-th fault type occurs to the total number of faults during the process of the data center equipment malfunctioning.
[0148] By using the stability assessment model, the stability coefficient of each data center device is obtained. The stability coefficient is used to assess the stability status of the data center device and reflects the possibility of the data center device maintaining a stable and normal operating state under the current operation and maintenance parameters.
[0149] The first threshold is less than the second threshold. It is used to assess the impact of current abnormal operation and maintenance parameters on the normal operation of each data center equipment, so as to determine and predict the stability status of each data center equipment. In order to control the data center equipment with a value greater than the second threshold in a timely manner based on the relationship between the stability coefficient of each data center equipment and the first and second thresholds, so as to avoid triggering continuous failures.
[0150] The first and second thresholds are empirical values set by R&D / testing personnel when judging the stability of the above-mentioned equipment. These values are derived from experience. When the set values are small, the equipment's stability may be predicted to be poor even if no abnormality occurs. When the set values are large, the equipment's stability may be predicted to be good even if an abnormality is highly likely. This affects the accuracy of the data judgment.
[0151] Based on the value range corresponding to the stability coefficient of each device, the stability status of each data center device under the current operation and maintenance parameters can be determined, and the data center device can be controlled or maintained according to the predicted stability status of the data center device.
[0152] To facilitate timely maintenance of data center equipment, administrators analyze the stability coefficients of each data center device under current abnormal operational parameters, select the device with the lowest stability coefficient, and then, based on the identified optimal fault type, impose constraints on fault repair for that optimal fault type, requiring the following conditions to be met:
[0153] formula: , The set time threshold is the safe constraint time for fault repair, and in conjunction with the current stability coefficient of the equipment in the computer room, it can handle faults in a timely manner.
[0154] By selecting the lowest stability coefficient from the stability coefficients corresponding to each data center device, and combining it with the optimal fault type located by the current abnormal operation and maintenance parameters, the system ensures that the data center device with the lowest stability is most likely to fail. This constrains the repair time of the fault source that causes the current abnormal operation and maintenance parameters, thus ensuring that each data center device can operate stably and normally, and maintaining the normal operation of the data center devices.
[0155] Based on the same inventive concept, this application also proposes an intelligent platform operation and maintenance data cleaning system, including a training correlation module, used to extract the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, analyze the change amount of the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, so as to train the correlation between each fault type and each operation and maintenance parameter, wherein the correlation is determined by the correlation coefficient.
[0156] The abnormal data determination module is used to extract the operation and maintenance data corresponding to each current operation and maintenance parameter and the operation and maintenance data under normal operation status, and analyze the abnormal operation and maintenance parameters.
[0157] The fault location module determines the optimal combination of initial screening fault types based on the correlation coefficient between different fault types and various operation and maintenance parameters, and locates the optimal fault type corresponding to the current abnormal operation and maintenance parameter from the optimal combination of initial screening fault types.
[0158] The data cleaning module uses a calibration model for operation and maintenance parameters to calibrate the operation and maintenance data corresponding to the operation and maintenance parameters that are related to the optimal fault type, thereby cleaning the detected operation and maintenance data and obtaining the calibrated operation and maintenance data.
[0159] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in the claims, they should all fall within the protection scope of the present invention.
Claims
1. A method for cleaning operation and maintenance data of an intelligent platform, extracting operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, characterized in that: include: Analyze the changes in the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type to train the correlation between each fault type and each operation and maintenance parameter. The correlation is determined by the correlation coefficient. The correlation coefficient is determined based on the changes in the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type or the ratio between the changes in the operation and maintenance data corresponding to each operation and maintenance parameter and the operation and maintenance data of the operation and maintenance parameter under normal operation. Extract the operation and maintenance data corresponding to each current operation and maintenance parameter and the operation and maintenance data under normal operation status, and analyze the abnormal operation and maintenance parameters; Based on the correlation coefficient between different fault types and various operation and maintenance parameters, the optimal fault type corresponding to the current abnormal operation and maintenance parameter is located. Based on the correlation coefficient between the optimal fault type and each operation and maintenance parameter, the operation and maintenance data corresponding to the operation and maintenance parameters that are correlated with the optimal fault type are corrected in order to clean the detected operation and maintenance data and obtain the corrected operation and maintenance data of each operation and maintenance parameter. Based on the fault types corresponding to the current abnormal operation and maintenance parameters, a method for constructing initial screening combinations of fault types is provided, including: A1. Extract the fault types associated with the operation and maintenance parameters of each anomaly as the initial fault types, wherein the number of fault types associated with the operation and maintenance parameters of the anomaly is at least one. A2. Compare the initial fault types associated with any two abnormal operation and maintenance parameters, and obtain at least two fault types associated with the operation and maintenance parameters of the two abnormal parameters as the initial screening fault types. A3. Analyze the number of maintenance parameters that have a correlation coefficient greater than the set threshold between each initial screening fault type and all current abnormal maintenance parameters. Determine whether the number of maintenance parameter types with a correlation coefficient greater than the set threshold is less than the number of current abnormal maintenance data types. If it is less, proceed to the next initial screening fault type until the abnormal maintenance parameters corresponding to the selected initial screening fault type cover all current abnormal maintenance parameters or all initial screening fault types have been screened. A4. Construct at least one set of initial screening fault type combinations, wherein the types of operation and maintenance parameters corresponding to the initial screening fault types that make up the initial screening fault type combinations cover all current abnormal operation and maintenance parameter types. When the number of initial screening fault type combinations is greater than or equal to 1, the methods for determining the optimal initial screening fault type combination include: Analyze the variation coefficients corresponding to the abnormal operation and maintenance parameters, and filter out the operation and maintenance parameters corresponding to the largest variation coefficients; The number of times the fault type associated with the operation and maintenance parameter with the largest coefficient of variation appeared in all the initial screening fault type combinations was statistically analyzed. Based on the relationship between the frequency of occurrence and the frequency threshold, the fault type corresponding to at least one maintenance parameter with the largest variation coefficient is selected from the initial screening fault type combination and used as the optimal initial screening fault type combination.
2. The method for cleaning operation and maintenance data of an intelligent platform according to claim 1, characterized in that: The correlation between each fault type and each operation and maintenance parameter is determined by judging the change of each operation and maintenance parameter under each fault type, comparing it with the set screening threshold, and / or comparing it with the operation and maintenance data of the corresponding operation and maintenance parameter under normal operation conditions, to determine the correlation coefficient.
3. The method for cleaning operation and maintenance data of an intelligent platform according to claim 1, characterized in that: When the number of initial screening fault type combinations is greater than or equal to 1, obtain the reference abnormal operation and maintenance data corresponding to the operation and maintenance parameters that are related to the initial screening fault types; In the initial screening of fault type combinations, the weight corresponding to the correlation coefficient between the operation and maintenance parameters of the same anomaly and each initial screening fault type is determined based on the correlation coefficient between the operation and maintenance parameters of the anomaly and each initial screening fault type. Based on the reference abnormal operation and maintenance data corresponding to the operation and maintenance parameters of each abnormality, analyze all fault types in each initial screening fault type combination in turn, and predict the abnormal operation and maintenance data corresponding to the operation and maintenance parameters of each current abnormality. For each initial screening fault type combination, calculate the difference between the predicted abnormal operation and maintenance data corresponding to the operation and maintenance parameters of each current abnormality and the operation and maintenance data corresponding to the operation and maintenance parameters of each current abnormality. Sum the differences and select the initial screening fault type combination with the smallest summed difference as the optimal initial screening fault type combination.
4. The method for cleaning operation and maintenance data of an intelligent platform according to claim 3, characterized in that: The predicted abnormal operation and maintenance data corresponding to each abnormal operation and maintenance parameter is obtained by multiplying the weight of each initial screening fault type in the initial screening fault type combination with the reference abnormal operation and maintenance data of the operation and maintenance parameters of the abnormality under the initial screening fault type, and then summing them up.
5. The method for cleaning operation and maintenance data of an intelligent platform according to claim 1, characterized in that: Based on the correlation coefficient between the optimal fault type and each operation and maintenance parameter, the operation and maintenance data corresponding to the operation and maintenance parameters that are correlated with the optimal fault type are corrected. This includes: obtaining the operation and maintenance data affected by the fault type at any given time, and correcting the operation and maintenance data corresponding to the operation and maintenance parameters at the current time based on the number of optimal fault types and the correlation coefficient between each optimal fault type and each operation and maintenance parameter.
6. The method for cleaning operation and maintenance data of an intelligent platform according to claim 5, characterized in that: The step of correcting the operation and maintenance data corresponding to the operation and maintenance parameters that are related to the optimal fault type includes: obtaining the operation and maintenance data under the influence of the fault type at any time and the operation and maintenance data under normal operation status, and correcting the operation and maintenance data corresponding to the operation and maintenance parameters at the current time based on the duration of the fault type and the correlation coefficient between the optimal fault type and each operation and maintenance parameter.
7. A data cleaning system for operation and maintenance of an intelligent platform, characterized in that: include: The training correlation module is used to extract the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, analyze the change of the operation and maintenance data corresponding to each operation and maintenance parameter under each fault type, so as to train the correlation between each fault type and each operation and maintenance parameter. The correlation is determined by the correlation coefficient. The abnormal data determination module is used to extract the operation and maintenance data corresponding to each current operation and maintenance parameter and the operation and maintenance data under normal operation status, and analyze the abnormal operation and maintenance parameters. The fault location module determines the optimal combination of initial screening fault types based on the correlation coefficient between different fault types and various operation and maintenance parameters, and locates the optimal fault type corresponding to the current abnormal operation and maintenance parameter from the optimal combination of initial screening fault types. The data cleaning module uses a calibration model for operation and maintenance parameters to calibrate the operation and maintenance data corresponding to the operation and maintenance parameters that are related to the optimal fault type, so as to clean the detected operation and maintenance data and obtain the calibrated operation and maintenance data. Based on the fault types corresponding to the current abnormal operation and maintenance parameters, a method for constructing initial screening combinations of fault types is provided, including: A1. Extract the fault types associated with the operation and maintenance parameters of each anomaly as the initial fault types, wherein the number of fault types associated with the operation and maintenance parameters of the anomaly is at least one. A2. Compare the initial fault types associated with any two abnormal operation and maintenance parameters, and obtain at least two fault types associated with the operation and maintenance parameters of the two abnormal parameters as the initial screening fault types. A3. Analyze the number of maintenance parameters that have a correlation coefficient greater than the set threshold between each initial screening fault type and all current abnormal maintenance parameters. Determine whether the number of maintenance parameter types with a correlation coefficient greater than the set threshold is less than the number of current abnormal maintenance data types. If it is less, proceed to the next initial screening fault type until the abnormal maintenance parameters corresponding to the selected initial screening fault type cover all current abnormal maintenance parameters or all initial screening fault types have been screened. A4. Construct at least one set of initial screening fault type combinations, wherein the types of operation and maintenance parameters corresponding to the initial screening fault types that make up the initial screening fault type combinations cover all current abnormal operation and maintenance parameter types. When the number of initial screening fault type combinations is greater than or equal to 1, the methods for determining the optimal initial screening fault type combination include: Analyze the variation coefficients corresponding to the abnormal operation and maintenance parameters, and filter out the operation and maintenance parameters corresponding to the largest variation coefficients; The number of times the fault type associated with the operation and maintenance parameter with the largest coefficient of variation appeared in all the initial screening fault type combinations was statistically analyzed. Based on the relationship between the frequency of occurrence and the frequency threshold, the fault type corresponding to at least one maintenance parameter with the largest variation coefficient is selected from the initial screening fault type combination and used as the optimal initial screening fault type combination.
8. The intelligent platform operation and maintenance data cleaning system according to claim 7, characterized in that: When the number of initial screening fault type combinations is greater than or equal to 1, the method for the fault location module to determine the optimal initial screening fault type combination includes: analyzing the variation coefficients corresponding to the abnormal operation and maintenance parameters, and screening the operation and maintenance parameters corresponding to the largest variation coefficient. The number of times the fault type associated with the operation and maintenance parameter with the largest coefficient of variation appeared in all the initial screening fault type combinations was statistically analyzed. Based on the relationship between the frequency of occurrence and the frequency threshold, the fault type corresponding to at least one maintenance parameter with the largest variation coefficient is selected from the initial screening fault type combination and used as the optimal initial screening fault type combination.
Citation Information
Patent Citations
Equipment operation and maintenance management system
CN117575559A
Heating ventilation air conditioner fault disposal recommendation method based on operation data
CN118816336A
Fault analysis method and system for computer room equipment
CN120067756A