Intelligent diagnosis method and system for data center infrastructure equipment state
By receiving intelligent diagnostic commands, collecting data, and building models, and combining dual thresholds and multi-device joint diagnosis, the problem of high false negative rate in the status diagnosis of data center infrastructure equipment has been solved, and high-precision identification of implicitly coupled faults has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ATHUB CO LTD
- Filing Date
- 2025-08-29
- Publication Date
- 2026-07-07
AI Technical Summary
Traditional methods have a high false negative rate in the fault diagnosis of data center infrastructure equipment and cannot effectively identify implicitly coupled faults, resulting in a high false positive and false negative rate.
By receiving intelligent diagnostic commands, collecting and encrypting data, constructing a fault identification model group, and utilizing dual-threshold preliminary diagnosis and multi-device joint diagnosis, the false negative rate is reduced and the diagnostic accuracy of implicitly coupled faults is improved.
Significantly reduce the false negative rate of data center infrastructure equipment status diagnosis, improve the accuracy of diagnosis of implicitly coupled faults, and ensure the accuracy and efficiency of fault identification.
Smart Images

Figure CN121070667B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment fault diagnosis technology, and in particular to an intelligent diagnostic method and system for the status of data center infrastructure equipment. Background Technology
[0002] With the explosive growth of cloud computing, big data and artificial intelligence businesses, data centers have become the core artery of the digital economy. Once their infrastructure, such as servers, networks and storage, fails, it will not only lead to large-scale service interruptions and economic losses, but may also trigger cascading risks. Therefore, real-time and accurate diagnosis of the health status of these infrastructures has become the key to ensuring the continuous, safe and efficient operation of data centers.
[0003] Traditional methods rely on manual inspections or simple threshold alarms to diagnose infrastructure faults. While these methods can detect some obvious faults, they are not good at identifying some hidden faults, resulting in a high false alarm and false negative rate. Summary of the Invention
[0004] This invention provides a method and system for intelligent diagnosis of the status of data center infrastructure equipment. Its main purpose is to reduce the false negative rate of data center infrastructure equipment status diagnosis and improve the accuracy of diagnosis of implicit coupling faults in infrastructure.
[0005] To achieve the above objectives, the present invention provides an intelligent diagnostic method for the status of data center infrastructure equipment, comprising:
[0006] Receive intelligent diagnostic instructions, determine the basic device set based on the intelligent diagnostic instructions, wherein the basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices or cooling devices;
[0007] Data is collected based on the basic equipment set to obtain the encrypted equipment dataset and the original data tag set. Using the original data tag set and the basic equipment set, the encrypted equipment dataset is transmitted to the pre-built diagnostic center to obtain the qualified equipment dataset.
[0008] A model is constructed on the basic equipment set to obtain a fault identification model group. Qualified equipment data is extracted sequentially from the qualified equipment dataset, and the current identification model is identified in the fault identification model group based on the qualified equipment data.
[0009] The current identification model is used to identify faults in qualified equipment data to obtain equipment fault values and fault type probability groups. The maximum and minimum fault values are set based on the current identification model.
[0010] Based on the maximum and minimum fault values, a preliminary diagnosis of the equipment fault values is performed to obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault.
[0011] If the preliminary diagnosis result is that a fault has occurred, the faulty equipment corresponding to the equipment fault value is identified, and the faulty equipment, fault type probability group and qualified equipment data are merged to obtain equipment fault data;
[0012] If the initial diagnosis result is a suspected fault, then a multi-device joint diagnosis is performed on the equipment fault value to obtain a joint fault value. If the joint fault value is greater than the maximum fault value, then the equipment fault data is confirmed.
[0013] By aggregating equipment failure data to obtain an equipment failure dataset, and displaying the dataset through a pre-built visualization window, intelligent diagnosis of the status of data center infrastructure equipment is achieved.
[0014] Optionally, the step of collecting data based on the basic device set to obtain the encrypted device dataset and the original data tag set includes: sequentially extracting basic devices from the basic device set, monitoring the data of the basic devices to obtain the original device data; performing hash transformation on the original device data to obtain encrypted device data, using a preset verification formula to tag the encrypted device data to obtain the original data tag; and summarizing the encrypted device data and the original data tag to obtain the encrypted device dataset and the original data tag set.
[0015] Optionally, the step of using the original data tag set and basic device set to transmit the encrypted device dataset to a pre-built diagnostic center to obtain a qualified device dataset includes:
[0016] The transmission order is set for the basic device set to obtain the transmission device queue;
[0017] Transmission and reception signals are generated based on the transmission device queue, and the transmission and reception signals are broadcast to the basic device set to obtain the transmission device set;
[0018] Based on the transmission device set, the encrypted device dataset is transmitted to the diagnostic center to obtain the target device dataset. The target device dataset includes multiple target device data, and the target device data in the target device dataset corresponds one-to-one with the transmission devices in the device set to be transmitted.
[0019] Extract target device data sequentially from the target device dataset, verify the target device data based on the original data label set, and obtain the verification result, which is either verification successful or verification failed.
[0020] If the verification result is a verification failure, the transmission device corresponding to the target device data will be recorded as an unqualified device.
[0021] If the verification result is successful, the target device data will be recorded as qualified device data.
[0022] Summarize the non-conforming devices to obtain a set of non-conforming devices, and identify the non-conforming device dataset corresponding to the set of non-conforming devices in the encrypted device dataset.
[0023] The non-conforming device set and the non-conforming device dataset are respectively denoted as the transmission device set and the encrypted device dataset, and the step of transmitting the encrypted device dataset to the diagnostic center based on the transmission device set is returned until the non-conforming device set is empty or the number of returns equals the preset return threshold.
[0024] The qualified equipment data is aggregated to obtain the qualified equipment dataset.
[0025] Optionally, the verification of the target device data based on the original data tag set to obtain the verification result includes:
[0026] The target device data is verified using a verification formula to obtain verification data tags;
[0027] Identify the target data tag corresponding to the target device data in the original data tag set, and determine whether the verification data tag is equal to the target data tag;
[0028] If the verification data tag is not equal to the target data tag, the verification result will be recorded as verification failure;
[0029] If the verification data tag is equal to the target data tag, the verification result is recorded as successful.
[0030] Optionally, setting the transmission order of the basic device set to obtain a transmission device queue includes:
[0031] Query the historical transmission delay set of the basic equipment set, and calculate the mean transmission delay and the variance of transmission delay based on the historical transmission delay set;
[0032] Extract the devices to be connected sequentially from the basic equipment set;
[0033] Obtain the historical connection probability of the device to be connected, and calculate and predict the transmission duration based on the historical connection probability, the mean transmission delay, and the variance of the transmission delay;
[0034] The predicted transmission durations are aggregated to obtain a set of predicted transmission durations, where the predicted transmission durations in the set of predicted transmission durations correspond one-to-one with the basic devices in the set of basic devices.
[0035] The transmission device queue is obtained by sorting the basic device set based on the predicted transmission duration set.
[0036] Optionally, the step of building a model for the basic equipment set to obtain a fault identification model group includes: obtaining equipment type groups for the basic equipment set; classifying the basic equipment set according to the equipment type groups to obtain multiple similar equipment sets, wherein each similar equipment set in the multiple similar equipment sets corresponds one-to-one with the equipment type in the equipment type group; sequentially extracting similar equipment sets from the multiple similar equipment sets, and obtaining a historical equipment dataset based on the similar equipment sets; labeling the historical equipment dataset using a preset fault label group to obtain a training equipment dataset, wherein the fault label group includes: no fault, type I fault, and type II fault; training a pre-built neural network using the training equipment dataset to obtain an equipment fault identification model; and summarizing the equipment fault identification models to obtain the equipment fault identification model group.
[0037] Optionally, setting the maximum and minimum fault values based on the current identification model includes: obtaining current environmental parameters based on qualified equipment data; identifying a family of training data in the training equipment dataset based on the current environmental parameters, and sequentially extracting family of training data from the family of training data; validating the current identification model using the family of training data to obtain the model accuracy value; summarizing the model accuracy values to obtain a model accuracy value set, calculating the mean of the model accuracy value set to obtain the current accuracy value; and calculating the maximum and minimum fault values based on the current accuracy value.
[0038] Optionally, the step of performing multi-device joint diagnosis on equipment fault values to obtain joint fault values includes: identifying suspected devices based on equipment fault values; extracting a set of devices to be associated based on suspected devices; obtaining the number of connection nodes between each device to be associated and the suspected devices in the set of devices to be associated, thus obtaining a set of connection nodes; identifying a set of associated nodes in the set of connection nodes according to a preset number of associated connections; determining a set of associated devices in the set of devices to be associated based on the set of associated nodes, and obtaining the associated fault value of each associated device in the set of associated devices, thus obtaining a set of associated fault values; calculating a joint diagnostic factor based on the set of associated fault values and the set of associated nodes, and using the joint diagnostic factor to compensate for suspected devices to obtain a joint fault value, wherein the joint fault value is expressed as:
[0039]
[0040] in, Indicates the combined fault value. This indicates the equipment fault value corresponding to the suspected device. This represents a predefined symbolic function. Represents the natural constant. Indicates combined diagnostic factors, This indicates taking the absolute value.
[0041] Optionally, the step of calculating the joint diagnostic factor based on the associated fault value set and the associated node number set includes:
[0042] The combined diagnostic factors are calculated using the following formula:
[0043]
[0044] in, This indicates the number of associated fault values in the associated fault value set or the number of associated nodes in the associated node set. Represents the number of nodes in the set of associated nodes. Number of associated nodes, This represents the sum of the number of all associated nodes in the associated node set. Represents the first fault value in the associated fault value set. One associated fault value, Represents the first fault value in the preset set of associated fault values. The maximum fault value among the associated fault values.
[0045] To achieve the above objectives, the present invention also provides an intelligent diagnostic system for the status of data center infrastructure equipment, comprising:
[0046] The diagnostic instruction receiving module is used to receive intelligent diagnostic instructions, determine the basic device set based on the intelligent diagnostic instructions, wherein the basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices or cooling devices. Data is collected based on the basic device set to obtain an encrypted device dataset and an original data tag set. Using the original data tag set and the basic device set, the encrypted device dataset is transmitted to the pre-built diagnostic center to obtain a qualified device dataset.
[0047] The identification model building module is used to build a model for the basic equipment set to obtain a fault identification model group. It extracts qualified equipment data from the qualified equipment dataset in sequence, identifies the current identification model in the fault identification model group based on the qualified equipment data, uses the current identification model to identify faults in the qualified equipment data, and obtains equipment fault values and fault type probability groups. It also sets the maximum and minimum fault values based on the current identification model.
[0048] The equipment fault diagnosis module is used to perform preliminary diagnosis of equipment fault values based on the maximum and minimum fault values, and obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault.
[0049] The fault data display module is used to identify the faulty equipment corresponding to the equipment fault value, merge the faulty equipment, fault type probability group and qualified equipment data to obtain equipment fault data, perform multi-device joint diagnosis on the equipment fault value to obtain joint fault value, if the joint fault value is greater than the maximum fault value, the equipment fault data is confirmed, the equipment fault data is summarized to obtain equipment fault dataset, and the equipment fault dataset is displayed based on a pre-built visualization window.
[0050] To address the aforementioned problems, the present invention also provides an electronic device, comprising: a memory storing at least one instruction; and a processor executing the instructions stored in the memory to implement the aforementioned intelligent diagnostic method for the status of data center infrastructure equipment.
[0051] To address the aforementioned issues, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the aforementioned intelligent diagnostic method for the status of data center infrastructure equipment.
[0052] To address the problems described in the background section, this invention first collects data from a basic equipment set to obtain an encrypted equipment dataset and an original data tag set. Using the original data tag set and the basic equipment set, the encrypted equipment dataset is transmitted to the diagnostic center to obtain a qualified equipment dataset. This step involves hashing and encrypting the original equipment data at the acquisition end and generating CRC tags, then transmitting it in a time-series broadcast manner using dynamic queuing. This ensures data integrity and tamper-proofing, while also eliminating transmission losses through a retransmission mechanism, ensuring that the qualified dataset received by the diagnostic center is high-fidelity and complete. Next, a model is built from the basic equipment set to obtain a fault identification model group. Qualified equipment data is sequentially extracted from the qualified equipment dataset, and the current identification model is identified within the fault identification model group based on this data. This step trains a dedicated fault identification model for each equipment type. By forming model groups, each type of basic equipment can obtain the most targeted fault discrimination capability, fundamentally improving the accuracy of subsequent fault identification. Furthermore, based on the maximum and minimum fault values, a preliminary diagnosis of equipment fault values is performed to obtain preliminary diagnostic results. This step uses dual thresholds to classify equipment status into faulty, suspected faulty, and non-faulty states in one go. This avoids the high error rate caused by a single threshold and reserves space for further joint diagnosis of suspected devices, thus balancing fault diagnosis efficiency and accuracy. If the preliminary diagnosis result is a suspected fault, a multi-device joint diagnosis is performed on the equipment fault value to obtain a joint fault value. This step introduces multi-device joint diagnosis based on the number of connected nodes for suspected devices, dynamically amplifying or diluting the fault probability through risk propagation, significantly improving the detection rate of latent and coupled faults and reducing false negatives. Therefore, this invention can reduce the false negative rate of data center infrastructure equipment status diagnosis and improve the accuracy of diagnosing latent coupled faults in infrastructure. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating an intelligent diagnostic method for the status of data center infrastructure equipment provided in an embodiment of the present invention.
[0054] Figure 2 A functional block diagram of a data center infrastructure equipment status intelligent diagnostic system provided in an embodiment of the present invention;
[0055] Figure 3 This is a schematic diagram of the structure of an electronic device that implements the intelligent diagnostic method for the status of data center infrastructure equipment, according to an embodiment of the present invention.
[0056] Explanation of reference numerals in the attached figures:
[0057] 10. Electronic device; 11. Processor; 12. Memory; 13. Bus.
[0058] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0059] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0060] This application provides a method for intelligent diagnostics of the status of data center infrastructure equipment. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for intelligent diagnostics of the status of data center infrastructure equipment can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0061] Reference Figure 1 The diagram shown is a flowchart illustrating an intelligent diagnostic method for the status of data center infrastructure equipment according to an embodiment of the present invention. In this embodiment, the intelligent diagnostic method for the status of data center infrastructure equipment includes:
[0062] S1. Receive intelligent diagnostic instructions and determine the basic device set based on the intelligent diagnostic instructions. The basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices, or cooling devices.
[0063] It is clear that the intelligent diagnostic command refers to a human-initiated command to perform status diagnostics on the infrastructure equipment of a data center. The infrastructure equipment set refers to multiple infrastructure equipment in a data center specified in the intelligent diagnostic command, including, for example, server equipment, network equipment, storage equipment, power supply equipment, cooling equipment, etc.
[0064] Understandably, the server equipment refers to hardware devices in a data center that provide computing services, such as rack servers and blade servers. The network equipment refers to hardware that enables communication and data transmission between devices, such as switches and routers. The storage equipment refers to hardware responsible for persistent data storage, such as network-attached storage (SAN) and network-attached storage (NAS). The power supply equipment refers to equipment that provides stable power to the infrastructure. The cooling equipment refers to hardware that regulates the ambient temperature of the data center, such as precision air conditioners and liquid cooling systems.
[0065] It should be noted that in the aforementioned set of basic equipment, each basic equipment is pre-installed with data acquisition sensors. These sensors are used to collect operational data of the basic equipment during operation, such as current, voltage, fan speed, and network traffic. Since different basic equipment have different functions and operating environments, the data acquisition sensors installed on different types of basic equipment are also different. Consequently, the corresponding encryption device datasets for different basic equipment will also be different. For example, if a basic equipment is a server device, a CPU temperature sensor and a memory utilization sensor need to be installed in it to collect data related to the computing load. If a basic equipment is a power supply device, a current sensor and a voltage fluctuation sensor need to be installed in it to monitor the stability of the power supply.
[0066] S2. Collect data based on the basic device set to obtain the encrypted device dataset and the original data tag set. Use the original data tag set and the basic device set to transmit the encrypted device dataset to the pre-built diagnostic center to obtain the qualified device dataset.
[0067] As is clear, the encrypted device dataset refers to a collection of encrypted device data, and each encrypted device data corresponds to a basic device in the basic device set. The encrypted device data refers to the operational data of the basic devices obtained through collection. The raw data tag set refers to a collection of raw data tags, and each raw data tag corresponds to a basic device in the basic device set. The raw data tag refers to a tag value (such as a CRC checksum) generated after verifying the encrypted device data. This raw data tag is used for subsequent verification of the encrypted device data, so that encrypted device data lost during transmission can be identified in a timely manner. The diagnostic center refers to a computer processing system that centrally processes and analyzes device data, and relevant personnel can use this diagnostic center to understand the operating status of each basic device in the data center.
[0068] Furthermore, the qualified device dataset refers to a collection of multiple qualified device data, and the qualified device data refers to the encrypted device data transmitted to the diagnostic center. The qualified device data in the qualified device dataset does not correspond one-to-one with the encrypted device data in the encrypted device dataset, because the encrypted device dataset will be lost during transmission. The encrypted device data that has been lost will not be included in the qualified device dataset.
[0069] Specifically, the step of collecting data based on the basic device set to obtain the encrypted device dataset and the original data tag set includes:
[0070] The basic equipment is extracted sequentially from the basic equipment set, and data monitoring is performed on the basic equipment to obtain the raw equipment data;
[0071] The original device data is hashed to obtain encrypted device data. The encrypted device data is then marked using a preset verification formula to obtain the original data mark.
[0072] The encrypted device data and the original data tags are summarized separately to obtain the encrypted device dataset and the original data tag set.
[0073] It should be explained that the raw device data refers to the operational data of the basic equipment collected by the acquisition sensors within the basic equipment over a period of time. The encrypted device data refers to the raw device data after hash transformation. Hash transformation refers to transforming the raw device data using a hash function, where the hash function can be SHA-256. The purpose of hash transformation here is to ensure data integrity and prevent tampering during transmission. The verification formula refers to the CRC verification formula. The step of using a preset verification formula to mark the encrypted device data means: substituting the encrypted device data into the verification formula, and the output of the verification formula is the raw data mark.
[0074] Specifically, the process of transmitting the encrypted device dataset to a pre-built diagnostic center using the original data tag set and basic device set to obtain a qualified device dataset includes:
[0075] The transmission order is set for the basic device set to obtain the transmission device queue;
[0076] Transmission and reception signals are generated based on the transmission device queue, and the transmission and reception signals are broadcast to the basic device set to obtain the transmission device set;
[0077] Based on the transmission device set, the encrypted device dataset is transmitted to the diagnostic center to obtain the target device dataset. The target device dataset includes multiple target device data, and the target device data in the target device dataset corresponds one-to-one with the transmission devices in the device set to be transmitted.
[0078] Extract target device data sequentially from the target device dataset, verify the target device data based on the original data label set, and obtain the verification result, which is either verification successful or verification failed.
[0079] If the verification result is a verification failure, the transmission device corresponding to the target device data will be recorded as an unqualified device.
[0080] If the verification result is successful, the target device data will be recorded as qualified device data.
[0081] Summarize the non-conforming devices to obtain a set of non-conforming devices, and identify the non-conforming device dataset corresponding to the set of non-conforming devices in the encrypted device dataset.
[0082] The non-conforming device set and the non-conforming device dataset are respectively denoted as the transmission device set and the encrypted device dataset, and the step of transmitting the encrypted device dataset to the diagnostic center based on the transmission device set is returned until the non-conforming device set is empty or the number of returns equals the preset return threshold.
[0083] The qualified equipment data is aggregated to obtain the qualified equipment dataset.
[0084] It is clear that the transmission device queue refers to the set of basic devices after being ordered. The reason for setting the transmission order of the basic device set is to avoid network congestion caused by multiple devices transmitting simultaneously and to optimize data transmission efficiency. The transmission and reception signal refers to the instruction signal containing the transmission timing of each device. The transmission signal is generated as follows: multiple transmission moments are generated according to the arrangement order of each basic device in the transmission device queue, where each transmission moment corresponds to one basic device, and the multiple transmission moments satisfy the following: the time interval is fixed and they do not overlap. Then, the multiple transmission moments are allocated to each basic device according to the arrangement order in the transmission queue, and then encapsulated into a broadcast signal to obtain the transmission and reception signal.
[0085] It should be explained that the transmission device set refers to the set of basic devices that receive transmission and reception signals. The target device dataset refers to the encrypted device dataset transmitted to the diagnostic center. When a basic device in the basic device set receives a transmission and reception signal, it will transmit the encrypted device data corresponding to that basic device to the diagnostic center according to the transmission time allocated to that basic device by the transmission and reception signal, thereby obtaining the target device data corresponding to that basic device. The non-compliant device dataset refers to a collection of multiple non-compliant device data, and the non-compliant device data refers to the target device data corresponding to the non-compliant devices.
[0086] Furthermore, the number of returns refers to the number of times the step of returning the encrypted device dataset to the diagnostic center based on the transmission device set is executed. The return threshold is a constant set by the user. When the number of returns is equal to the return threshold, it indicates that the device data transmission still fails after multiple retransmissions. In this case, in order to avoid infinitely consuming resources, there is no need to return again.
[0087] In detail, the verification of the target device data based on the original data tag set to obtain the verification result includes: verifying the target device data using a verification formula to obtain a verification data tag; identifying the target data tag corresponding to the target device data in the original data tag set, and determining whether the verification data tag is equal to the target data tag; if the verification data tag is not equal to the target data tag, the verification result is recorded as verification failure; if the verification data tag is equal to the target data tag, the verification result is recorded as verification success.
[0088] It is clear that the verification data marker refers to the output value of the verification formula after substituting the target device data. The target data marker refers to the original data marker corresponding to the target device data. If the verification data marker is not equal to the target data marker, it indicates that the target device data has been lost during transmission, meaning that the verification result is a failure.
[0089] In detail, the step of setting the transmission order of the basic device set to obtain the transmission device queue includes: querying the historical transmission delay set of the basic device set, calculating the mean transmission delay and the variance of transmission delay based on the historical transmission delay set; sequentially extracting devices to be connected from the basic device set; obtaining the historical connection probability of the devices to be connected, calculating the predicted transmission duration based on the historical connection probability, the mean transmission delay, and the variance of transmission delay; summarizing the predicted transmission durations to obtain a predicted transmission duration set, wherein the predicted transmission durations in the predicted transmission duration set correspond one-to-one with the basic devices in the basic device set; and sorting the basic device set based on the predicted transmission duration set to obtain the transmission device queue.
[0090] It should be explained that the historical transmission delay set refers to the collection of transmission delays of each basic device in the basic device set over a past period. This historical transmission delay set is obtained by sequentially extracting basic devices from the basic device set, acquiring multiple historical transmission delays of each basic device over a past period, and summing up the multiple historical transmission delays corresponding to each basic device to obtain the historical transmission delay set. The transmission delay mean and transmission delay variance refer to the average value and variance of the historical transmission delay set, respectively.
[0091] Furthermore, the "device to be connected" refers to the basic equipment extracted from the basic equipment collection. The historical connection probability refers to the probability that the device to be connected successfully establishes a connection with the diagnostic center within a certain period. This historical connection probability can be obtained through historical data statistics. Optionally, the historical connection probability is obtained by recording the total number of successful connections established between the device to be connected and the diagnostic center within a past period, and the number of attempts made by the device to establish a connection. Then, the number of attempts to establish a connection is divided by the total number of successful connections to obtain the historical connection probability. For example, if a device to be connected attempted to connect 10 times in the past 10 seconds, and 5 of those connections were successful, then the historical connection probability of the device to be connected is 5 divided by 10, which is 0.5. The above-mentioned sorting of the basic equipment set based on the predicted transmission duration set to obtain the transmission equipment queue refers to sorting all basic equipment in the basic equipment set according to their corresponding predicted transmission duration from smallest to largest, thereby obtaining an ordered transmission equipment queue. The purpose of this operation is to allow the equipment with the shortest predicted transmission duration and the most stable network connection to transmit first, so as to shorten the overall data aggregation time, reduce the probability of network congestion, and improve the overall efficiency and reliability of encrypted equipment data successfully delivered to the diagnostic center.
[0092] Importantly, the predicted transmission time mentioned above is expressed as:
[0093]
[0094] in, Indicates the predicted transmission duration. Represents the natural constant. Represents the probability of historical connections. Represents the normal distribution function. Indicates the average transmission delay. This represents the variance of transmission delay.
[0095] S3. Construct a model for the basic equipment set to obtain a fault identification model group. Extract qualified equipment data sequentially from the qualified equipment dataset and identify the current identification model in the fault identification model group based on the qualified equipment data.
[0096] It is understood that the fault identification model group comprises multiple fault identification models, wherein each fault identification model refers to a model for identifying faults in a specific type of basic equipment, and the specific type refers to the subsequent equipment type. The input of the fault identification model is qualified equipment data, and the output is a fault probability value and a probability value for the occurrence of each fault type (i.e., a vector, where each element of the vector corresponds to the probability value for the occurrence of a fault type), wherein the fault type probability refers to the subsequent no fault, Class I fault, and Class II fault.
[0097] In detail, the step of building a model for the basic equipment set to obtain a fault identification model group includes: obtaining equipment type groups for the basic equipment set; classifying the basic equipment set according to the equipment type groups to obtain multiple similar equipment sets, wherein each similar equipment set in the multiple similar equipment sets corresponds one-to-one with the equipment type in the equipment type group; sequentially extracting similar equipment sets from the multiple similar equipment sets, and obtaining a historical equipment dataset based on the similar equipment sets; labeling the historical equipment dataset using a preset fault label group to obtain a training equipment dataset, wherein the fault label group includes: no fault, Class I fault, and Class II fault; training a pre-built neural network using the training equipment dataset to obtain an equipment fault identification model; and summarizing the equipment fault identification models to obtain the equipment fault identification model group.
[0098] It should be explained that the "device type group" refers to a combination of multiple device types. Here, "device type" refers to one type within a set of basic devices. For example, if a set of basic devices includes multiple storage devices and multiple power supply devices, then the set can be divided into two device types: storage and power supply. The "same-type device set" refers to a collection of multiple basic devices with the same device type. For example, if a set of basic devices includes multiple storage devices, namely storage device A, storage device B, and storage device C, and these storage devices have the same data acquisition sensors, meaning they collect the same type of raw device data, then storage device A, storage device B, and storage device C can be placed in a same-type device set. The "historical device dataset" refers to the collection of device data recorded by each device in the same-type device set in previous periods. This historical device dataset is obtained by sequentially extracting the same-type devices from the set and acquiring multiple historical device data points acquired by each device in previous periods. This historical device data includes normal device data and faulty device data. Normal device data refers to data from devices that have not experienced faults, and faulty device data refers to data from devices that have experienced faults. Whether a fault has occurred is determined by relevant personnel. The training device dataset refers to the labeled historical device dataset, and the training device data in the training device dataset corresponds one-to-one with the historical device data in the historical device dataset.
[0099] Furthermore, the fault label refers to the type of fault that occurs in similar devices, which is set manually. The fault label includes: no fault, Class I fault, and Class II fault. Here, no fault means that no fault has occurred in similar devices. Class I and Class II faults refer to the specific categories of faults that have occurred in similar devices. For example, if a similar device is a power supply device, the power supply device may have faults such as abnormal voltage or overload current. The abnormal voltage can be recorded as a Class I fault, and the overload current as a Class II fault. If it is necessary to further refine the fault, Class III faults, Class IV faults, etc. can also be set. If the power supply device does not have a fault during the period when historical device data is acquired, the historical device data is marked with a no-fault label. If the power supply device detects voltage fluctuations exceeding a threshold in the historical data, the historical device data is marked with a Class I fault label.
[0100] It is understood that the neural network is, for example, a convolutional neural network (CNN) or a long short-term memory network (LSTM). The training refers to adjusting the network weights through the backpropagation algorithm to minimize the prediction error. The above training steps are existing technology and will not be described in detail here.
[0101] S4. Use the current identification model to identify faults in qualified equipment data, obtain equipment fault values and fault type probability groups, and set the maximum and minimum fault values based on the current identification model.
[0102] It should be explained that the "using the current identification model to identify faults in qualified equipment data" refers to: inputting the qualified equipment data into the current identification model, and the output of the current identification model is the equipment fault value and fault type probability set. The equipment fault value quantifies the probability of a fault occurring in the underlying equipment corresponding to the qualified equipment data; the higher the fault value, the higher the probability of a fault occurring. The fault type probability set refers to the combination of probabilities of different fault categories occurring in the underlying equipment corresponding to the qualified equipment data. For example, if a fault type probability set is 0.2 and 0.4, where 0.2 and 0.4 correspond to Class I and Class II faults respectively, it means that the probabilities of the qualified equipment data occurring a Class I fault and a Class II fault are 0.2 and 0.4, respectively.
[0103] Furthermore, the maximum fault value and the minimum fault value refer to dynamically set fault probability thresholds, which serve as diagnostic boundaries to distinguish the equipment status (fault occurred, suspected fault, or no fault).
[0104] In detail, the step of setting the maximum and minimum fault values based on the current identification model includes: obtaining current environmental parameters based on qualified equipment data; identifying a family of training data in the training equipment dataset based on the current environmental parameters, and sequentially extracting family of training data from the family of training data; validating the current identification model using the family of training data to obtain the model accuracy value; summarizing the model accuracy values to obtain a model accuracy value set, calculating the mean of the model accuracy value set to obtain the current accuracy value; and calculating the maximum and minimum fault values based on the current accuracy value.
[0105] It should be explained that the current environmental parameters refer to the environmental data when qualified equipment data is obtained, such as the temperature and humidity of the data center when qualified equipment data is obtained. Because these environmental parameters will have a certain impact on the operation of the basic equipment set, the equipment data of the basic equipment under different environmental parameters will be different from the equipment data under other environmental parameters, even if the basic equipment does not malfunction. The same family training dataset refers to the collection of training equipment data with the same environmental parameters as the current environmental parameters. The model accuracy value refers to the accuracy of the current recognition model in predicting the same family training data. The model accuracy value is calculated as follows: input the same family training data into the current recognition model to obtain the output fault value, and then identify the fault label corresponding to the same family training data. If the fault label is no fault, the output fault value is subtracted from the value 1 to obtain the model accuracy value. If the fault label is not no fault, the output fault value is recorded as the model accuracy rate. The current accuracy value refers to the average value of each model accuracy value in the model accuracy value set.
[0106] Furthermore, the method for calculating the maximum and minimum fault values based on the current accurate value is as follows: obtain a baseline fault threshold, which is a threshold set manually during the training of the current recognition model. When the output value of the current recognition model is greater than the baseline fault threshold, it indicates that the basic equipment corresponding to the input data has failed. Calculate the maximum fault value based on the baseline fault threshold and the current accurate value, where the maximum fault value is expressed as the baseline fault threshold divided by the current accurate value. The minimum fault value is calculated by multiplying the maximum fault value by 0.4.
[0107] S5. Based on the maximum and minimum fault values, perform a preliminary diagnosis of the equipment fault values to obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault.
[0108] Understandably, "fault occurred" means that the basic equipment corresponding to the equipment fault value has failed, "suspected fault" means that the basic equipment corresponding to the equipment fault value may have failed, and "no fault" means that the basic equipment corresponding to the equipment fault value has not failed. Therefore, if the preliminary diagnosis result is a suspected fault, further judgment is required.
[0109] It is clear that the preliminary diagnosis of equipment fault values based on the maximum and minimum fault values means that if the equipment fault value is greater than the maximum fault value, the preliminary diagnosis result is recorded as a fault occurring; if the equipment fault value is less than the maximum fault value but greater than the minimum fault value, the preliminary diagnosis result is recorded as a suspected fault; and if the equipment fault value is less than the minimum fault value, the preliminary diagnosis result is recorded as no fault.
[0110] S6. If the preliminary diagnosis result is that a fault has occurred, the faulty equipment corresponding to the equipment fault value is identified, and the faulty equipment, fault type probability group and qualified equipment data are merged to obtain equipment fault data.
[0111] It is clear that the faulty equipment refers to the basic equipment corresponding to the equipment fault value. The equipment fault data refers to an array that combines the faulty equipment, the fault type probability group, and the qualified equipment data.
[0112] S7. If the preliminary diagnosis result is a suspected fault, then perform a multi-device joint diagnosis on the equipment fault value to obtain a joint fault value. If the joint fault value is greater than the maximum fault value, then the equipment fault data is confirmed.
[0113] It is clear that the joint fault value refers to the fault value obtained after joint diagnosis by multiple devices. The larger the joint fault value, the greater the probability that the basic equipment corresponding to the joint fault value has failed. The step of confirming the equipment fault data is the same as the above-mentioned step of merging the faulty equipment, fault type probability group, and qualified equipment data to obtain the equipment fault data.
[0114] Furthermore, the aforementioned multi-device joint diagnosis means that if it is not possible to determine whether the corresponding basic equipment has malfunctioned based solely on the data of a single qualified device, then it is necessary to combine the data of other basic devices for judgment.
[0115] In detail, the step of performing multi-device joint diagnosis on equipment fault values to obtain joint fault values includes: identifying suspected devices based on equipment fault values; extracting a set of devices to be associated based on suspected devices; obtaining the number of connection nodes between each device to be associated and the suspected devices in the set of devices to be associated, thus obtaining a set of connection nodes; identifying a set of associated nodes in the set of connection nodes according to a preset number of associated connections; determining a set of associated devices in the set of devices to be associated based on the set of associated nodes, and obtaining the associated fault value of each associated device in the set of associated devices, thus obtaining a set of associated fault values; calculating a joint diagnostic factor based on the set of associated fault values and the set of associated nodes, and using the joint diagnostic factor to compensate for suspected devices to obtain a joint fault value, wherein the joint fault value is expressed as:
[0116]
[0117] in, Indicates the combined fault value. This indicates the equipment fault value corresponding to the suspected device. This represents a predefined symbolic function. Represents the natural constant. Indicates combined diagnostic factors, This indicates taking the absolute value.
[0118] It is clear that the suspected device refers to the basic device corresponding to the device fault value when the preliminary diagnosis result is a suspected fault. The set of devices to be associated refers to a collection of multiple qualified devices other than the suspected device, where qualified devices are the basic devices corresponding to qualified device data. The number of connection nodes refers to the number of basic devices that separate the suspected device from the device to be associated. For example, if a suspected device is a power supply device D and a device to be associated is a server device E, and power supply device D and server device E are not directly connected but separated by multiple basic devices, such as: power supply device D → switch → router → server device E, then the number of connection nodes between power supply device D and server device E is 2. The joint diagnostic factor refers to the numerical value that quantifies the weight of the associated device's impact on the suspected device's fault. The larger the joint diagnostic factor, the more significant the impact of the associated device's fault status on the suspected device. If there is no connection between two basic devices, the number of connection nodes is recorded as a preset maximum value (optionally, set to 1000).
[0119] Understandably, the number of associated connections refers to a manually set constant. If the number of connected nodes is less than this number, it indicates a high functional correlation between the two devices corresponding to this number of connected nodes. If one of these two devices fails, the other device may also fail. The set of associated nodes includes multiple associated node counts, and the associated node count refers to the number of connected nodes greater than the number of associated connections. The set of associated devices refers to the collection of multiple devices to be associated corresponding to the set of associated node counts. The associated fault value refers to the device fault value corresponding to the associated device. The symbolic function is: if the independent variable (i.e., parameter) If the number is negative, then the sign function is numerical. If the independent variable is positive, then the sign function is numerical. .
[0120] It needs to be explained that the joint diagnostic factor essentially measures the intensity of risk transmission from the associated equipment group to the suspected equipment. When the joint diagnostic factor is positive and larger, it indicates that the probability of failure of the surrounding equipment is high and the risk is positively superimposed, so the failure value of the suspected equipment needs to be increased (i.e., the larger the joint failure value). When the joint diagnostic factor is negative and smaller, it indicates that the surrounding equipment is in good condition and the risk is diluted, so the failure value of the suspected equipment should be reduced (i.e., the smaller the joint failure value).
[0121] In detail, the calculation of the joint diagnostic factor based on the associated fault value set and the associated node number set includes:
[0122] The combined diagnostic factors are calculated using the following formula:
[0123]
[0124] in, This indicates the number of associated fault values in the associated fault value set or the number of associated nodes in the associated node set. Represents the number of nodes in the set of associated nodes. Number of associated nodes, This represents the sum of the number of all associated nodes in the associated node set. Represents the first fault value in the associated fault value set. One associated fault value, Represents the first fault value in the preset set of associated fault values. The maximum fault value among the associated fault values.
[0125] It is clear that the first fault value in the associated fault value set The maximum fault value of an associated fault value refers to the maximum fault value of the current identification model corresponding to the i-th associated fault value.
[0126] S8. Summarize equipment fault data to obtain equipment fault dataset, and display the equipment fault dataset based on a pre-built visualization window to complete intelligent diagnosis of the status of data center infrastructure equipment.
[0127] Understandably, the visualization window refers to a graphical user interface (GUI), such as a device status monitoring dashboard. The aforementioned display of device fault datasets based on a pre-built visualization window refers to displaying device fault data in real-time on the monitoring dashboard in the form of charts, lists, or heatmaps, allowing maintenance personnel to intuitively understand the device status.
[0128] To address the problems described in the background section, this invention first collects data from a basic equipment set to obtain an encrypted equipment dataset and an original data tag set. Using the original data tag set and the basic equipment set, the encrypted equipment dataset is transmitted to the diagnostic center to obtain a qualified equipment dataset. This step involves hashing and encrypting the original equipment data at the acquisition end and generating CRC tags, then transmitting it in a time-series broadcast manner using dynamic queuing. This ensures data integrity and tamper-proofing, while also eliminating transmission losses through a retransmission mechanism, ensuring that the qualified dataset received by the diagnostic center is high-fidelity and complete. Next, a model is built from the basic equipment set to obtain a fault identification model group. Qualified equipment data is sequentially extracted from the qualified equipment dataset, and the current identification model is identified within the fault identification model group based on this data. This step trains a dedicated fault identification model for each equipment type. By forming model groups, each type of basic equipment can obtain the most targeted fault discrimination capability, fundamentally improving the accuracy of subsequent fault identification. Furthermore, based on the maximum and minimum fault values, a preliminary diagnosis of equipment fault values is performed to obtain preliminary diagnostic results. This step uses dual thresholds to classify equipment status into faulty, suspected faulty, and non-faulty states in one go. This avoids the high error rate caused by a single threshold and reserves space for further joint diagnosis of suspected devices, thus balancing fault diagnosis efficiency and accuracy. If the preliminary diagnosis result is a suspected fault, a multi-device joint diagnosis is performed on the equipment fault value to obtain a joint fault value. This step introduces multi-device joint diagnosis based on the number of connected nodes for suspected devices, dynamically amplifying or diluting the fault probability through risk propagation, significantly improving the detection rate of latent and coupled faults and reducing false negatives. Therefore, this invention can reduce the false negative rate of data center infrastructure equipment status diagnosis and improve the accuracy of diagnosing latent coupled faults in infrastructure.
[0129] like Figure 2 The diagram shown is a functional block diagram of an intelligent diagnostic system for data center infrastructure equipment status provided in an embodiment of the present invention.
[0130] The intelligent diagnostic system 100 for data center infrastructure equipment status described in this invention can be installed in an electronic device. Depending on the functions implemented, the intelligent diagnostic system 100 may include a diagnostic instruction receiving module 101, an identification model construction module 102, an equipment fault diagnosis module 103, and a fault data display module 104. The module described in this invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, and which are stored in the memory of the electronic device.
[0131] The diagnostic instruction receiving module 101 is used to receive intelligent diagnostic instructions, determine the basic device set based on the intelligent diagnostic instructions, wherein the basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices or cooling devices, collect data according to the basic device set to obtain an encrypted device dataset and an original data tag set, and use the original data tag set and the basic device set to transmit the encrypted device dataset to the pre-built diagnostic center to obtain a qualified device dataset.
[0132] The identification model construction module 102 is used to construct a model for the basic equipment set to obtain a fault identification model group, extract qualified equipment data sequentially from the qualified equipment dataset, identify the current identification model in the fault identification model group based on the qualified equipment data, use the current identification model to identify faults in the qualified equipment data to obtain equipment fault values and fault type probability groups, and set the maximum fault value and minimum fault value based on the current identification model.
[0133] The equipment fault diagnosis module 103 is used to perform preliminary diagnosis of equipment fault values based on the maximum fault value and the minimum fault value, and obtain preliminary diagnosis results, wherein the preliminary diagnosis results include: fault occurred, suspected fault, and no fault.
[0134] The fault data display module 104 is used to identify the faulty device corresponding to the fault value, merge the faulty device, fault type probability group and qualified device data to obtain the equipment fault data, perform multi-device joint diagnosis on the equipment fault value to obtain the joint fault value, if the joint fault value is greater than the maximum fault value, then the equipment fault data is identified, the equipment fault data is summarized to obtain the equipment fault dataset, and the equipment fault dataset is displayed based on a pre-built visualization window.
[0135] In detail, the modules in the intelligent diagnostic system 100 for data center infrastructure equipment status described in this embodiment of the invention employ the same methods as described above. Figure 1 The method used here is the same as the intelligent diagnostic method for the status of data center infrastructure equipment described above, and can produce the same technical effect, so it will not be elaborated here.
[0136] like Figure 3 The diagram shown is a structural schematic of an electronic device for implementing an intelligent diagnostic method for the status of data center infrastructure equipment, according to an embodiment of the present invention.
[0137] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a data center infrastructure equipment status intelligent diagnostic method program.
[0138] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 1. Furthermore, the memory 11 includes both internal storage units and external storage devices of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a data center infrastructure equipment status intelligent diagnostic method program, but also to temporarily store data that has been output or will be output.
[0139] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (such as intelligent diagnostic methods for data center infrastructure equipment status) and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.
[0140] The bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0141] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0142] For example, although not shown, the electronic device 1 may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management system, thereby enabling functions such as charging management, discharging management, and power consumption management through the power management system. The power supply may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0143] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 1 and other electronic devices.
[0144] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual user interface.
[0145] The data center infrastructure equipment status intelligent diagnostic method program stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When run in the processor 10, it can achieve the following:
[0146] Receive intelligent diagnostic instructions, determine the basic device set based on the intelligent diagnostic instructions, wherein the basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices or cooling devices;
[0147] Data is collected based on the basic equipment set to obtain the encrypted equipment dataset and the original data tag set. Using the original data tag set and the basic equipment set, the encrypted equipment dataset is transmitted to the pre-built diagnostic center to obtain the qualified equipment dataset.
[0148] A model is constructed on the basic equipment set to obtain a fault identification model group. Qualified equipment data is extracted sequentially from the qualified equipment dataset, and the current identification model is identified in the fault identification model group based on the qualified equipment data.
[0149] The current identification model is used to identify faults in qualified equipment data to obtain equipment fault values and fault type probability groups. The maximum and minimum fault values are set based on the current identification model.
[0150] Based on the maximum and minimum fault values, a preliminary diagnosis of the equipment fault values is performed to obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault.
[0151] If the preliminary diagnosis result is that a fault has occurred, the faulty equipment corresponding to the equipment fault value is identified, and the faulty equipment, fault type probability group and qualified equipment data are merged to obtain equipment fault data;
[0152] If the initial diagnosis result is a suspected fault, then a multi-device joint diagnosis is performed on the equipment fault value to obtain a joint fault value. If the joint fault value is greater than the maximum fault value, then the equipment fault data is confirmed.
[0153] By aggregating equipment failure data to obtain an equipment failure dataset, and displaying the dataset through a pre-built visualization window, intelligent diagnosis of the status of data center infrastructure equipment is achieved.
[0154] Specifically, the processor 10's implementation method for the above instructions can be found in [reference needed]. Figures 1 to 3 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0155] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or system capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0156] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following:
[0157] Receive intelligent diagnostic instructions, determine the basic device set based on the intelligent diagnostic instructions, wherein the basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices or cooling devices;
[0158] Data is collected based on the basic equipment set to obtain the encrypted equipment dataset and the original data tag set. Using the original data tag set and the basic equipment set, the encrypted equipment dataset is transmitted to the pre-built diagnostic center to obtain the qualified equipment dataset.
[0159] A model is constructed on the basic equipment set to obtain a fault identification model group. Qualified equipment data is extracted sequentially from the qualified equipment dataset, and the current identification model is identified in the fault identification model group based on the qualified equipment data.
[0160] The current identification model is used to identify faults in qualified equipment data to obtain equipment fault values and fault type probability groups. The maximum and minimum fault values are set based on the current identification model.
[0161] Based on the maximum and minimum fault values, a preliminary diagnosis of the equipment fault values is performed to obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault.
[0162] If the preliminary diagnosis result is that a fault has occurred, the faulty equipment corresponding to the equipment fault value is identified, and the faulty equipment, fault type probability group and qualified equipment data are merged to obtain equipment fault data;
[0163] If the initial diagnosis result is a suspected fault, then a multi-device joint diagnosis is performed on the equipment fault value to obtain a joint fault value. If the joint fault value is greater than the maximum fault value, then the equipment fault data is confirmed.
[0164] By aggregating equipment failure data to obtain an equipment failure dataset, and displaying the dataset through a pre-built visualization window, intelligent diagnosis of the status of data center infrastructure equipment is achieved.
[0165] In the embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and actual implementations may have other classification methods.
[0166] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0167] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0168] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for intelligent diagnosis of the status of data center infrastructure equipment, characterized in that, The method includes: Receive intelligent diagnostic instructions, determine the basic device set based on the intelligent diagnostic instructions, wherein the basic device set includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices or cooling devices; Data is collected based on the basic equipment set to obtain the encrypted equipment dataset and the original data tag set. Using the original data tag set and the basic equipment set, the encrypted equipment dataset is transmitted to the pre-built diagnostic center to obtain the qualified equipment dataset. The process of collecting data based on the basic device set to obtain the encrypted device dataset and the original data tag set includes: The basic equipment is extracted sequentially from the basic equipment set, and data monitoring is performed on the basic equipment to obtain the raw equipment data; The original device data is hashed to obtain encrypted device data. The encrypted device data is then marked using a preset verification formula to obtain the original data mark. The encrypted device data and the original data tags are summarized separately to obtain the encrypted device dataset and the original data tag set; A model is constructed on the basic equipment set to obtain a fault identification model group. Qualified equipment data is extracted sequentially from the qualified equipment dataset, and the current identification model is identified in the fault identification model group based on the qualified equipment data. The current identification model is used to identify faults in qualified equipment data to obtain equipment fault values and fault type probability groups. The maximum and minimum fault values are set based on the current identification model. Based on the maximum and minimum fault values, a preliminary diagnosis of the equipment fault values is performed to obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault. If the preliminary diagnosis result is that a fault has occurred, the faulty equipment corresponding to the equipment fault value is identified, and the faulty equipment, fault type probability group and qualified equipment data are merged to obtain equipment fault data; If the initial diagnosis result is a suspected fault, then a multi-device joint diagnosis is performed on the equipment fault value to obtain a joint fault value. If the joint fault value is greater than the maximum fault value, then the equipment fault data is confirmed. By aggregating equipment failure data to obtain an equipment failure dataset, and displaying the dataset through a pre-built visualization window, intelligent diagnosis of the status of data center infrastructure equipment is achieved.
2. The intelligent diagnostic method for data center infrastructure equipment status as described in claim 1, characterized in that, The process of transmitting the encrypted device dataset to a pre-built diagnostic center using the original data tag set and basic device set to obtain a qualified device dataset includes: The transmission order is set for the basic device set to obtain the transmission device queue; Transmission and reception signals are generated based on the transmission device queue, and the transmission and reception signals are broadcast to the basic device set to obtain the transmission device set; Based on the transmission device set, the encrypted device dataset is transmitted to the diagnostic center to obtain the target device dataset. The target device dataset includes multiple target device data, and the target device data in the target device dataset corresponds one-to-one with the transmission devices in the device set to be transmitted. Extract target device data sequentially from the target device dataset, verify the target device data based on the original data label set, and obtain the verification result, which is either verification successful or verification failed. If the verification result is a verification failure, the transmission device corresponding to the target device data will be recorded as an unqualified device. If the verification result is successful, the target device data will be recorded as qualified device data. Summarize the non-conforming devices to obtain a set of non-conforming devices, and identify the non-conforming device dataset corresponding to the set of non-conforming devices in the encrypted device dataset. The non-conforming device set and the non-conforming device dataset are respectively denoted as the transmission device set and the encrypted device dataset, and the step of transmitting the encrypted device dataset to the diagnostic center based on the transmission device set is returned until the non-conforming device set is empty or the number of returns equals the preset return threshold. The qualified equipment data is aggregated to obtain the qualified equipment dataset.
3. The intelligent diagnostic method for data center infrastructure equipment status as described in claim 2, characterized in that, The verification of the target device data based on the original data tag set, to obtain the verification result, includes: The target device data is verified using a verification formula to obtain verification data tags; Identify the target data tag corresponding to the target device data in the original data tag set, and determine whether the verification data tag is equal to the target data tag; If the verification data tag is not equal to the target data tag, the verification result will be recorded as verification failure; If the verification data tag is equal to the target data tag, the verification result is recorded as successful.
4. The intelligent diagnostic method for data center infrastructure equipment status as described in claim 3, characterized in that, The step of setting the transmission order of the basic device set to obtain the transmission device queue includes: Query the historical transmission delay set of the basic equipment set, and calculate the mean transmission delay and the variance of transmission delay based on the historical transmission delay set; Extract the devices to be connected sequentially from the basic equipment set; Obtain the historical connection probability of the device to be connected, and calculate and predict the transmission duration based on the historical connection probability, the mean transmission delay, and the variance of the transmission delay; The predicted transmission durations are aggregated to obtain a set of predicted transmission durations, where the predicted transmission durations in the set of predicted transmission durations correspond one-to-one with the basic devices in the set of basic devices. The transmission device queue is obtained by sorting the basic device set based on the predicted transmission duration set.
5. The intelligent diagnostic method for data center infrastructure equipment status as described in claim 4, characterized in that, The process of building a model of the basic equipment set to obtain a fault identification model group includes: Retrieve the device type group of the basic equipment set; The basic equipment set is classified according to the equipment type group to obtain multiple equipment sets of the same type. Among them, the equipment sets of the same type correspond one-to-one with the equipment types in the equipment type group. Extract similar device sets sequentially from multiple similar device sets, and obtain historical device datasets based on these similar device sets; The historical equipment dataset is labeled using a pre-defined fault label group to obtain the training equipment dataset. The fault label group includes: no fault, Class I fault, and Class II fault. A pre-built neural network is trained using a training equipment dataset to obtain an equipment fault identification model; The equipment fault identification models are summarized to obtain the equipment fault identification model group.
6. The intelligent diagnostic method for the status of data center infrastructure equipment as described in claim 5, characterized in that, The setting of maximum and minimum fault values based on the current identification model includes: Obtain current environmental parameters based on data from qualified equipment; Based on the current environmental parameters, identify the same training dataset in the training device dataset, and extract the same training data in the same training dataset in sequence. The current recognition model is validated using training data from the same family to obtain the model's accuracy. Summarize the accurate values of the model to obtain the accurate value set, and calculate the mean of the accurate value set to obtain the current accurate value; Calculate the maximum and minimum fault values based on the current accurate values.
7. The intelligent diagnostic method for data center infrastructure equipment status as described in claim 6, characterized in that, The process of performing multi-device joint diagnosis on equipment fault values to obtain joint fault values includes: Based on the equipment fault values, suspected equipment is identified, and a set of equipment to be associated is extracted based on the suspected equipment. Obtain the number of connection nodes between each device to be associated and the suspected device in the set of devices to be associated, and obtain the set of connection node counts; Identify the set of associated nodes in the set of connected nodes based on the preset number of associated connections; Based on the set of associated nodes, determine the set of associated devices in the set of devices to be associated, and obtain the associated fault value of each associated device in the set of associated devices to obtain the set of associated fault values; A joint diagnostic factor is calculated based on the associated fault value set and the associated node set. This joint diagnostic factor is then used to compensate for suspected equipment, yielding a joint fault value, which is expressed as follows: in, Indicates the combined fault value. This indicates the equipment fault value corresponding to the suspected device. This represents a predefined symbolic function. Represents the natural constant. Indicates combined diagnostic factors, This indicates taking the absolute value.
8. The intelligent diagnostic method for the status of data center infrastructure equipment as described in claim 7, characterized in that, The calculation of the joint diagnostic factor based on the associated fault value set and the associated node number set includes: The combined diagnostic factors are calculated using the following formula: in, This indicates the number of associated fault values in the associated fault value set or the number of associated nodes in the associated node set. Represents the number of nodes in the set of associated nodes. Number of associated nodes, This represents the sum of the number of all associated nodes in the associated node set. Represents the first fault value in the associated fault value set. One associated fault value, Represents the first fault value in the preset set of associated fault values. The maximum fault value among the associated fault values.
9. A data center infrastructure equipment status intelligent diagnostic system, characterized in that, The system includes: A diagnostic instruction receiving module is used to receive intelligent diagnostic instructions, determine a set of basic devices based on the intelligent diagnostic instructions, wherein the set of basic devices includes multiple basic devices, and the basic devices are server devices, network devices, storage devices, power supply devices, or cooling devices. Data is collected based on the set of basic devices to obtain an encrypted device dataset and an original data tag set. The process of collecting data based on the set of basic devices to obtain the encrypted device dataset and the original data tag set includes: The basic equipment is extracted sequentially from the basic equipment set, and data monitoring is performed on the basic equipment to obtain the raw equipment data; The original device data is hashed to obtain encrypted device data. The encrypted device data is then marked using a preset verification formula to obtain the original data mark. The encrypted device data and the original data tags are summarized separately to obtain the encrypted device dataset and the original data tag set; Using the original data tag set and basic device set, the encrypted device dataset is transmitted to the pre-built diagnostic center to obtain a qualified device dataset; The identification model building module is used to build a model for the basic equipment set to obtain a fault identification model group. It extracts qualified equipment data from the qualified equipment dataset in sequence, identifies the current identification model in the fault identification model group based on the qualified equipment data, uses the current identification model to identify faults in the qualified equipment data, and obtains equipment fault values and fault type probability groups. It also sets the maximum and minimum fault values based on the current identification model. The equipment fault diagnosis module is used to perform preliminary diagnosis of equipment fault values based on the maximum and minimum fault values, and obtain preliminary diagnosis results, which include: fault occurred, suspected fault, and no fault. The fault data display module is used to identify the faulty equipment corresponding to the equipment fault value, merge the faulty equipment, fault type probability group and qualified equipment data to obtain equipment fault data, perform multi-device joint diagnosis on the equipment fault value to obtain joint fault value, if the joint fault value is greater than the maximum fault value, the equipment fault data is confirmed, the equipment fault data is summarized to obtain equipment fault dataset, and the equipment fault dataset is displayed based on a pre-built visualization window.