Failure prediction method, device and equipment for PCIE (Peripheral Component Interface Express) card and readable storage medium

By monitoring the operating status of PCIE cards and using pre-trained fault prediction models to detect and output the fault prediction sequence, the problem of time-consuming PCIE cards in the existing technology is solved, and fault detection and troubleshooting is achieved in advance, and fault prediction is reduced and maintenance time and cost is reduced, and equipment life is extended.

CN120045373APending Publication Date: 2025-05-27INSPUR (SHANDONG) COMPUTER TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510167641.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology lacks effective PCIE card fault prediction methods, which leads to time-consuming failure detection and troubleshooting, affects system performance, may lead to data loss or damage, and will have a long maintenance time and high cost, which will affect the life of the equipment.

Method used

By monitoring the running status of the PCIE card, the running status data is detected using the pre-trained fault prediction model, the fault prediction sequence obtained by sorting each parameter type is output, and troubleshooting is carried out in this order.

Benefits of technology

It realizes the advance prediction of possible failures of PCIE cards, reduces maintenance time and costs, extends the service life of the equipment, avoids further deterioration of the fault, and improves production efficiency and equipment reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045373A_ABST
    Figure CN120045373A_ABST
Patent Text Reader

Abstract

The invention discloses a PCIE (Peripheral Component Interface Express) card fault prediction method, which comprises the following steps: monitoring the running state of a PCIE card to obtain running state data; detecting the operation state data by using a pre-trained fault prediction model to obtain a detection result; and when the detection result is that the PCIE card is in the abnormal mode, outputting a fault prediction sequence obtained by sorting various parameter types through a fault prediction model so as to carry out PCIE card troubleshooting according to the fault prediction sequence. By applying the fault prediction method of the PCIE card provided by the invention, the time and cost required by maintenance are reduced, and the service life of equipment is prolonged. The invention further discloses a fault prediction device and equipment for the PCIE card and a storage medium, which have corresponding technical effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies, and in particular, to a method, device, equipment, and computer-readable storage medium for predicting faults of a PCIE card. Background Art

[0002] Currently, servers are required to support the function of predicting component alarms, but there is currently no method for predicting faults of PCIE (peripheral component interconnect express, a high-speed serial computer expansion bus standard) cards, which are important components during the operation of servers.

[0003] Due to the lack of a method for predicting faults of PCIE cards, both the process of fault detection and troubleshooting of PCIE cards are time-consuming. Faults of PCIE cards can cause a significant decline in system performance, such as slow program operation, long response times, lags, and crashes. Severe PCIE card faults may lead to data loss or corruption, especially when processing important files or performing critical tasks. In addition, system crashes and data loss also pose a threat to data security. The repair time of PCIE card faults is long, the cost is high, and it affects the service life of the equipment.

[0004] In summary, how to effectively solve the problems of lack of fault prediction for PCIE cards, long repair time, high cost, and affecting the service life of equipment for PCIE cards is an urgent problem for those skilled in the art currently. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for predicting faults of a PCIE card, which reduces the time and cost required for repair and extends the service life of the equipment; another purpose of the present invention is to provide a device, equipment, and computer-readable storage medium for predicting faults of a PCIE card.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] A method for predicting faults of a PCIE card includes:

[0008] Monitoring the operating state of the PCIE card to obtain operating state data;

[0009] Using a pre-trained fault prediction model to detect the operating state data to obtain a detection result;

[0010] When the detection result indicates that the PCIE card is in an abnormal mode, outputting a fault prediction order sorted by various parameter types through the fault prediction model to perform fault troubleshooting on the PCIE card according to the fault prediction order.

[0011] In a specific embodiment of the present invention, it further includes the training process of the fault prediction model, and the training process of the fault prediction model includes:

[0012] Obtain the historical operation status data and operation modes respectively corresponding to multiple PCIE cards;

[0013] Use each historical operation status data and each operation mode to train a pre-constructed initial prediction model to obtain a mode prediction model;

[0014] Collect the historical usage data sets respectively corresponding to multiple PCIE cards; wherein, the historical usage data sets contain parameter groups corresponding to multiple usage duration values respectively;

[0015] Use the historical usage data sets to train the mode prediction model to obtain the fault prediction model.

[0016] In a specific embodiment of the present invention, using the historical usage data sets to train the mode prediction model to obtain the fault prediction model includes:

[0017] Calculate the overall mean of each parameter in the historical usage data sets;

[0018] Calculate the mean of each parameter type corresponding to each parameter type in the historical usage data sets;

[0019] According to the mean of each parameter type and the overall mean, calculate the sum of squared deviations corresponding to each parameter type;

[0020] Calculate the mean of the parameter groups respectively corresponding to each PCIE card in the historical usage data sets;

[0021] According to the mean of each parameter group and each parameter value included in each parameter group, calculate the sum of squared errors corresponding to each PCIE card respectively;

[0022] Determine the fault prediction order according to the sum of squared deviations and the sum of squared errors;

[0023] Use the fault prediction order to adjust the mode prediction model to obtain the fault prediction model.

[0024] In a specific embodiment of the present invention, determining the fault prediction order according to the sum of squared deviations and the sum of squared errors includes:

[0025] Sort each parameter type according to the sum of squared deviations to obtain a first sorting result;

[0026] Sort each parameter type according to the sum of squared errors to obtain a second sorting result;

[0027] Obtain the first influence rate of the preset sum of squared deviations on fault prediction, and obtain the second influence rate of the preset sum of squared errors on fault prediction;

[0028] Determine the fault prediction order according to the first sorting result, the second sorting result, the first influence rate, and the second influence rate.

[0029] In a specific embodiment of the present invention, after obtaining the fault prediction model, it further includes:

[0030] Re-collect the historical usage data sets corresponding to multiple PCIe cards respectively, and determine the re-collected historical usage data sets as the test data sets;

[0031] Use the test data sets to test the fault prediction model to obtain test results;

[0032] When it is determined according to the test results that there are parameter types in the fault prediction order that have no influence on fault prediction, remove the parameter types that have no influence on fault prediction from the fault prediction order.

[0033] In a specific embodiment of the present invention, after obtaining the fault prediction model, it further includes:

[0034] When adding new parameter types, based on the parameter types after adding the new parameter types, repeat the step of calculating the mean values of the parameter types corresponding to the parameter types in the historical usage data set to update the fault prediction model.

[0035] In a specific embodiment of the present invention, after outputting the fault prediction order sorted by each parameter type through the fault prediction model to perform PCIe card fault troubleshooting according to the fault prediction order, it further includes:

[0036] When performing PCIe card fault troubleshooting according to the fault prediction order and taking corresponding maintenance strategies for maintenance, continue to run the PCIe card;

[0037] When it is monitored that the PCIe card is in an abnormal mode, output a prompt message indicating that the PCIe card life has reached the upper limit.

[0038] A fault prediction device for a PCIe card, including:

[0039] An operating state data acquisition module, configured to monitor the operating state of the PCIe card to obtain operating state data;

[0040] A detection result acquisition module, configured to use a pre-trained fault prediction model to detect the operating state data to obtain detection results;

[0041] A fault prediction module, configured to, when the detection result indicates that the PCIe card is in an abnormal mode, output a fault prediction order obtained by sorting each parameter type through the fault prediction model, so as to troubleshoot PCIe card faults according to the fault prediction order.

[0042] A fault prediction device for a PCIe card, comprising:

[0043] A memory, configured to store a computer program;

[0044] A processor, configured to implement the steps of the fault prediction method for the PCIe card as described above when executing the computer program.

[0045] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the fault prediction method for the PCIe card as described above are implemented.

[0046] The fault prediction method for the PCIe card provided by the present invention monitors the operating state of the PCIe card to obtain operating state data; uses a pre-trained fault prediction model to detect the operating state data to obtain a detection result; when the detection result indicates that the PCIe card is in an abnormal mode, outputs a fault prediction order obtained by sorting each parameter type through the fault prediction model, so as to troubleshoot PCIe card faults according to the fault prediction order.

[0047] The beneficial effects of the present invention are as follows: By pre-training a fault prediction model for predicting faults of the PCIe card, it is possible to predict in advance the possible faults of the PCIe card according to the operating state of the PCIe card, so as to troubleshoot PCIe card faults according to the fault prediction order obtained by sorting each parameter type output by the fault prediction model, avoid the generation of downtime, and then greatly reduce the production interruption caused by faults in the production line or equipment, and improve production efficiency. It avoids the further deterioration of faults and reduces the time and cost required for maintenance. Predictive maintenance is more efficient than regular maintenance, and only performs maintenance when it is truly necessary, avoiding unnecessary maintenance. Fault prediction can help detect potential equipment faults and take corresponding measures to avoid the occurrence of accidents. By promptly discovering faults and performing maintenance, it can ensure the safe operation of the equipment, reduce potential risks, and extend the service life of the equipment.

[0048] Correspondingly, the present invention also provides a fault prediction device, equipment, and computer-readable storage medium corresponding to the above-mentioned fault prediction method for the PCIe card, which have the above technical effects and will not be elaborated here. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a flowchart of an implementation of the fault prediction method for the PCIE card in the embodiments of the present invention;

[0051] Figure 2 It is a block diagram of the structure of a fault prediction device for a PCIE card in the embodiments of the present invention;

[0052] Figure 3 It is a block diagram of the structure of a fault prediction device for a PCIE card in the embodiments of the present invention;

[0053] Figure 4 It is a specific structural schematic diagram of a fault prediction device for a PCIE card provided in the embodiments of the present invention. Specific implementation manners

[0054] To enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following further elaborates on the present invention in conjunction with the drawings and specific implementation manners. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0055] Refer to Figure 1 , Figure 1 It is a flowchart of an implementation of the fault prediction method for the PCIE card in the embodiments of the present invention. This method may include the following steps:

[0056] S101: Monitor the operating state of the PCIE card to obtain operating state data.

[0057] During the operation of the PCIE card, monitor the operating state of the PCIE card to obtain operating state data. For example, a real-time monitoring system can be established to continuously and real-time monitor the usage and performance of the PCIE card. This can be achieved by using monitoring tools, software, or scripts. The monitoring system should be able to regularly obtain and record the indicators and data related to the PCIE card as the operating state data.

[0058] The operating status data may include the bus width currently used by the PCIE card, the current transmission rate of the PCIE card, the link status between the PCIE card and the host, the number of errors occurring during the sending and receiving processes of the PCIE card, the status of the device driver used by the PCIE card, various functions and features supported by the PCIE card (such as power management, error detection and correction, interrupt handling, etc.), the status of the bus controller to which the PCIE card is connected, and so on.

[0059] S102: Detect the operating status data using the pre-trained fault prediction model to obtain a detection result.

[0060] The fault prediction model is pre-trained, and the fault prediction model can perform matching detection on the operating status data and abnormal patterns. After monitoring the operating status data of the PCIE card, the pre-trained fault prediction model is used to detect the operating status data to obtain a detection result.

[0061] S103: When the detection result indicates that the PCIE card is in an abnormal mode, output the fault prediction order obtained by sorting each parameter type through the fault prediction model, so as to troubleshoot the PCIE card faults in accordance with the fault prediction order.

[0062] The pre-trained fault prediction model can also output the fault prediction order obtained by sorting each parameter type when it detects that the PCIE card is in an abnormal mode. The sorting of each parameter type is the sorting result obtained by sorting the influence probabilities of each parameter type on the PCIE card faults during training. When the detection result obtained by using the fault prediction model to detect the operating status data of the PCIE card indicates that the PCIE card is in an abnormal mode, output the fault prediction order obtained by sorting each parameter type through the fault prediction model, so as to remind the operation and maintenance personnel to troubleshoot the PCIE card faults in accordance with the fault prediction order.

[0063] By sorting the influence probabilities of each parameter type on the PCIE card faults during training to obtain the fault prediction order and troubleshooting the PCIE card faults in accordance with the fault prediction order, the timely troubleshooting of the important fault influence causes is realized, the generation of downtime is avoided, and further the production interruption caused by faults in the production line or equipment is greatly reduced, improving the production efficiency. The further deterioration of the faults is avoided, and the time and cost required for maintenance are reduced. By timely discovering and maintaining the faults, the safe operation of the equipment can be guaranteed, potential risks are reduced, and the service life of the equipment is extended.

[0064] As can be seen from the above technical solutions, a fault prediction model for predicting faults of a PCIE card is obtained through pre-training, which can predict possible faults of the PCIE card in advance according to the operating state of the PCIE card. Thus, the PCIE card fault troubleshooting is carried out according to the fault prediction order sorted by various parameter types output by the fault prediction model, avoiding the generation of downtime, and further greatly reducing the production interruption caused by faults in the production line or equipment, improving production efficiency. The further deterioration of faults is avoided, and the time and cost required for maintenance are reduced. Predictive maintenance is more efficient than regular maintenance, and maintenance is only carried out when it is truly necessary, avoiding unnecessary maintenance. Fault prediction can help detect potential equipment faults and take corresponding measures to avoid the occurrence of accidents. By promptly discovering faults and performing maintenance, the safe operation of the equipment can be guaranteed, potential risks can be reduced, and the service life of the equipment can be extended.

[0065] It should be noted that based on the above embodiments, the embodiments of the present invention also provide corresponding improvement solutions. In the subsequent embodiments, the same steps or corresponding steps involved in the above embodiments can be referred to each other, and the corresponding beneficial effects can also be referred to each other, and will not be elaborated one by one in the following improvement embodiments.

[0066] In a specific embodiment of the present invention, the method may further include the training process of the fault prediction model, and the training process of the fault prediction model may include the following steps:

[0067] Step 1: Obtain the historical operating state data and operating modes respectively corresponding to multiple PCIE cards;

[0068] Step 2: Use each piece of historical operating state data and each operating mode to train a pre-constructed initial prediction model to obtain a mode prediction model;

[0069] Step 3: Collect the historical usage data sets respectively corresponding to multiple PCIE cards; wherein, the historical usage data set contains parameter groups corresponding to multiple usage duration values;

[0070] Step 4: Use the historical usage data set to train the mode prediction model to obtain a fault prediction model.

[0071] For convenience of description, the above four steps can be combined for explanation.

[0072] Pre-build an initial prediction model, obtain the historical operation status data and operation modes corresponding to multiple PCIe cards respectively, and use machine learning algorithms to train the pre-built initial prediction model with each piece of historical operation status data and each operation mode to obtain a mode prediction model. The mode prediction model can predict whether a PCIe card is in an abnormal mode based on the input operation status data. Through the analysis and mining of the historical operation status data, the mode prediction model can identify different types of fault modes. For example, it can determine that the PCIe card is in a fault mode by identifying page errors, segment errors, etc. such as the temperature of the PCIe card. Collect the historical usage data sets corresponding to multiple PCIe cards respectively. The historical usage data set contains parameter groups corresponding to multiple usage duration values. Use the historical usage data set to train the mode prediction model to obtain a fault prediction model. Thus, the fault prediction model can output the fault troubleshooting order for each parameter type when the PCIe card is in an abnormal mode. When it is determined that the PCIe card is in an abnormal mode, the system will generate a PCIe card alarm to remind the operation and maintenance personnel of possible PCIe card problems. The alarm information can be sent to the operation and maintenance personnel through methods such as mobile applications and emails. The operation and maintenance personnel can also view the log content through the man-machine interaction interface and take corresponding solutions according to the description of the log so that the operation and maintenance personnel can take repair measures in time. The alarm log page of the management system can alarm according to the fault level and the fault location of the PCIe card and give handling opinions.

[0073] Through training, a fault prediction model capable of outputting the fault prediction order sorted by each parameter type is obtained. When the detection result obtained by using the fault prediction model to detect the operation status data of the PCIe card is that the PCIe card is in an abnormal mode, the fault prediction order sorted by each parameter type is output through the fault prediction model. Thus, it can remind the operation and maintenance personnel to conduct PCIe card fault troubleshooting according to the fault prediction order, can predict in advance before the fault occurs, and take corresponding measures to avoid or reduce the impact of the fault on the normal operation of the system, improving the reliability and stability of the computing device.

[0074] The historical usage data set can include the used duration, temperature, voltage, power consumption, and usage frequency of the PCIe card.

[0075] It should be noted that the PCIe cards to which the training data sets collected during the training of the initial prediction model belong and the PCIe cards to which the training data sets collected during the training of the mode prediction model belong can be the same PCIe cards or different PCIe cards. The embodiments of the present invention do not limit this.

[0076] In a specific embodiment of the present invention, training the pre-built initial prediction model with each piece of historical operation status data and each operation mode to obtain a mode prediction model may include the following steps:

[0077] Step 1: Calculate the overall mean of each parameter in the historical usage dataset;

[0078] Step 2: Calculate the mean of each parameter type corresponding to each parameter type in the historical usage dataset;

[0079] Step 3: Calculate the sum of squared deviations corresponding to each parameter type according to the mean of each parameter type and the overall mean;

[0080] Step 4: Calculate the mean of each parameter group corresponding to each PCIe card in the historical usage dataset;

[0081] Step 5: Calculate the sum of squared errors corresponding to each PCIe card according to the mean of each parameter group and each parameter value included in each parameter group;

[0082] Step 6: Determine the fault prediction order according to the sum of squared deviations and the sum of squared errors;

[0083] Step 7: Adjust the pattern prediction model using the fault prediction order to obtain the fault prediction model.

[0084] For ease of description, the above seven steps can be combined for illustration.

[0085] Calculate the overall mean of each parameter in the historical usage dataset, calculate the mean of each parameter type corresponding to each parameter type in the historical usage dataset, calculate the sum of squared deviations corresponding to each parameter type according to the mean of each parameter type and the overall mean, calculate the mean of each parameter group corresponding to each PCIe card in the historical usage dataset, calculate the sum of squared errors corresponding to each PCIe card according to the mean of each parameter group and each parameter value included in each parameter group, determine the fault prediction order according to the sum of squared deviations and the sum of squared errors, adjust the pattern prediction model using the fault prediction order to obtain the fault prediction model.

[0086] Refer to Table 1. Table 1 is a parameter statistical table of a PCIe card in an embodiment of the present invention.

[0087]

[0088] As shown in Table 1, the overall mean of each parameter is the value obtained by calculating the mean of the above 18 parameter values, the mean of the parameter type is the value obtained by calculating the mean of each column in the three columns of parameters respectively, and the formula for the sum of squared deviations can be expressed as:

[0089] ;

[0090] where SSA is the sum of squared deviations, is the mean of the parameter type corresponding to the jth type of parameter.

[0091] As shown in Table 1, the mean value of the parameter group corresponding to each PCIE card is the value obtained by calculating the mean of each row of parameters. The calculation formula for the sum of squared errors corresponding to each PCIE card can be expressed as:

[0092] ;

[0093] where SSE is the sum of squared errors, is the mean value of the parameter group corresponding to the i-th PCIE card.

[0094] The larger the sum of squared deviations, the greater the impact of the corresponding type of parameter on the failure of the PCIE card, thereby identifying the parameters that have a significant impact on the life of the PCIE card. The smaller the sum of squared deviations, the smaller the impact of the corresponding type of parameter on the failure of the PCIE card. The larger the sum of squared errors, the greater the impact of the parameters that deviate significantly from the mean value of the parameter group on the failure of the PCIE card. By determining the failure prediction order based on each sum of squared deviations and each sum of squared errors, and adjusting the pattern prediction model using the failure prediction order, a failure prediction model is obtained. Thus, the failure prediction model can output the failure prediction order sorted according to the impact degree of each parameter type on the PCIE card respectively, so as to remind the operation and maintenance personnel to conduct troubleshooting on the PCIE card according to the failure prediction order, realizing the timely troubleshooting of the important reasons for the failure.

[0095] In a specific embodiment of the present invention, determining the failure prediction order according to each sum of squared deviations and each sum of squared errors may include the following steps:

[0096] Step 1: Sort each parameter type according to each sum of squared deviations to obtain a first sorting result;

[0097] Step 2: Sort each parameter type according to each sum of squared errors to obtain a second sorting result;

[0098] Step 3: Obtain a first influence rate of the sum of squared deviations on failure prediction preset, and obtain a second influence rate of the sum of squared errors on failure prediction preset;

[0099] Step 4: Determine the failure prediction order according to the first sorting result, the second sorting result, the first influence rate and the second influence rate.

[0100] For the convenience of description, the above four steps can be combined for explanation.

[0101] Sort each type of parameter according to the sum of squared deviations to obtain the first sorting result, and sort each type of parameter according to the sum of squared errors to obtain the second sorting result. Since the second sorting results corresponding to different PCIe cards may vary, the initial second sorting results of multiple PCIe cards can be statistically analyzed in advance. The quantity statistics are respectively performed on various initial second sorting results, and the corresponding proportions of various initial second sorting results are determined according to the quantity statistics results. The second sorting result with a larger proportion is determined as the final second sorting result. Obtain the first influence rate of the preset sum of squared deviations on fault prediction, and obtain the second influence rate of the preset sum of squared errors on fault prediction. Determine the fault prediction order according to the first sorting result, the second sorting result, the first influence rate, and the second influence rate.

[0102] The first proportional values respectively corresponding to each sequence bit in the first sorting result can be preset, and the sum value of each first proportional value is 100%. The second proportional values respectively corresponding to each sequence bit in the second sorting result are preset, and the sum value of each second proportional value is 100%. By performing a multiplication operation on the first proportional value corresponding to each type of parameter and the first influence rate, a first operation result is obtained, and by performing a multiplication operation on the second proportional value corresponding to each type of parameter and the second influence rate, a second operation result is obtained. Calculate the sum value of the first operation result and the second operation result to obtain the target operation result corresponding to this type of parameter. Sort the target operation results respectively corresponding to each type of parameter in size to obtain the fault prediction order. Determine the fault prediction order through the proportional weighted calculation result, so that the higher the target operation result of a type of parameter, the higher the fault troubleshooting priority. Furthermore, faults can be detected and maintained in a timely manner, ensuring the safe operation of the device, reducing potential risks, and extending the service life of the device.

[0103] It should be noted that the first influence rate and the second influence rate can be set and adjusted according to the actual situation, and only the sum of the first influence rate and the second influence rate needs to be 100%. The embodiments of the present invention do not limit this.

[0104] In a specific implementation manner of the present invention, after obtaining the fault prediction model, the method may further include the following steps:

[0105] Step 1: Re-collect the historical usage data sets respectively corresponding to multiple PCIe cards, and determine the re-collected historical usage data sets as the test data sets;

[0106] Step 2: Use the test data sets to test the fault prediction model to obtain the test results;

[0107] Step 3: When there are parameter types in the failure prediction order determined according to the test results that have no impact on failure prediction, remove the parameter types that have no impact on failure prediction from the failure prediction order.

[0108] For ease of description, the above three steps can be combined for illustration.

[0109] After obtaining the failure prediction model, re-collect the historical usage data sets corresponding to multiple PCIe cards respectively, and determine the re-collected historical usage data sets as the test data sets. Use the test data sets to test the failure prediction model to obtain test results. When there are parameter types in the failure prediction order determined according to the test results that have no impact on failure prediction, remove the parameter types that have no impact on failure prediction from the failure prediction order. By promptly removing the parameter types that have no impact on failure prediction from the failure prediction order and selecting the remaining parameter types as the prediction basis, the troubleshooting of irrelevant parameter types is avoided, and the effectiveness of failure troubleshooting is improved.

[0110] In a specific embodiment of the present invention, after obtaining the failure prediction model, the method may further include the following steps:

[0111] When adding new parameter types, based on the parameter types after adding the new parameter types, repeat the step of calculating the mean values of the parameter types corresponding to the respective parameter types in the historical usage data set to update the failure prediction model to which it belongs.

[0112] After obtaining the failure prediction model, when adding new parameter types, based on the parameter types after adding the new parameter types, repeat the calculation of the mean values of the parameter types corresponding to the respective parameter types in the historical usage data set to obtain an updated failure prediction order, thereby realizing the update of the failure prediction model to which it belongs. By promptly adding parameter types that can have an impact on failure prediction according to the test results and updating the failure prediction order, the troubleshooting of irrelevant parameter types is avoided, and the effectiveness of failure troubleshooting is improved.

[0113] In addition, failure analysis can be performed on the prediction results of the failure prediction model, including verifying the accuracy and timeliness of the early warning. If false alarms or missed alarms are found, the failure prediction model is updated in a timely manner.

[0114] In a specific embodiment of the present invention, after outputting the failure prediction order sorted by the respective parameter types through the failure prediction model to perform PCIe card failure troubleshooting according to the failure prediction order, the method may further include the following steps:

[0115] Step 1: After troubleshooting the PCIe card according to the fault prediction sequence and taking corresponding maintenance strategies for maintenance, continue to run the PCIe card.

[0116] Step 2: When it is detected that the PCIe card is in an abnormal mode, output a prompt message indicating that the lifespan of the PCIe card has reached the upper limit.

[0117] For ease of description, the above two steps can be combined for explanation.

[0118] After troubleshooting the PCIe card according to the fault prediction sequence and taking corresponding maintenance strategies for maintenance, continue to run the PCIe card. When it is detected that the PCIe card is in an abnormal mode, it indicates that the PCIe card itself has reached the upper limit of its lifespan. Output a prompt message indicating that the lifespan of the PCIe card has reached the upper limit, so as to timely prompt the operation and maintenance personnel to replace the PCIe card.

[0119] The maintenance strategies for the PCIe card can include:

[0120] Heat dissipation maintenance: After the PCIe card gives an early warning, the maintenance personnel check whether the laboratory temperature exceeds the ideal value. If the laboratory temperature is too high, the internal heat dissipation of the server is affected, resulting in the PCIe card working in a high-temperature environment and affecting the lifespan of the PCIe card.

[0121] Voltage maintenance: After the PCIe card gives an early warning, the maintenance personnel measure whether the voltage of the server's PCIe card is normal with an electric meter.

[0122] Electrostatic protection: After the PCIe card gives an early warning, detect whether the static electricity charge in the laboratory exceeds the normal standard with an electrometer.

[0123] PCIe card optimization: After the PCIe card gives an early warning, check in the server system whether there are unclosed programs continuously accessing the PCIe card, resulting in a high usage frequency and affecting the service life of the PCIe card.

[0124] Protection and security: After an early warning is generated, the maintenance personnel check whether anyone has operated on the server, opened the chassis for operation, and caused misoperation on the PCIe card.

[0125] The maintenance personnel analyze the reason for the early warning based on the detection results of various maintenance methods. If the reason for the PCIe card to give an early warning is due to abnormal external environment, the environmental parameters can be adjusted, such as reducing the temperature, adjusting the voltage, reducing the static electricity charge, closing unnecessary programs, etc. After the environmental parameters are adjusted to the expected values, the server continues to run and check whether the PCIe card early warning still exists. If the PCIe card early warning is eliminated, it indicates that the PCIe card can still be used in a normal environment and there is no need to directly replace the PCIe card, thus reducing costs.

[0126] If all environmental factors are normal, it indicates that the actual service life of the PCIE card is about to reach its maximum life. The PCIE card needs to be replaced in a timely manner to avoid downtime and reduce production interruptions caused by failures in the production line or equipment, thereby improving production efficiency.

[0127] Corresponding to the above method embodiments, the present invention further provides a fault prediction device for a PCIE card. The fault prediction device for a PCIE card described below can be mutually corresponded and referred to with the fault prediction method for a PCIE card described above.

[0128] See Figure 2 , Figure 2 is a structural block diagram of a fault prediction device for a PCIE card in an embodiment of the present invention. The device may include:

[0129] An operating status data acquisition module 21, configured to monitor the operating status of the PCIE card to obtain operating status data;

[0130] A detection result acquisition module 22, configured to use a pre-trained fault prediction model to detect the operating status data to obtain a detection result;

[0131] A fault prediction module 23, configured to, when the detection result indicates that the PCIE card is in an abnormal mode, output a fault prediction order sorted by various parameter types through the fault prediction model, so as to perform fault troubleshooting on the PCIE card according to the fault prediction order.

[0132] The beneficial effects of the present invention are as follows: By pre-training a fault prediction model for predicting faults of the PCIE card, it is possible to predict in advance the possible faults of the PCIE card according to the operating status of the PCIE card, so as to perform fault troubleshooting on the PCIE card according to the fault prediction order sorted by various parameter types output by the fault prediction model, avoid downtime, and then greatly reduce production interruptions caused by failures in the production line or equipment, improving production efficiency. It avoids the further deterioration of faults and reduces the time and cost required for maintenance. Predictive maintenance is more efficient than regular maintenance, and maintenance is only carried out when it is truly necessary, avoiding unnecessary maintenance. Fault prediction can help detect potential equipment faults and take corresponding measures to avoid the occurrence of accidents. By timely discovering faults and performing maintenance, the safe operation of the equipment can be ensured, potential risks can be reduced, and the service life of the equipment can be extended.

[0133] In a specific embodiment of the present invention, the device may further include a model training module, and the model training module includes:

[0134] A historical data and operating mode acquisition sub-module, configured to acquire historical operating status data and operating modes respectively corresponding to multiple PCIE cards;

[0135] The mode prediction model obtaining sub-module is used to train a pre-constructed initial prediction model by using each piece of historical operation state data and each operation mode to obtain a mode prediction model;

[0136] The historical usage dataset collection sub-module is used to collect historical usage datasets corresponding to multiple PCIe cards respectively; wherein, the historical usage dataset contains parameter groups corresponding to multiple usage duration values respectively;

[0137] The fault prediction model obtaining sub-module is used to train the mode prediction model by using the historical usage dataset to obtain a fault prediction model.

[0138] In a specific embodiment of the present invention, the fault prediction model obtaining sub-module includes:

[0139] The overall mean calculation unit is used to calculate the overall mean of each parameter in the historical usage dataset;

[0140] The parameter type mean calculation unit is used to calculate the parameter type means corresponding to each parameter type in the historical usage dataset;

[0141] The sum of squared deviations calculation unit is used to calculate the sum of squared deviations corresponding to each parameter type according to each parameter type mean and the overall mean;

[0142] The parameter group mean calculation unit is used to calculate the parameter group means corresponding to each PCIe card in the historical usage dataset;

[0143] The sum of squared errors calculation unit is used to calculate the sum of squared errors corresponding to each PCIe card respectively according to each parameter group mean and each parameter value included in each parameter group;

[0144] The fault prediction order determination unit is used to determine the fault prediction order according to each sum of squared deviations and each sum of squared errors;

[0145] The fault prediction model obtaining unit is used to adjust the mode prediction model by using the fault prediction order to obtain a fault prediction model.

[0146] In a specific embodiment of the present invention, the fault prediction order determination unit includes:

[0147] The first sorting result obtaining sub-unit is used to sort each parameter type according to each sum of squared deviations to obtain a first sorting result;

[0148] The second sorting result obtaining sub-unit is used to sort each parameter type according to each sum of squared errors to obtain a second sorting result;

[0149] An influence rate acquisition subunit, configured to acquire a first influence rate of a preset sum of squared deviations on fault prediction, and acquire a second influence rate of a preset sum of squared errors on fault prediction;

[0150] A fault prediction order determination subunit, configured to determine a fault prediction order according to a first sorting result, a second sorting result, the first influence rate, and the second influence rate.

[0151] In a specific embodiment of the present invention, the device may further include:

[0152] A test data set determination module, configured to, after obtaining a fault prediction model, re-collect historical usage data sets corresponding to multiple PCIe cards respectively, and determine the re-collected historical usage data sets as test data sets;

[0153] A test result acquisition module, configured to test the fault prediction model by using the test data set to obtain a test result;

[0154] A parameter type elimination module, configured to, when there is a parameter type that has no influence on fault prediction among the parameter types included in the fault prediction order determined according to the test result, eliminate the parameter type that has no influence on fault prediction from the fault prediction order.

[0155] In a specific embodiment of the present invention, the device may further include:

[0156] A model update module, configured to, after obtaining a fault prediction model, when adding a new parameter type, repeat the step of calculating the mean value of each parameter type corresponding to each parameter type in the historical usage data set based on each parameter type after adding the new parameter type, so as to update the fault prediction model to which it belongs.

[0157] In a specific embodiment of the present invention, the device may further include:

[0158] A continued operation control module, configured to, after performing a fault troubleshooting on the PCIe card according to the fault prediction order and taking corresponding maintenance strategies for maintenance, continue to run the PCIe card;

[0159] A prompt message output module, configured to output a prompt message indicating that the life of the PCIe card has reached the upper limit when it is monitored that the PCIe card is in an abnormal mode.

[0160] Corresponding to the above method embodiment, see Figure 3 , Figure 3 is a schematic diagram of a fault prediction device for a PCIe card provided by the present invention. The device may include:

[0161] A memory 332, configured to store a computer program;

[0162] A processor 322, which is configured to execute a computer program to implement the steps of the method for predicting faults of a PCIe card in the above method embodiments.

[0163] Specifically, please refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of a device for predicting faults of a PCIe card provided in this embodiment. The device for predicting faults of the PCIe card may vary greatly due to different configurations or performances, and may include a processor (central processing units, CPU) 322 (for example, one or more processors) and a memory 332. The memory 332 stores one or more computer programs 342 or data 344. Among them, the memory 332 may be a transient storage or a persistent storage. The programs stored in the memory 332 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the data processing device. Further, the processor 322 may be configured to communicate with the memory 332 and execute a series of instruction operations in the memory 332 on the device 301 for predicting faults of the PCIe card.

[0164] The device 301 for predicting faults of the PCIe card may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.

[0165] The steps in the method for predicting faults of the PCIe card described above may be implemented by the structure of the device for predicting faults of the PCIe card.

[0166] Corresponding to the above method embodiments, the present invention further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the following steps may be implemented:

[0167] Monitor the operating state of the PCIe card to obtain operating state data; use the pre-trained fault prediction model to detect the operating state data to obtain a detection result; when the detection result indicates that the PCIe card is in an abnormal mode, output the fault prediction order sorted by each parameter type through the fault prediction model, so as to perform fault troubleshooting on the PCIe card according to the fault prediction order.

[0168] The computer-readable storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0169] For the introduction of the computer-readable storage medium provided by the present invention, please refer to the above method embodiments, and the present invention will not be elaborated herein.

[0170] Corresponding to the above method embodiments, the present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the steps of the previous fault prediction method of the PCIE card.

[0171] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices, equipment, and computer-readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for relevant parts.

[0172] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the technical solutions and core ideas of the present invention. It should be pointed out that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A PCIE card fault prediction method, characterized in that: include: Monitor the running status of the PCIE card to obtain running status data; Using the pre-trained fault prediction model to detect the operating status data to obtain a detection result; When the detection result is that the PCIE card is in an abnormal mode, the fault prediction model outputs a fault prediction sequence obtained by sorting the parameter types, so as to perform PCIE card fault troubleshooting according to the fault prediction sequence.

2. The PCIE card fault prediction method according to claim 1, characterized in that: The method also includes a training process of the fault prediction model, wherein the training process of the fault prediction model includes: Obtain historical operation status data and operation modes corresponding to multiple PCIE cards respectively; The pre-built initial prediction model is trained using each historical operation status data and each operation mode to obtain a mode prediction model; Collecting historical usage data sets corresponding to multiple PCIE cards respectively; wherein the historical usage data sets include parameter groups corresponding to multiple usage duration values ​​respectively; The pattern prediction model is trained using the historical usage data set to obtain the fault prediction model.

3. The PCIE card fault prediction method according to claim 2, characterized in that: The pattern prediction model is trained using the historical usage data set to obtain the fault prediction model, including: Calculating the overall mean of each parameter in the historical usage data set; Calculate the parameter type mean value corresponding to each parameter type in the historical usage data set; According to the mean of each parameter type and the overall mean, the sum of squares of deviations corresponding to each parameter type is calculated; Calculate the mean value of the parameter group corresponding to each PCIE card in the historical usage data set; According to the mean value of each parameter group and the parameter values ​​contained in each parameter group, the sum of square errors corresponding to each PCIE card is calculated respectively; Determining the fault prediction order according to the sum of squares of each deviation and the sum of squares of each error; The fault prediction sequence is used to adjust the pattern prediction model to obtain the fault prediction model.

4. The PCIE card fault prediction method according to claim 3, characterized in that: Determining the fault prediction order according to the sum of squares of the deviations and the sum of squares of the errors includes: Sort the parameter types according to the sum of squares of the deviations to obtain a first sorting result; Sort the parameter types according to the sum of squares of the errors to obtain a second sorting result; Obtaining a first influence rate of a preset sum of squared deviations on fault prediction, and obtaining a second influence rate of a preset sum of squared errors on fault prediction; The fault prediction order is determined according to the first sorting result, the second sorting result, the first impact rate, and the second impact rate.

5. The PCIE card fault prediction method according to claim 3 or 4, characterized in that: After obtaining the fault prediction model, the method further includes: Re-collecting historical usage data sets corresponding to the multiple PCIE cards respectively, and determining the re-collected historical usage data sets as the test data sets; Using the test data set to test the fault prediction model, and obtain a test result; When it is determined according to the test result that there are parameter types that have no effect on fault prediction among the parameter types included in the fault prediction sequence, the parameter types that have no effect on fault prediction are removed from the fault prediction sequence.

6. The PCIE card fault prediction method according to claim 3, characterized in that: After obtaining the fault prediction model, the method further includes: When a new parameter type is added, based on each parameter type after the new parameter type is added, the step of calculating the parameter type mean value corresponding to each parameter type in the historical usage data set is repeated to update the fault prediction model.

7. The PCIE card fault prediction method according to claim 1, characterized in that: After the fault prediction model outputs a fault prediction sequence obtained by sorting the parameter types, and PCIE card fault troubleshooting is performed according to the fault prediction sequence, the method further includes: After troubleshooting the PCIE card fault according to the fault prediction sequence and taking corresponding maintenance strategies for maintenance, continuing to operate the PCIE card; When it is detected that the PCIE card is in an abnormal mode, a prompt message indicating that the life of the PCIE card has reached an upper limit is output.

8. A PCIE card fault prediction device, characterized in that: include: The operation status data acquisition module is used to monitor the operation status of the PCIE card and obtain the operation status data; A detection result acquisition module is used to detect the operating status data using the pre-trained fault prediction model to obtain a detection result; The fault prediction module is used to output a fault prediction sequence obtained by sorting various parameter types through the fault prediction model when the detection result shows that the PCIE card is in an abnormal mode, so as to perform PCIE card fault troubleshooting according to the fault prediction sequence.

9. A PCIE card fault prediction device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the PCIE card fault prediction method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the fault prediction method for the PCIE card according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Fault recovery method and device for shared card in storage cluster

    CN120567653A

  • Fault recovery method and device for shared card under storage cluster

    CN120567653B