Fault monitoring and self-recovery method and system for industrial personal computer

By deploying microcontroller units and neural network models on industrial control machines, fault detection and self-recovery of industrial control machines is realized, and the problems of fault detection in the existing technology are solved, which reduces downtime and maintenance costs, and ensures that the industrial control machines operate stably in complex environments.

CN120085635APending Publication Date: 2025-06-03SUZHOU APQI INTERNET OF THINGS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510238492.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing industrial-controlled machine fault detection system lacks real-time and accuracy in abnormal identification in complex and extreme environments, resulting in misjudgment or loss of fault detection, increasing downtime and maintenance costs.

Method used

By deploying a microcontroller unit (MCU) on the industrial control machine, real-time monitoring is performed before and after the industrial control machine is started, parameter data of CPU, memory, hard disk, graphics card, motherboard and fan are collected, data preprocessing and fault prediction are used for training neural network models (such as KNN models), and fault recovery is carried out through an automated fault handling module.

Benefits of technology

It improves the real-time and accuracy of fault detection, reduces manual intervention, reduces downtime and maintenance costs, ensures that the industrial control machine operates stably in extremely complex industrial environments, avoids production interruptions, and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085635A_ABST
    Figure CN120085635A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial automation online monitoring, in particular to a fault monitoring and self-recovery method and system for an industrial personal computer, and the method comprises the steps: detecting and regulating the external environment temperature of the industrial personal computer to accord with a starting mechanism before power-on, and starting industrial personal computer equipment; hardware information of the industrial personal computer is read, circulation monitoring is carried out, whether hardware faults exist or not is checked, and parameter data of the industrial personal computer are synchronously collected; and through the prediction model, the processed parameter data is evaluated, whether non-hardware abnormity exists is judged, and fault self-recovery and abnormity processing are carried out. According to the invention, real-time monitoring is carried out through the MCU before and after the industrial personal computer is started, so that the real-time performance and accuracy of fault detection are improved; fault recovery is automatically carried out according to a fault analysis result, manual intervention is reduced, and downtime and maintenance cost are reduced; through environment monitoring and control, stable operation of the industrial personal computer in an extremely complex industrial environment is ensured; and through real-time monitoring and fault prediction, production interruption is avoided, and the production efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial automation online monitoring, and particularly to a method and system for fault monitoring and self - recovery of an industrial control computer. Background Art

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, industrial control hosts are increasingly widely used in unattended fields such as industrial automation, intelligent manufacturing, and robots. Industrial control computers need to operate stably in complex and changeable industrial environments such as high temperature, low temperature, humidity, dryness, and electromagnetic interference. However, after a failure occurs in existing industrial control computer devices, manual intervention is often required to resume operation, which greatly increases the downtime and maintenance costs.

[0003] Industrial control computer devices usually operate in extreme industrial environments such as high temperature, humidity, and strong electromagnetic interference. Traditional fault monitoring systems are difficult to adapt to complex environmental changes, resulting in loss or misjudgment of fault information. The probability of faults occurring is relatively high, and the detection system lacks real - time performance and accuracy.

[0004] Existing fault detection methods usually rely on regular manual inspections and rule - based diagnostic methods, which are not only time - consuming and laborious but also have the possibility of misdiagnosis and missed diagnosis. Especially in the startup phase and the operating system boot phase, the system has not loaded all necessary hardware or software support, making it difficult to effectively identify abnormalities.

[0005] Based on the problems in the existing technology, the present invention provides a method and system for fault monitoring and self - recovery of an industrial control computer. Summary of the Invention

[0006] The object of the present invention is to provide a method and system for fault monitoring and self - recovery of an industrial control computer to solve the technical problem that the fault detection system in the existing technology lacks real - time performance and accuracy in identifying abnormalities in complex and extreme environments.

[0007] The technical solution of the present invention is: A method for fault monitoring and self - recovery of an industrial control computer includes:

[0008] Before power - on, detect and regulate the external environmental temperature of the industrial control computer to meet the startup mechanism, and start the industrial control computer device;

[0009] Read the hardware information of the industrial control computer and enter a loop for monitoring, check whether there is a fault code, and synchronously collect the parameter data of the CPU, memory, hard disk, graphics card, motherboard, and fan of the industrial control computer;

[0010] When there is a fault code, recover by restarting. If self - recovery cannot be achieved by restarting, perform exception handling;

[0011] The prediction model obtained by training the neural network evaluates the parameters after processing the acquired data, and outputs whether the display device status is faulty or normal.

[0012] Preferably, the parameter data of the industrial control computer and the corresponding fault status labels are recorded one by one; the fault status labels are represented binary, 0 represents the normal state, and 1 represents the fault state.

[0013] Preprocess the collected parameters, including data denoising and normalization:

[0014] Divide the preprocessed data into a training set and a prediction set;

[0015] The neural network uses the KNN model, dynamically determines the value of K based on the data point density of the training set. The algorithm in the KNN model uses the Euclidean distance to measure the similarity between samples. Each time a prediction is made, the KNN model classifies according to the distance between the newly input sample and the training set samples.

[0016] For a new sample, first calculate the distance between the new sample and each sample in the training set, then select the K samples with the smallest distance to it, and make a prediction based on the labels of the neighbors; if the number of neighbors with a fault label of 1 is greater than the number of neighbors with a fault label of 0, then predict as faulty, otherwise predict as normal. Evaluate whether the model can capture faults in time through the recall rate of the device.

[0017] Preferably, when monitoring and collecting the working parameters of the industrial control computer, automatically adjust the sampling frequency and sampling points according to the current operating state of the industrial control computer;

[0018] When the system is working normally, the sampling frequency is reduced to reduce unnecessary data redundancy; when the system shows abnormalities or enters the fault mode, the sampling frequency is automatically increased to obtain more hardware status information.

[0019] Preferably, during the training process of the KNN model, weight according to the distance;

[0020] Assign a greater weight to the neighbors with a closer distance, and calculate the weight of each neighbor using the inverse distance weight;

[0021] For each prediction sample x new , we perform a weighted vote according to the labels of the K closest neighbors; assume the labels of the K neighbors are y 1 , y 2 ,…, y k , and their corresponding weights are w 1 , w 2 ,…, w k , the calculation method of the result of the weighted vote is as follows:

[0022]

[0023] If the result is greater than 0.5, it is predicted as a fault; otherwise, it is predicted as normal.

[0024] Preferably, parameter data of the CPU, memory, hard disk, graphics card, motherboard, and fan of the industrial control computer are collected. The specific parameters include:

[0025] CPU base frequency, CPU real-time frequency, CPU load ratio, CPU operating temperature, CPU power; memory ratio;

[0026] Hard disk operating temperature, hard disk read and write speed, storage space ratio, partition space ratio, S.M.A.R.T. status;

[0027] GPU load, GPU core temperature, video memory occupancy, ECC error, GPU power;

[0028] Motherboard temperature, power supply voltage, power supply current, power supply power;

[0029] Fan: actual speed and target speed.

[0030] Preferably, the collected parameters are preprocessed, including data denoising and normalization processing:

[0031] Through the Z-Score model, outliers in the working environment parameters of the industrial control computer collected are detected, and the detected outliers are processed to remove noise data to reduce the impact of abnormal data on fault prediction. The processing methods include deleting samples, replacing outliers, and difference processing;

[0032] The Min-Max method is used to normalize the data after denoising processing, and all data is scaled to the range of [0,1], so that all feature data is converted into dimensionless values, ensuring that the contribution of each feature in the model is balanced.

[0033] Preferably, after the industrial control computer is powered on, the micro control unit reads the hardware information and enters the monitoring loop. The micro control unit reads the fault codes of port 80H or PCH GPIO through the LPC / eSPI interface and detects whether there are fault codes, including hard disk fault codes and BIOS error codes;

[0034] If there are fault codes, the micro control unit determines whether the hardware boots normally. If the BIOS cannot boot the system, a restart attempt is initiated; if the system still cannot be restored after restarting more than three times, the micro control unit enters the exception handling process and records the exception information in the local storage;

[0035] If a hard disk anomaly is detected, the backup hard disk or the backup system disk is controlled to power on and start, and the device anomaly information is sent through the network interface to prompt the user that the main hard disk is abnormal and needs to be processed in time.

[0036] Preferably, when a micro-fault occurs in the model evaluation system device, indicating a non-hard disk abnormality in the device, the fault information is displayed on the screen; and the device abnormality information and the corresponding POST 80H fault code are sent through the network interface.

[0037] Preferably, before power-on, if the environmental temperature of the industrial control computer is lower than the normal operating range, the heating device is started to make the environmental temperature reach the normal temperature range of the industrial control computer, the heating is stopped and the industrial control computer is started; if the temperature parameter is not detected before power-on, the industrial control computer is directly powered on and started by default.

[0038] A fault monitoring and self-recovery system for an industrial control computer, used to implement the method for fault monitoring and self-recovery of an industrial control computer, includes an external environment data acquisition module, an industrial control computer data acquisition module, a data transmission module, an MCU data analysis module, and a fault processing module;

[0039] The external environment data acquisition module includes temperature and humidity sensors, and the temperature and humidity sensors are connected to the data transmission module and transmit data;

[0040] The industrial control computer data acquisition module runs on the industrial control computer and includes an LPC / eSPI and a status data acquisition module running on the operating system. The industrial control computer data acquisition module is used to collect the startup phase error information of the industrial control computer and the status data of the industrial control computer after the operating system runs;

[0041] The data transmission module uses a hardware interface and is connected to the industrial control computer and the data analysis module, and transmits data to the MCU data analysis module through the hardware interface;

[0042] The MCU data analysis module conducts power supply and data acquisition interaction through the data transmission module, starts before the industrial control computer starts, and can process and analyze the data before and after the industrial control computer is powered on;

[0043] The fault processing module automatically processes the fault according to the fault information given by the MCU data analysis module, and the processing content includes external environment heating, power-on and power-off control, standby system switching, and system backup and restoration;

[0044] After the system is normally started, the MCU enters a continuous monitoring and protection stage to ensure the stability of the system and perform fault recovery as needed.

[0045] The data acquisition module continuously collects the hardware status of the industrial control computer and sends the data to the MCU for processing. After receiving the data, the MCU will analyze the fault and predict the fault. After detecting the fault, the MCU will transfer the fault information to the fault processing module for corresponding fault handling.

[0046] Compared with the prior art, the advantages of the present invention are as follows:

[0047] Before and after the industrial control computer is started, the present invention performs real-time monitoring through the MCU, improving the real-time performance and accuracy of fault detection; automatically performing fault recovery according to the fault analysis results, reducing manual intervention, and reducing downtime and maintenance costs; ensuring the stable operation of the industrial control computer in an extremely complex industrial environment through environmental monitoring and control; avoiding production interruptions and improving production efficiency through real-time monitoring and fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The present invention will be further described below in conjunction with the drawings and embodiments:

[0049] Figure 1 It is a block diagram of the method for fault monitoring and self-recovery of the industrial control computer described in the present invention;

[0050] Figure 2 It is a schematic diagram of the parameter fault status label of the industrial control computer described in the present invention;

[0051] Figure 3 It is a schematic structural diagram of the fault monitoring and self-recovery system for the industrial control computer described in the present invention;

[0052] Figure 4 It is a schematic flow diagram of the method for fault monitoring and self-recovery of the industrial control computer described in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The content of the present invention will be further described in detail below in conjunction with specific embodiments:

[0054] As Figure 1 shown, a method for fault monitoring and self-recovery of an industrial control computer, in combination with the attached Figure 4 , the specific process is as follows:

[0055] Temperature detection and control process before the system is powered on.

[0056] Before the system is powered on, the micro control unit is first turned on. When the micro control unit starts, its doorman card reads the temperature data fed back by the external sensor through the UART interface, and judges whether the current ambient temperature is suitable for starting the industrial control computer (IPC) device by reading the external temperature sensor (such as NTC thermistor, digital temperature sensor, etc.);

[0057] If the temperature is within the normal operating range of the IPC, directly enter the normal device startup process.

[0058] If the temperature is lower than the normal operating range of the IPC (such as -10 degrees Celsius), the doorman card of the microcontroller unit controls the heating device (such as a heating pad, an electric heating tube, etc.) to start through the GPIO interface. When the external environmental temperature reaches the normal operating temperature range of the IPC, the heating device is controlled to turn off and the IPC device is started.

[0059] If temperature data cannot be obtained (such as the sensor is not installed or malfunctioning), the microcontroller unit (MCU) directly defaults to controlling the device to enter the startup process.

[0060] After the industrial computer device system starts, the microcontroller unit enters the continuous monitoring and protection stage to ensure the stability of the system and perform fault recovery as needed.

[0061] After the industrial computer starts, on the one hand, it monitors the hardware status during the operation of the industrial computer.

[0062] When the industrial computer is powered on, the microcontroller unit (MCU) reads the fault codes of port 80H or PCH GPIO through the LPC / eSPI interface, reads the hardware information and enters the monitoring loop to check whether there are fault codes in the hardware (such as CPU, memory, hard disk fault codes, BIOS error codes, etc.);

[0063] If there is a fault code, the MCU will determine whether the hardware can be normally booted. If the BIOS cannot boot the system, a restart attempt will be initiated.

[0064] If the system fails to recover after restarting more than three times, the MCU enters the exception handling process and records the exception information in the local storage.

[0065] Furthermore, in the case of a hard disk exception, the system automatically controls the backup hard disk or the backup system disk to be powered on for restart, and sends device exception information through the network interface to prompt the user that the main hard disk is abnormal and needs to be processed in time.

[0066] On the other hand, it collects the parameter data in the working environment of the industrial computer and judges whether there are non-hard disk exceptions in the device according to the real-time parameter data during operation. The specific content is as follows:

[0067] After the industrial computer is powered on, it collects industrial computer data:

[0068] CPU-related data includes: the rated operating frequency (base frequency) of the CPU, the current operating frequency (real-time frequency) of the CPU, the current load of the CPU, usually expressed as a percentage (CPU load), CPU temperature, and CPU power;

[0069] The memory occupancy, usually expressed in GB or as a percentage;

[0070] Hard disk related parameters include: hard disk temperature, read and write speed, storage space occupancy, partition space occupancy, S.M.A.R.T. status (Self-Monitoring, Analysis and Reporting Technology status, including hard disk health and error prediction).

[0071] Graphics card related parameters include: the current workload of the graphics card (GPU load), computing engine occupancy (usage of each computing engine of the GPU), GPU core temperature, video memory occupancy, ECC errors, and GPU power.

[0072] Motherboard related parameters include: motherboard temperature, supply voltage, supply current, supply power,

[0073] Fan related parameters include: actual fan speed, fan target speed.

[0074] Each parameter data collected is recorded together with the corresponding fault status label. The label uses binary annotation, where 0 represents the normal state and 1 represents the fault state. The parameter fault status representation refers to the appendix Figure 2 .

[0075] In addition, when continuously monitoring the operating status of the industrial control computer, the sampling frequency and sampling points are automatically adjusted based on the current operating status of the industrial control computer. When the system is working properly, the sampling frequency will be automatically reduced to reduce unnecessary data redundancy; when the system has an abnormality or enters the fault mode, the sampling frequency will be automatically increased to more closely monitor the hardware status.

[0076] The original data collected is denoised and normalized.

[0077] Abnormal data caused by sensor failures or external interferences is removed through denoising, reducing the impact of abnormal data on fault prediction and improving the data quality and the accuracy of subsequent models.

[0078] Abnormal data includes abnormally high temperatures (temperature data far higher than the actual device temperature) or very low power readings (such as negative values).

[0079] For each feature x ij , Z-Score is used to detect outliers. If the feature value x ij of a certain sample exceeds the set threshold, it is considered an outlier and should be removed or processed.

[0080] The Z-Score outlier detection formula is:

[0081]

[0082] where x ij represents the j-th feature value of the i-th sample, u j is the mean of the j-th feature, and σ jis the standard deviation. If ∣z ij ∣>t (for example, t = 3), where t = 3 means that if a certain eigenvalue exceeds the mean by more than 3 times the standard deviation, then the data point is considered an outlier.

[0083] Using the above method to perform noise reduction processing on the CPU, memory, hard disk, graphics card, motherboard, fan and other data of the collected sample data. For the detected outliers, the following processing methods can be taken:

[0084] Delete the sample: If a certain eigenvalue of a certain sample is an outlier, then delete the sample;

[0085] Replace the outlier: Replace the outlier with the mean, median or other statistics;

[0086] Interpolation processing: Perform interpolation processing according to the values of adjacent samples.

[0087] After performing noise reduction processing on the collected industrial control computer device data, perform normalization processing on the data. Different sensor data usually have different dimensions (for example, the unit of temperature is degrees Celsius, power is watts, and load is percentage). In order to prevent the influence of certain features on model training from being too large, it is necessary to perform normalization processing on the data.

[0088] Use the Min-Max method to normalize the data, which scales all data to the range of [0,1]. The formula for Min-Max normalization is:

[0089]

[0090] where: x is the original data point, x min represents the minimum value of this feature, x max represents the maximum value of this feature, and x' is the normalized data.

[0091] After normalization processing, all feature data are converted into dimensionless numerical values, ensuring that the contribution of each feature in the model is balanced.

[0092] Build a neural network model to analyze and process the data to estimate whether there are non-hardware anomalies in the device.

[0093] Use the KNN model as the basic model, divide the normalized data set into a training set and a test set, and use 80% of the data as the training set and 20% of the data as the test set.

[0094] In the KNN model, the value of K (the number of neighbors) directly affects the classification effect of the KNN model. To adapt to different datasets, the value of K is dynamically determined based on the data point density, and K = sqrt(sample number). The KNN algorithm uses the Euclidean distance to measure the similarity between samples, and its calculation formula is:

[0095]

[0096] where x i =(x i1 ,x i2 ,x i3 ,…,x in ) and x j =(x j1 ,x j2 ,x j3 ,…,x n ) are the feature vectors of two samples, n is the dimension of the features, and the Euclidean distance calculates the straight-line distance between two points.

[0097] KNN is an instance-based learning algorithm that does not require an explicit training process; each time a prediction is made, KNN will classify based on the distance between the newly input sample and the training set samples.

[0098] For a new sample x new , first calculate the distance d(x i ,x new ,x i ) between the new sample and each sample x i in the training set, and then select the K samples with the smallest distance, and make a prediction based on the labels of these neighbors.

[0099] The specific prediction process is as follows:

[0100] For the new sample x new , sort it according to its distance from all samples in the training set, and select the K neighbors with the smallest distance.

[0101] Assume that the labels of the selected K neighbors are y 1 ,y 2 ,…,y k , and determine the label y new of x new according to the majority voting principle new That is:

[0102] y new =majority(y 1 ,y 2 ,…,y k );

[0103] If the number of faulty tags (tag value is 1) among the K neighbors is greater than the number of normal tags, then it is predicted as faulty, and the output is shown as 1; otherwise, it is predicted as normal, and the output is shown as 0.

[0104] Furthermore, based on distance weighting, the prediction result is further refined by assigning a greater weight to the nearer neighbors. The inverse distance weighting is used to calculate the weight of each neighbor:

[0105] Inverse distance weighting: The weight of a neighbor is proportional to the reciprocal of its distance, that is, the nearer the distance, the greater the weight.

[0106]

[0107] For each predicted sample x new , weighted voting is performed according to the tags of the K nearest neighbors.

[0108] Assume the tags of the K neighbors are y 1 , y 2 , …, y k , and their corresponding weights are w 1 , w 2 , …, w k ,

[0109] The calculation method of the weighted voting result is as follows:

[0110] If the result is greater than 0.5, then it is predicted as faulty, and the output is shown as 1; otherwise, it is predicted as normal, and the output is shown as 0.

[0111] For the prediction of industrial computer faults, the recall rate of the device is concerned. The recall rate of the device is used to evaluate whether the model can capture faults in a timely manner and quantitatively evaluate the reliability of the evaluation model.

[0112] If the model predicts and determines that the device is not a hard disk anomaly, then the fault information is displayed and the device anomaly information and the corresponding POST 80H fault code are sent through the network interface.

[0113] Furthermore, the present invention provides a fault monitoring and self - recovery system for an industrial computer, which is used to implement the above - mentioned method for fault monitoring and self - recovery of an industrial computer, so as to solve the current situation that the current industrial computer can only locate faults after the device crashes or only monitors the hardware status at the operating system layer and requires manual intervention for recovery.

[0114] The system includes: an external environment data acquisition module, an industrial computer data acquisition module, a data transmission module, an MCU data analysis module, and a fault processing module.

[0115] The external environment data acquisition module includes temperature and humidity sensors, which are connected to the data transmission module via RS-485 to transmit data.

[0116] The industrial control computer data acquisition module runs on the industrial control computer and includes LPC / eSPI and the status data acquisition module running on the operating system. The industrial control computer data acquisition module collects the error information during the startup phase of the industrial control computer and the status data of the industrial control computer after the operating system runs.

[0117] The data acquisition module continuously collects the hardware status and data of the industrial control computer and delivers them to the MCU for processing. After receiving the data, the MCU will analyze the faults and predict the faults. After detecting a fault, the MCU will transmit the fault information to the fault handling module for corresponding fault handling.

[0118] The data transmission module uses a hardware interface to connect to the industrial control computer and the data analysis module, and transmits the data to the MCU data analysis module through the hardware interface.

[0119] The MCU data analysis module uses an independent MCU board card to interact with the data transmission module for power supply and data acquisition. This module starts before the industrial control computer starts and can process and analyze the data before and after the industrial control computer is powered on.

[0120] The fault handling module automatically processes the faults according to the fault information given by the data analysis module, including functions such as external environment heating, power on / off control, standby system switching, and system backup and restoration.

[0121] As shown in the Figure 3 accompanying figure, a schematic diagram of an embodiment of the system structure diagram is provided, and the working principle is as follows:

[0122] Monitor the operating environment of the industrial control computer to ensure that the industrial control computer can be normally powered on and work in a harsh environment;

[0123] The method of the present invention analyzes the faults by reading the fault codes during the boot process of the industrial control computer, monitors and intervenes in the boot process, and then automatically recovers the faults of the industrial control computer according to the fault analysis results.

[0124] After the industrial control computer enters the operating system, the MCU will interact with the assistant software running on the industrial control computer, report the recorded abnormal information to the assistant, and at the same time the assistant will monitor the operating system, CPU, motherboard, memory, hard disk, graphics card, network card and software environment of the industrial control computer. When a fault occurs in the device, it will send a command to the MCU for fault recovery.

[0125] In summary, the present invention conducts real-time monitoring before and after the industrial control computer is started by the MCU, improving the real-time performance and accuracy of fault detection; automatically performing fault recovery according to the fault analysis results, reducing manual intervention, and decreasing the downtime and maintenance costs; ensuring the stable operation of the industrial control computer in an extremely complex industrial environment through environmental monitoring and control; and avoiding production interruptions and improving production efficiency through real-time monitoring and fault prediction.

[0126] The above embodiments are only used to illustrate the technical concept and features of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.

Claims

1. A method for fault monitoring and self-recovery for an industrial computer, characterized in that: include: Before powering on, detect and adjust the external ambient temperature of the industrial computer to meet the startup mechanism, and start the industrial computer equipment; Read the hardware information of the industrial computer and enter the loop monitoring to check whether there is a fault code, and simultaneously collect the parameter data of the CPU, memory, hard disk, graphics card, motherboard and fan of the industrial computer; When a fault code exists, it is recovered by restarting. If the restart cannot recover automatically, the exception is handled; The prediction model obtained by training the neural network evaluates the parameters after the acquired data is processed, and the output shows the equipment status as faulty or normal.

2. A method for fault monitoring and self-recovery for an industrial computer according to claim 1, characterized in that: Collect the parameter data of the industrial computer and the corresponding fault status label and record them one by one; the fault status label is binary, 0 represents normal state, and 1 represents fault state; Preprocess the collected parameters, including data denoising and normalization: Divide the preprocessed data into training set and prediction set; The neural network uses the KNN model, which dynamically determines the K value based on the density of data points in the training set. The algorithm in the KNN model uses the Euclidean distance to measure the similarity between samples. Each time a prediction is made, the KNN model classifies the newly input sample based on the distance between the sample in the training set. For a new sample, we first calculate the distance between the new sample and each sample in the training set, then select the K samples with the smallest distance to them and make predictions based on the labels of their neighbors. If the number of fault labels of 1 among the K neighbors is greater than the number of fault labels of 0, then it is predicted to be a fault, otherwise it is predicted to be normal. The recall rate of the device is used to evaluate whether the model can capture the fault in time.

3. The method for fault monitoring and self-recovery for industrial computers according to claim 1, characterized in that: When monitoring and collecting the working parameters of the industrial computer, the sampling frequency and sampling point are automatically adjusted according to the current operating status of the industrial computer; When the system is operating normally, the sampling frequency is reduced to reduce unnecessary data redundancy; when the system is abnormal or enters failure mode, the sampling frequency is automatically increased to obtain more hardware status information.

4. A method for fault monitoring and self-recovery for an industrial computer according to claim 2, characterized in that: During the KNN model training process, weighting is performed based on distance; Give greater weight to neighbors with closer distances, and use inverse distance weighting to calculate the weight of each neighbor; For each prediction sample x new , we perform weighted voting based on the labels of the K nearest neighbors; assuming that the labels of the K neighbors are y1, y2, …, y k , their corresponding weights are w1,w2,…,w k , the result of weighted voting is calculated as follows: If the result is greater than 0.5, the prediction is failure; otherwise, the prediction is normal.

5. The method for fault monitoring and self-recovery for industrial computers according to claim 1, characterized in that: Collect parameter data of the industrial computer's CPU, memory, hard disk, graphics card, motherboard, and fan. Specific parameters include: CPU base frequency, CPU real-time frequency, CPU load ratio, CPU operating temperature, CPU power; memory ratio; Hard disk operating temperature, hard disk read and write speed, storage space ratio, partition space ratio, SMART status; GPU load, GPU core temperature, video memory usage, ECC errors, GPU power; Mainboard temperature, power supply voltage, power supply current, power supply power; Fan: actual speed and target speed.

6. A method for fault monitoring and self-recovery for an industrial computer according to claim 2, characterized in that: Preprocess the collected parameters, including data denoising and normalization: Through the Z-Score model, the abnormal values ​​of the collected industrial computer working environment parameters are detected, and the detected abnormal values ​​are processed to remove noise data to reduce the impact of abnormal data on fault prediction. The processing methods include deleting samples, replacing abnormal values, and difference processing; The Min-Max method is used to normalize the denoised data and scale all data to the range of [0,1] so that all feature data are converted into dimensionless values ​​to ensure that the contribution of each feature in the model is balanced.

7. The method for fault monitoring and self-recovery for industrial computers according to claim 1, characterized in that: After the industrial computer is powered on, the microcontroller reads the hardware information and enters the monitoring loop. The microcontroller reads the fault code of port 80H or PCH GPIO through the LPC / eSPI interface and detects whether there is a fault code, including hard disk fault code and BIOS error code; If a fault code is present, the microcontroller unit determines whether the hardware boots normally, and if the BIOS cannot boot the system, a restart attempt is initiated; If the system cannot be restored after more than three restarts, the microcontroller unit enters the exception handling process and records the exception information to local storage; If a hard disk abnormality is detected, the backup hard disk or backup system disk is controlled to power on and start, and a device abnormality message is sent through the network interface to prompt the user that the main hard disk is abnormal and needs to be processed in time.

8. The method for fault monitoring and self-recovery for industrial computers according to claim 2, characterized in that: When the model evaluation system device has a micro-fault, it indicates that the device has a non-hard disk abnormality, and the fault information is displayed on the screen; and the device abnormality information and the corresponding POST 80H fault code are sent through the network interface.

9. The method for fault monitoring and self-recovery for industrial computers according to claim 1, characterized in that: Before power-on, if the ambient temperature of the industrial computer is lower than the normal working range, the heating device will be started to make the ambient temperature reach the normal temperature range of the industrial computer, and the heating will be stopped and the industrial computer will be started; if no temperature parameters are detected before power-on, the industrial computer will be powered on directly by default.

10. A fault monitoring and self-recovery system for an industrial computer, used to implement a fault monitoring and self-recovery method for an industrial computer as claimed in any one of claims 1 to 9, characterized in that: It includes external environment data acquisition module, industrial computer data acquisition module, data transmission module, MCU data analysis module and fault handling module; The external environment data acquisition module includes a temperature and humidity sensor, which is connected to the data transmission module and transmits data; The industrial computer data acquisition module runs on the industrial computer, including LPC / eSPI and a status data acquisition module running on the operating system. The industrial computer data acquisition module is used to collect error information during the startup phase of the industrial computer and status data of the industrial computer after the operating system is running; The data transmission module uses a hardware interface to connect with the industrial computer and the data analysis module, and transmits data to the MCU data analysis module through the hardware interface; The MCU data analysis module interacts with the data transmission module for power supply and data collection. It is started before the industrial computer is started and can process and analyze the data before and after the industrial computer is started. The fault processing module automatically processes the fault according to the fault information provided by the MCU data analysis module. The processing content includes outsourcing environment heating, power on / off control, standby system switching and system backup and restoration. After the system starts normally, the MCU enters the continuous monitoring and protection phase to ensure system stability and perform fault recovery as needed; The data acquisition module continuously collects the hardware status and data of the industrial computer and sends it to the MCU for processing. After receiving the data, the MCU will analyze and predict the fault. After detecting the fault, the MCU will pass the fault information to the fault handling module for corresponding fault handling.

Citation Information

Cited By

  • Industrial personal computer operation monitoring system and method based on data analysis

    CN120704205A

  • An industrial computer operation monitoring system and method based on data analysis

    CN120704205B