Real-time fault detection and diagnosis system for intelligent controller hardware
Through the real-time fault detection and diagnosis system, multi-dimensional fault monitoring and prediction of the intelligent controller hardware is realized, solving the problems of low fault detection efficiency and insufficient prediction in the existing technology, and improving the reliability and stability of the equipment.
Patent Information
- Application Number
- CN202510360652.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
The existing intelligent controller hardware fault detection and diagnosis methods are inefficient, difficult to achieve real-time monitoring and accurate diagnosis, and cannot effectively predict faults, resulting in unstable equipment operation and potential safety hazards.
A real-time fault detection and diagnosis system is designed, including hardware status monitoring, fault detection, fault diagnosis and processing, fault prediction and data storage modules. Through multi-dimensional data analysis and machine learning models, real-time status monitoring, fault potential screening, fault type judgment and severity assessment of intelligent controller hardware is realized, and fault processing is carried out in combination with environmental interference and task priority.
Improve the accuracy and adaptability of fault detection, ensure that critical tasks are not interrupted, extend the service life of the equipment, reduce maintenance costs, and enhance system reliability and stability.
Smart Images

Figure CN120255471A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of controller hardware fault detection, and specifically to a real-time fault detection and diagnosis system for intelligent controller hardware. Background Art
[0002] In the context of the rapid development of modern industrial automation and intelligent control systems, intelligent controller hardware is widely used in various devices, covering multiple fields such as industrial production, smart home, and automotive electronics, and has become a key component to ensure the stable operation of devices and achieve intelligent control. However, the reliability issue of intelligent controller hardware has always been an important factor restricting its further development and wide application.
[0003] Intelligent controller hardware is usually composed of many complex components, including microprocessors, memory chips, sensors, communication modules, etc. These components work together to achieve various functions of the device. During actual operation, these components are vulnerable to various factors and may malfunction. For example, in an industrial production environment, intelligent controller hardware will face harsh conditions such as high temperature, high humidity, and strong electromagnetic interference; in a smart home scenario, frequent power fluctuations may also damage it. Once a failure occurs, it will not only cause abnormal operation of the device, affect production efficiency and user experience, but may also lead to safety accidents and cause huge economic losses in severe cases.
[0004] Existing fault detection and diagnosis methods have many deficiencies. On the one hand, some detection means rely on manual regular inspections. This method is not only inefficient and difficult to achieve real-time monitoring, but also the accuracy of manual detection is easily affected by subjective factors, and potential fault hazards may be missed. On the other hand, most traditional fault diagnosis systems are based on simple threshold judgments and only monitor individual parameters, unable to comprehensively analyze the operating states of various components of intelligent controller hardware, and have limited diagnostic capabilities for complex faults. With the continuous expansion of the functions of intelligent controller hardware and the increasing performance requirements, the integration of its internal components is getting higher and higher, and the causes of faults are more complex and diverse. For example, when multiple components have minor faults at the same time, traditional detection systems may not be able to accurately determine the root cause of the fault, resulting in untimely or incorrect fault handling. At the same time, in the face of sudden faults, existing fault handling strategies often lack consideration of the overall operating tasks of the system, and may interrupt important tasks due to blind fault handling, further affecting the stability and reliability of the system.
[0005] At present, the fault prediction technology for intelligent controller hardware is relatively weak. In most cases, it can only handle faults after they occur, unable to predict the possibility of faults in advance, difficult to take effective preventive measures, and unable to meet the requirements of modern industry for high reliability and stability of equipment. Therefore, it is urgent to develop a system that can monitor the operating status of intelligent controller hardware in real time, accurately detect potential faults, deeply diagnose the types and severity of faults, and can predict faults in advance. This is of great significance for improving the reliability and stability of intelligent controller hardware and promoting the development of related industries. Summary of the Invention
[0006] The purpose of the present invention is to provide a real-time fault detection and diagnosis system for intelligent controller hardware to solve the problems raised in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical solution: A real-time fault detection and diagnosis system for intelligent controller hardware, the system includes:
[0008] A hardware status monitoring module for collecting the operating data of each component of the intelligent controller hardware in real time. The operating data includes voltage value, current value, temperature value, and clock frequency, and each component of the intelligent controller hardware is marked as each target component;
[0009] A fault detection module for determining the monitoring points of each target component of the intelligent controller hardware, obtaining the real-time operating parameters of each monitoring point according to the set monitoring period, and screening out the components with potential faults; The specific screening method is: obtaining the normal operating parameter range of each target component monitoring point from the historical database, comparing the real-time operating parameters with the normal range, and calculating the deviation rate where x represents the target component number, x = 1, 2,..., y, y is a positive integer greater than 2, n is the monitoring point number, n = 1, 2,..., m, m is a positive integer greater than 2, p xn is the real-time operating parameter, n xn is the normal operating parameter; setting a deviation rate threshold R, if r xn > R, then the component corresponding to this monitoring point has potential faults;
[0010] A fault diagnosis and processing module for deeply detecting the components with potential faults, analyzing the types and severity of faults, and taking corresponding processing measures; The fault processing module further includes:
[0011] A fault information collection unit for controlling the detection equipment to detect the components with potential faults and obtaining the circuit topology structure information, signal transmission waveform data, and chip internal register status information of the components;
[0012] A fault analysis unit, which is used to judge the fault type and evaluate the severity of the fault according to the collected fault information;
[0013] A fault handling unit, which is used to select a corresponding handling strategy according to the fault type and severity.
[0014] Preferably, the specific method for judging the fault type is as follows:
[0015] Based on the circuit topology information, analyze the current path and voltage distribution. If the current of a certain line increases abnormally and the voltage decreases abnormally, it is judged as a short - circuit fault; if there is no current passing through a certain line and the voltage increases abnormally, it is judged as an open - circuit fault;
[0016] Compare the signal transmission waveform data with the normal waveform template and calculate the waveform similarity where a i is the real - time waveform sampling value, b i is the normal waveform sampling value, N is the number of sampling points, and a similarity threshold S0 is set; calculate the frequency f real of the real - time waveform and the frequency deviation rate normal from the frequency f of the normal waveform. Set a frequency deviation rate threshold R f ; when r f >R f and the waveform similarity S < S0, it is judged as a signal transmission fault;
[0017] Judge the overheat fault according to the temperature value and the temperature change trend. When the temperature exceeds the upper limit of the normal operating temperature and continues to rise, it is judged as an overheat fault.
[0018] Preferably, the specific method for evaluating the severity of the fault is as follows:
[0019] Establish a fault impact factor model, considering the functional importance I x of the faulty component in the entire intelligent controller hardware, the influence range A x of the fault on other components, and the degree P x to which the fault causes the system performance to decline;
[0020] Calculate the fault severity index F x =I x ×A x ×P x ; Set different severity index intervals corresponding to different fault severity levels.
[0021] Preferably, when the hardware status monitoring module collects data, calculate the environmental interference coefficient of each target component. The specific calculation method is as follows:
[0022] Obtain the electromagnetic interference intensity E, humidity value H, and dust concentration value D of the environment where the intelligent controller hardware is located from the environmental monitoring database;
[0023] Set the electromagnetic interference intensity weight W E 、humidity weight W H 、dust concentration weight W D respectively, and W E +W H +W D = 1;
[0024] Calculate the environmental interference coefficient where E0, H0, and D0 are the normal reference values of the corresponding environmental factors respectively.
[0025] Preferably, when the fault detection module screens components with potential fault hazards, it considers the influence of the environmental interference coefficient on the operating parameters. The specific method is:
[0026] According to the environmental interference coefficient C, adjust the normal operating parameter range. The new lower limit upper limit where k is the adjustment coefficient;
[0027] According to the adjusted normal operating parameter range, recalculate the deviation rate and perform fault judgment.
[0028] Preferably, the adjustment coefficient k is a dynamic value and is adjusted according to different environmental interference factors and component types; establish an adjustment coefficient mapping table, which records the adjustment coefficient values corresponding to different combinations of environmental interference factors and component types; when the environmental interference coefficient C is calculated, look up the corresponding adjustment coefficient k from the mapping table according to the current environmental interference factors and component types.
[0029] Preferably, when the fault handling module selects a handling strategy, it considers the operation task priority of the intelligent controller hardware. The specific method is:
[0030] Obtain the task list currently being executed by the intelligent controller hardware and the priority of each task;
[0031] When a fault occurs, if the fault handling will affect the high-priority tasks being executed, preferably pause the fault handling first and wait until the high-priority tasks are completed or interruption is allowed before proceeding with the handling; if the fault handling does not affect the high-priority tasks, then immediately execute the fault handling operation.
[0032] Preferably, when considering the operation task priority to select a handling strategy, the remaining execution time of the task is also evaluated; when a fault occurs, if the fault handling will affect the high-priority tasks being executed, calculate the remaining execution time t remain of the high-priority tasks, and at the same time evaluate the time t required for fault handlingrepair ; if t repair < t remain and the high - priority task allows a short interruption within the remaining execution time, then fault handling is performed during the interruption of the high - priority task.
[0033] Preferably, it further includes a fault prediction module for predicting faults of each component of the intelligent controller hardware. The specific method is as follows:
[0034] Collect the historical operation data and fault data of each component, establish a fault prediction model based on machine learning, and construct a feature vector where p i is an operating parameter, and C is an environmental interference coefficient;
[0035] Use the trained model to predict the probability P of each component having a fault in the next period of time default ; Set a fault prediction probability threshold P0. When P default > P0, a fault warning is issued.
[0036] Preferably, the system further includes a data storage module for storing the collected operation data, fault data, and system configuration information.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] In the fault diagnosis link, the fault information acquisition unit in the fault diagnosis and handling module obtains multi - dimensional data such as the circuit topology structure information, signal transmission waveform data, and internal register status information of the chip of the component. The fault analysis unit accurately judges the fault type based on this information. For example, based on the circuit topology structure, analyzing the current path and voltage distribution to identify short - circuit and open - circuit faults, calculating the similarity and frequency deviation rate by comparing the signal transmission waveform data with the normal template to judge signal transmission faults, and judging overheating faults based on the temperature value and change trend. Compared with the traditional method that only relies on a single parameter threshold judgment, it can analyze the essence of faults more comprehensively and deeply, accurately distinguish various complex fault types, and provide a key basis for subsequent accurate fault handling.
[0039] In evaluating the severity of faults, a fault impact factor model is established to comprehensively consider the functional importance of faulty components, the scope of influence on other components, and the degree of system performance degradation. The fault severity index is calculated and graded. This enables maintenance personnel to quickly understand the impact of faults on the entire intelligent controller hardware system, prioritize the handling of severe faults, reasonably arrange maintenance resources, avoid exacerbating system problems due to improper fault handling order, and significantly improve the scientificity and efficiency of fault handling. The hardware status monitoring module calculates the environmental interference coefficient and adjusts the normal operating parameter range in the fault detection module according to this coefficient, while dynamically adjusting the coefficient k. This method effectively reduces the interference of environmental factors on fault detection, improves the accuracy and adaptability of fault detection. Whether in a complex environment with strong electromagnetic interference, high humidity, or high dust concentration, it can more accurately judge the operating status of components and reduce the occurrence of misjudgments.
[0040] In the selection of fault handling strategies, the priority of the hardware operation tasks of the intelligent controller and the remaining execution time of the tasks are combined. High-priority tasks are guaranteed to run first. Fault handling is carried out when it does not affect high-priority tasks or high-priority tasks allow interruption and the fault handling time is shorter than the remaining task time, avoiding the interruption of critical tasks caused by improper fault handling, ensuring the stability and reliability of the overall system operation, and enhancing the system's response ability in complex multi-task scenarios. The fault prediction module collects historical operation and fault data of components, constructs a machine learning-based model for fault probability prediction and sets a warning threshold. This function realizes the transformation from passive fault handling to active fault prevention. Maintenance personnel can perform maintenance or replacement on components in advance based on the warning, reducing the probability of faults, minimizing equipment downtime, improving production efficiency, reducing maintenance costs, extending the service life of the intelligent controller hardware, and enhancing the reliability and sustainable operation ability of the entire system. Description of the Drawings
[0041] Figure 1 It is the working principle diagram of the real-time fault detection and diagnosis system described in the present invention;
[0042] Figure 2 It is the working flow chart of the fault type judgment method;
[0043] Figure 3 It is the flow chart of environmental interference coefficient calculation and fault detection. Detailed Embodiments
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Please refer to Figures 1-3 , the present invention provides a technical solution: a real-time fault detection and diagnosis system for the hardware of an intelligent controller, the system includes:
[0046] The hardware status monitoring module is responsible for collecting the operation data of each component of the intelligent controller hardware in real time, covering voltage value, current value, temperature value and clock frequency. Each component of the intelligent controller hardware is marked as each target component. During the data collection process, the environmental interference coefficient of each target component will also be calculated. Specifically, obtain the electromagnetic interference intensity E, humidity value H and dust concentration value D of the environment where the intelligent controller hardware is located from the environmental monitoring database, and set the electromagnetic interference intensity weight W E , humidity weight W H , dust concentration weight W D , and satisfy W E +W H +W D =1. Calculate the environmental interference coefficient C through the formula , where E0, H0, D0 are the normal reference values of the corresponding environmental factors respectively.
[0047] The fault detection module will determine the monitoring points of each target component of the intelligent controller hardware, and obtain the real-time operation parameters of each monitoring point according to the set monitoring period. Obtain the normal operation parameter range of each target component monitoring point from the historical database, compare the real-time operation parameters with the normal range, and calculate the deviation rate through the formula , where x represents the target component number, x = 1, 2,..., y (y is a positive integer greater than 2), n is the monitoring point number, n = 1, 2,..., m (m is a positive integer greater than 2), p xn is the real-time operation parameter, n xn is the normal operation parameter. Set the deviation rate threshold R. If r xn >R, then the component corresponding to this monitoring point has a potential fault. And, when this module screens the components with potential faults, it will consider the influence of the environmental interference coefficient on the operation parameters, and adjust the normal operation parameter range according to the environmental interference coefficient C. The new normal operation parameter lower limit upper limit where k is the adjustment coefficient, and then recalculate the deviation rate and perform fault judgment according to the adjusted normal operation parameter range.
[0048] After the fault diagnosis and handling module determines the components with potential fault hazards, it conducts in-depth detection on them, analyzes the fault types and severity levels, and takes corresponding handling measures. This module further includes a fault information acquisition unit, a fault analysis unit, and a fault handling unit. The fault information acquisition unit controls the detection equipment to detect the components with potential fault hazards, and obtains the circuit topology structure information, signal transmission waveform data, and internal register status information of the chips of the components; the fault analysis unit determines the fault type based on the collected fault information and evaluates the fault severity level; the fault handling unit selects the corresponding handling strategy according to the fault type and severity level, and when selecting the handling strategy, it will consider the operation task priorities of the intelligent controller hardware. The fault prediction module collects the historical operation data and fault data of each component, establishes a fault prediction model based on machine learning, and constructs a feature vector x = (p1, p2, …, p n , C), where p i is an operation parameter and C is an environmental interference coefficient. The probability P default of each component having a fault within a certain period in the future is predicted using the trained model. A fault prediction probability threshold P0 is set. When P default > P0, a fault warning is issued.
[0049] The data storage module is used to store the collected operation data, fault data, and system configuration information, providing data support for the operation, fault analysis, and fault prediction of the system.
[0050] The present invention will be further described below in conjunction with Embodiments 1 to 5:
[0051] Embodiment 1:
[0052] In the fault analysis unit, when determining the fault type based on the circuit topology structure information, a circuit analysis tool is used to perform a detailed analysis on the obtained circuit topology. Taking the power supply circuit of a certain complex intelligent controller hardware as an example, this circuit supplies power to multiple core components. When it is monitored that the current of a line connecting a certain key chip increases abnormally, and at the same time the voltage across both ends of this line decreases abnormally, through the current and voltage monitoring functions of the circuit analysis tool and in combination with the circuit design principle, it is determined that this line has a short-circuit fault. Because under normal circumstances, according to Ohm's law when the resistance remains unchanged, the current should decrease as the voltage decreases, but at this time the current increases, indicating that the resistance of the line has changed, most likely a short-circuit situation has occurred. Similarly, if it is monitored that no current passes through a certain line and the voltage increases abnormally, in combination with the circuit structure and electrical characteristics, it is determined as an open-circuit fault.
[0053] For signal transmission waveform data, taking the communication interface circuit as an example, this circuit is responsible for data transmission between the intelligent controller hardware and external devices. In the normal working state, the communication interface sends and receives signals according to a specific communication protocol, and its signal transmission waveform has a specific shape and frequency. First, real-time signal transmission waveform data is obtained through a signal acquisition device and compared with the normal waveform template stored in the system. Assume that the normal waveform template is a standard waveform collected during multiple stable communication processes and processed through data. According to the formula
[0054]
[0055] calculate the waveform similarity S, where a i is the real-time waveform sampling value, b i is the normal waveform sampling value, and N is the number of sampling points. At the same time, use the frequency detection algorithm to calculate the real-time waveform frequency f real and the normal waveform frequency f normal , and then according to the formula
[0056]
[0057] calculate the frequency deviation rate r f . Set the similarity threshold S0 and the frequency deviation rate threshold R f , r f >R f and when the waveform similarity S < S0, it is judged as a signal transmission fault. For example, in a certain monitoring, the frequency deviation rate between the real-time waveform frequency and the normal frequency of the communication interface reaches 20% (greater than the set frequency deviation rate threshold R f = 10%), and at the same time, the waveform similarity is only 0.6 (less than the set similarity threshold S0 = 0.8). At this time, the system determines that there is a signal transmission fault in this communication interface.
[0058] In terms of judging overheating faults, taking the power chip in the intelligent controller hardware as an example, this chip generates a large amount of heat during operation. The temperature value of the chip is monitored in real time through a temperature sensor, and the temperature change trend within a period of time is recorded. Set the upper limit of the normal operating temperature of the power chip as T max When it is monitored that the chip temperature exceeds T max and continues to rise, it is judged as an overheating fault. For example, the upper limit of the normal operating temperature of a certain power chip is 80 °C. During operation, the temperature sensor detects that the chip temperature reaches 85 °C, and in subsequent monitoring, the temperature continues to rise. The system then determines that this power chip has an overheating fault.
[0059] Example 2:
[0060] When establishing a fault impact factor model, for each component in the intelligent controller hardware, its functional importance I is determined respectively. x The influence range A of the fault on other components x The degree P to which the fault causes the system performance to decline x Take the central processing unit (CPU) of the intelligent controller hardware as an example. Since the CPU is responsible for the core tasks of operation and control of the entire system, its functional importance I CPU is given a relatively high value. Assuming that according to system function analysis and expert experience, it is set to 0.9 (the value range is 0 - 1, and 1 represents the most important). When the CPU fails, it may cause the entire system to malfunction, and the influence range A CPU on other components is very large, and after evaluation, it is set to 0.8 (the value range is also 0 - 1, and 1 represents the largest influence range). If the CPU failure causes the system operation speed to drop by 50%, then the degree P CPU to which the fault causes the system performance to decline is 0.5 (the value range is determined according to the performance decline ratio).
[0061] According to the formula F x = I x × A x × P x calculate the fault severity index F x . For the above - mentioned CPU failure situation, the calculated fault severity index F CPU = 0.9 × 0.8 × 0.5 = 0.36.
[0062] Set different severity index intervals to correspond to different fault severity levels. For example, divide the severity index interval into: [0, 0.2) is the minor fault level, [0.2, 0.5) is the moderate fault level, [0.5, 1] is the severe fault level. Since the severity index F CPU of the CPU failure is 0.36, so this CPU failure is determined to be a moderate fault.
[0063] For other components, such as the storage chip, assume its functional importance I 存储芯片 = 0.6. When the storage chip fails, it may only affect the storage and reading of some data, and the influence range A 存储芯片 on other components is 0.4. If the failure causes the system data read - write speed to drop by 30%, then P 存储芯片 = 0.3. Calculate the fault severity index F 存储芯片 of the storage chip as F = 0.6 × 0.4 × 0.3 = 0.072, and this fault belongs to the minor fault level. In this way, the fault severity of different components can be quantitatively evaluated, providing a scientific basis for the selection of fault handling strategies.
[0064] Example 3:
[0065] This example focuses on the calculation process of the environmental interference coefficient, as well as its specific applications in adjusting the normal operating parameter range and fault detection. Meanwhile, the dynamic adjustment mechanism of the adjustment coefficient is elaborated to improve the accuracy of fault detection. When the hardware status monitoring module calculates the environmental interference coefficient, it obtains the electromagnetic interference intensity E, humidity value H, and dust concentration value D of the environment where the intelligent controller hardware is located from the environmental monitoring database through sensors or data interfaces. For example, in a certain industrial production environment, the electromagnetic interference intensity E = 50 V / m is measured by an electromagnetic interference monitor, the humidity value H = 60% is measured by a humidity sensor, and the dust concentration value D = 0.5 mg / m 3 .
[0066] The electromagnetic interference intensity weight W E is respectively set to 0.4, the humidity weight W H is set to 0.3, the dust concentration weight W D is set to 0.3, and it satisfies W E +W H +W D = 1. Assuming that the normal reference value of the electromagnetic interference intensity E0 = 30 V / m, the normal reference value of the humidity H0 = 50%, and the normal reference value of the dust concentration D0 = 0.3 mg / m 3 , according to the formula calculate the environmental interference coefficient C, that is
[0067]
[0068] When the fault detection module screens for components with potential fault hazards, it considers the influence of the environmental interference coefficient on the operating parameters. Taking a certain resistor component as an example, assuming that its normal operating parameter range is [95Ω, 105Ω], the normal operating parameter n 电阻 = 100Ω of the monitoring point of this resistor component is obtained from the historical database, and the adjustment coefficient k = 0.1 (a fixed value is assumed here first to illustrate the adjustment process, and the dynamic adjustment mechanism will be elaborated later). According to the environmental interference coefficient C = 1.53, calculate the new lower limit of the normal operating parameter upper limit
[0069] When the real-time operating parameter p 电阻 = 120Ω of this resistor component is obtained, calculate the deviation rate according to the adjusted normal operating parameter range (If the deviation rate is calculated according to the unadjusted normal range, it is Set the deviation rate threshold R = 0.15. Since the deviation rate r 电阻= 0.2 > R, so it is determined that there are potential faults in this resistor component, and accurate fault judgment may not be possible when calculated according to the unadjusted range.
[0070] Regarding the implementation of the adjustment coefficient k as a dynamic value, an adjustment coefficient mapping table is established. For example, in the mapping table, for a high electromagnetic interference intensity (E > E 高阈值 ), high humidity (H > H 高阈值 ) and in an environment with a high dust concentration (D > D 高阈值 ), and at the same time for the case of the power component, the adjustment coefficient k = 0.2 is set; for a low electromagnetic interference intensity (E < E 低阈值 ), normal humidity (H 低阈值 < H < H 高阈值 ) and low dust concentration (D < D 低阈值 ) for ordinary logic components, the adjustment coefficient k = 0.05 is set, etc. After the environmental interference coefficient C is calculated, according to the current environmental interference factors (such as the comparison results of the actual values of E, H, D with the thresholds) and the component type (determined by information such as component identification or circuit location), the corresponding adjustment coefficient k is found from the mapping table, so as to achieve precise adjustment of the normal operating parameter range and improve the accuracy of fault detection.
[0071] Example 4:
[0072] When the fault handling module selects a handling strategy, first obtain the task list currently being executed by the intelligent controller hardware and the priorities of each task. For example, the intelligent controller hardware is executing three tasks: Task A is to control the operation of equipment on the production line in real time, and the priority is set to high; Task B is data collection and storage, with a medium priority; Task C is system self-check, with a low priority.
[0073] When a certain component fails, assume that the failed component affects the operation of Task A and Task B. If the fault handling will affect the currently executing high-priority Task A, the system will first determine whether Task A allows interruption. If Task A does not allow interruption, the system suspends the fault handling and waits until Task A is completed before performing the fault handling. If Task A allows interruption, at this time, calculate the remaining execution time t remain of the high-priority Task A, and at the same time evaluate the time t repair required for fault handling. For example, through the task scheduling system and fault handling experience data, estimate that the remaining execution time t remain of Task A = 10s, and the time t repair required for fault handling = 5s. Because t repair < t remain and Task A allows a short interruption within the remaining execution time, the system performs fault handling during the interruption of Task A.
[0074] If the fault handling does not affect high-priority tasks, taking the case where the faulty component only affects Task C as an example, since Task C has a lower priority, the system immediately performs the fault handling operation to prevent the fault from further expanding and affecting other tasks. In this way, while ensuring the normal operation of high-priority tasks, the fault is timely handled, improving the stability and reliability of the system.
[0075] Embodiment 5:
[0076] This embodiment details how to construct a prediction model by means of data collection and machine learning techniques, and then accurately predict and warn of potential faults in each component of the intelligent controller hardware.
[0077] The system continuously collects the historical operation data of each component of the intelligent controller hardware, including key parameters such as voltage value, current value, temperature value, clock frequency, etc. These data reflect the operation characteristics of the components under different working states. At the same time, the fault data of the components are collected, including information such as the time of fault occurrence, fault type, and operation parameter status at the time of fault, so as to analyze the law of fault occurrence later. For example, for a specific model of power transistor component, its operation data is recorded every 10 minutes during its continuous operation for 3 months, and a total of 5 faults occur during this period. The relevant data of each fault are detailedly recorded.
[0078] Combined with the environmental interference coefficient obtained by the hardware status monitoring module, a feature vector is constructed. Taking the power transistor component as an example, its feature vector combined with the environmental interference coefficient obtained by the hardware status monitoring module, a feature vector is constructed. Taking the power transistor component as an example, its feature vector x = (p 电压 , p 电流 , p 温度 , p 时钟频率 , C), where p 电压、 p 电流、 p 温度、 p 时钟频率 are the operation parameters of this component, and C is the environmental interference coefficient at the corresponding moment. In this way, the component operation state is combined with external environmental factors, providing a more comprehensive data basis for subsequent model training.
[0079] In terms of model construction, the long short-term memory network (LSTM) algorithm in deep learning is adopted. The LSTM network can effectively process time series data and capture long-term dependencies in the data, making it very suitable for scenarios of predicting future component failures based on historical operation data. When building the LSTM model, first determine the number of network layers and the number of neurons in each layer. After multiple experiments and optimizations, it is determined to use 3 LSTM layers, with 128, 64, and 32 neurons respectively in each layer, and finally connect a fully connected layer to output the prediction results. In the model training stage, the collected historical data is preprocessed. The operation parameter data is normalized and mapped to the interval [0, 1] to eliminate the influence of the dimension between different parameters and improve the efficiency and stability of model training. The environmental interference coefficient is also normalized. The preprocessed data is divided according to the ratio of 70% as the training set and 30% as the test set. The LSTM model is trained using the training set. During the training process, the mean squared error (MSE) is used as the loss function, and the Adam optimizer is used to adjust the weight parameters of the model. The learning rate is set to 0.001, and the number of training epochs is set to 200. Through continuous iterative training, the model gradually learns the complex relationship between the component operation parameters, environmental interference coefficient, and the probability of failure occurrence.
[0080] After training is completed, the test set is used to evaluate the model. Calculate metrics such as the accuracy, recall rate, and F1 value of the model prediction results to evaluate the performance of the model. After multiple tests and adjustments, the accuracy of this LSTM model on the test set reaches 90%, the recall rate reaches 85%, and the F1 value is 0.87, indicating that the model has good prediction performance.
[0081] Use the trained model to predict the probability P of each component failing within a certain period in the future default . Set the probability of the component failing within the next 48 hours. For the above-mentioned power transistor component, after model prediction, the probability P of it failing within the next 48 hours default = 0.3. Set the failure prediction probability threshold P0 = 0.25. Since P default > P0, the system issues a failure warning. At this time, the maintenance personnel can, based on the warning information, check, maintain, or replace the power transistor component in advance to avoid the sudden occurrence of a failure during operation, thereby ensuring the stable operation of the intelligent controller hardware system.
[0082] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0083] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A real-time fault detection and diagnosis system for intelligent controller hardware, characterized in that, The system includes: A hardware status monitoring module, which is used to collect the operation data of each component of the intelligent controller hardware in real time. The operation data includes voltage value, current value, temperature value and clock frequency, and marks each component of the intelligent controller hardware as each target component; A fault detection module is used to determine the monitoring points of each target component of the intelligent controller hardware, obtain the real-time operating parameters of each monitoring point according to the set monitoring period, and screen out the components with potential faults. The specific screening method is as follows: obtain the normal operating parameter range of the monitoring points of each target component from the historical database, compare the real-time operating parameters with the normal range, and calculate the deviation rate. Where x represents the target component number, x = 1, 2,..., y, y is a positive integer greater than 2, n is the monitoring point number, n = 1, 2,..., m, m is a positive integer greater than 2, p xn is the real-time operating parameter, n xn is the normal operating parameter; set the deviation rate threshold R, if r xn > R, then the component corresponding to this monitoring point has potential faults. A fault diagnosis and processing module, which is used to deeply detect the components with potential faults, analyze the fault types and severity, and take corresponding treatment measures; The fault processing module further includes: A fault information acquisition unit, which is used to control the detection equipment to detect the components with potential faults, and obtain the circuit topology information, signal transmission waveform data and chip internal register status information of the components; A fault analysis unit, which is used to judge the fault type and evaluate the fault severity according to the collected fault information; A fault processing unit, which is used to select the corresponding treatment strategy according to the fault type and severity.
2. The real-time fault detection and diagnosis system for the intelligent controller hardware according to claim 1, characterized in that, The specific method for judging the fault type is: Based on the circuit topology information, analyze the current path and voltage distribution. If the current of a certain line increases abnormally and the voltage decreases abnormally, it is judged as a short circuit fault; If there is no current passing through a certain line and the voltage increases abnormally, it is judged as an open circuit fault; Compare the signal transmission waveform data with the normal waveform template and calculate the waveform similarity where a i is the real-time waveform sampling value, b i is the normal waveform sampling value, N is the number of sampling points, and a similarity threshold S0 is set; calculate the real-time waveform frequency f real and the frequency deviation rate of the normal waveform frequency f normal Set the frequency deviation rate threshold R f ; when r f > R f and the waveform similarity S < S0, it is judged as a signal transmission failure; Judge the overheat fault according to the temperature value and the temperature change trend. When the temperature exceeds the upper limit of the normal working temperature and continues to rise, it is judged as an overheat fault.
3. A real-time fault detection and diagnosis system for the hardware of an intelligent controller, according to claim 2, characterized in that The specific method for evaluating the fault severity is: Establish a fault impact factor model, considering the functional importance I of the faulty component in the entire intelligent controller hardware x 、the influence range A of the fault on other components x 、the degree P of system performance degradation caused by the fault x ; Calculate the fault severity index F x = I x × A x × P x ; Set different severity index intervals corresponding to different fault severity levels.
4. A real-time fault detection and diagnosis system for the hardware of an intelligent controller, characterized in that, When the hardware status monitoring module collects data, calculate the environmental interference coefficient of each target component. The specific calculation method is: Obtain the electromagnetic interference intensity E, humidity value H and dust concentration value D of the environment where the intelligent controller hardware is located from the environmental monitoring database; Set the electromagnetic interference intensity weight W respectively E and the humidity weight W H and the dust concentration weight W D , and W E +W H +W D = 1; Calculation of environmental interference coefficient where E0, H0, and D0 are the normal reference values of the corresponding environmental factors respectively.
5. The real-time fault detection and diagnosis system for the intelligent controller hardware according to claim 4, characterized in that, When the fault detection module screens the components with potential faults, consider the influence of the environmental interference coefficient on the operation parameters. The specific method is: Adjust the normal operating parameter range according to the environmental interference coefficient C, and the new lower limit of the normal operating parameters upper limit where k is the adjustment coefficient; Recalculate the deviation rate according to the adjusted normal operation parameter range and conduct fault judgment.
6. The real-time fault detection and diagnosis system for the intelligent controller hardware according to claim 5, characterized in that, The adjustment coefficient k is a dynamic value, which is adjusted according to different environmental interference factors and component types; Establish an adjustment coefficient mapping table, which records the adjustment coefficient values corresponding to different combinations of environmental interference factors and component types; After the environmental interference coefficient C is calculated, find the corresponding adjustment coefficient k from the mapping table according to the current environmental interference factors and component types.
7. A real-time fault detection and diagnosis system for the intelligent controller hardware according to claim 1, characterized in that, When the fault processing module selects the treatment strategy, consider the operation task priority of the intelligent controller hardware. The specific method is: Obtain the task list currently being executed by the intelligent controller hardware and the priority of each task; When a fault occurs, if the fault processing will affect the high-priority tasks being executed, give priority to suspending the fault processing and wait until the high-priority tasks are completed or interruption is allowed before proceeding; If the fault processing does not affect the high-priority tasks, the fault processing operation will be executed immediately.
8. The real-time fault detection and diagnosis system for the intelligent controller hardware according to claim 7, characterized in that, When considering the processing strategy for running task priority selection, the remaining execution time of the task is also evaluated; when a failure occurs, if the failure handling will affect the high-priority task being executed, calculate the remaining execution time t of the high-priority task remain , and at the same time evaluate the time t required for failure handling repair ; if t repair < t remain and the high-priority task allows a short interruption within the remaining execution time, then the failure handling is performed during the interruption of the high-priority task.
9. The real-time fault detection and diagnosis system for the intelligent controller hardware according to claim 1, characterized in that, It also includes a fault prediction module, which is used to predict faults for each component of the intelligent controller hardware. The specific method is: Collect the historical operation data and fault data of each component, establish a fault prediction model based on machine learning, and construct a feature vector where p i is the operating parameter, and C is the environmental interference coefficient; Use the trained model to predict the probability P of each component failing within a certain period in the future default ; Set the fault prediction probability threshold P0. When P default > P0, issue a fault warning.
10. A real-time fault detection and diagnosis system for the hardware of an intelligent controller, characterized in that, The system also includes a data storage module, which is used to store the collected operation data, fault data and system configuration information.
Citation Information
Cited By
Vehicle remote diagnosis method and device, electronic equipment and storage medium
CN120630959A
Vehicle remote diagnosis method and device, electronic equipment and storage medium
CN120630959B