Chip failure detection method, device and system for relay protection device
By combining hardware and software testing in the relay protection device to conduct comprehensive failure detection on the chip, the problem of incomplete and inaccurate chip failure detection is solved, and the safety and reliability of the device are improved.
Patent Information
- Application Number
- CN202411934658.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-03
AI Technical Summary
The chip failure detection of the relay protection device is incomplete and inaccurate, and there are safety hazards.
By obtaining the chip's hardware detection data, including operating temperature, voltage and current, hardware failure detection is performed; if the hardware test result is failure, software failure test is performed, comprehensive failure detection is performed based on the hardware and software test results, failure probability and type are evaluated, and the reinforcement plan is determined.
It improves the accuracy and comprehensiveness of chip failure detection of relay protection devices, and enhances the safety and reliability of the device.
Smart Images

Figure CN120085140A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grids, and particularly to a method, device, and system for detecting chip failures in a relay protection device. Background Art
[0002] With the development of grid digitization and intelligence, relay protection devices have increasingly high requirements for the computing performance and storage capacity of the hardware chips used. However, the advancement of the manufacturing process of core chips has reduced their critical voltage, leading to an increased probability of storage unit or software anomalies and an increased risk of device failure.
[0003] Currently, although there are some technologies for detecting and correcting chip anomalies, they do not fully cover all core chips or key links, resulting in potential safety hazards in relay protection devices. For example, the CPU of a relay protection device works in a high-temperature environment for a long time, leading to a decline in the aging performance of the CPU chip. Using conventional temperature, voltage, and current detection methods to monitor the CPU status and give early warnings may have problems with inaccurate detection of CPU failures. Another example is that chips such as FPGAs in relay protection devices are often detected by methods such as program and data verification. When the hardware performance of the chip deteriorates and the program and data verification are normal, the failure detection of the chip is incomplete and inaccurate. Summary of the Invention
[0004] The present invention provides a method, device, and system for detecting chip failures in a relay protection device, which can solve the technical problem of incomplete and inaccurate detection of chip failures in relay protection devices, improve the accuracy and comprehensiveness of chip failure detection in relay protection devices, and enhance the safety and reliability of relay protection devices.
[0005] In a first aspect, the present invention provides a method for detecting chip failures in a relay protection device, the method comprising: obtaining hardware detection data of a chip in a relay protection device during operation in a current period; the hardware detection data includes operating temperature, operating voltage, and operating current; based on the hardware detection data and a preset threshold, performing a hardware failure detection on the relay protection device to determine a hardware test result of the chip in the relay protection device; if the hardware test result is that the chip has failed, performing a software failure test on the chip in the relay protection device to obtain a software test result; based on the hardware test result and the software test result, determining a failure detection result of the chip in the relay protection device; the failure detection result includes a failure probability and a failure type; based on the failure detection result, determining a reinforcement plan for the chip in the relay protection device.
[0006] In a second aspect, an embodiment of the present invention provides a chip failure detection device for a relay protection device, including: a communication module, specifically configured to obtain hardware detection data of a chip in the relay protection device during operation in the current period; the hardware detection data includes operating temperature, operating voltage, and operating current; a processing module, specifically configured to perform hardware failure detection on the relay protection device based on the hardware detection data and a preset threshold to determine a hardware test result of the chip in the relay protection device; if the hardware test result indicates chip failure, perform software failure testing on the chip in the relay protection device to obtain a software test result; determine a failure detection result of the chip in the relay protection device based on the hardware test result and the software test result; the failure detection result includes a failure probability and a failure type; determine a reinforcement scheme for the chip in the relay protection device based on the failure detection result.
[0007] In a third aspect, an embodiment of the present invention provides a chip failure detection system for a relay protection device. The chip failure detection system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the steps of the method described in the first aspect and any possible implementation manner in the first aspect.
[0008] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in the first aspect and any possible implementation manner in the first aspect.
[0009] The present invention provides a chip failure detection method, device, and system for a relay protection device. Aiming at the problems of incomplete and inaccurate chip detection, the present invention first performs hardware detection on the chip of the relay protection device and then performs software testing. Through the combination of hardware testing and software testing, comprehensive failure detection of the chip is carried out to evaluate the failure probability and failure type of the chip, improving the accuracy and comprehensiveness of chip failure detection of the relay protection device and enhancing the safety and reliability of the relay protection device. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 is a schematic flowchart of a chip failure detection method for a relay protection device provided by an embodiment of the present invention;
[0012] Figure 2 It is a schematic structural diagram of a chip failure detection device for a relay protection device provided by an embodiment of the present invention.
[0013] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0014] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented in order to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0015] In the description of the present invention, unless otherwise specified, " / " means "or". For example, A / B can represent A or B. Herein, "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" and "a plurality of" refer to two or more. The words such as "first" and "second" do not limit the quantity and execution order, and the words such as "first" and "second" do not necessarily limit being different.
[0016] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner for easy understanding.
[0017] In addition, the terms "including" and "having" mentioned in the description of the present application and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but may alternatively include other steps or modules not listed, or may alternatively include other steps or modules inherent to these processes, methods, products, or devices.
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will be described through specific embodiments in conjunction with the drawings of the present invention.
[0019] Such as Figure 1As shown in the figure, an embodiment of the present invention provides a method for detecting chip failures in a relay protection device. The method includes steps S101 - S105.
[0020] S101. Obtain the hardware detection data of the chip during the operation of the relay protection device in the current period.
[0021] In the embodiment of the present application, the hardware detection data includes the operating temperature, operating voltage, and operating current.
[0022] In the present invention, in - depth research on the failure mechanism of the core chips in the relay protection device, such as CPU, FPGA, etc., has been carried out. By analyzing the hardware failure mechanism and the effectiveness of anti - error design in the existing protection software and hardware architecture, the present invention can more accurately evaluate the risk of chip failure and provide strong support for subsequent reinforcement measures. According to the research results of the failure mechanism, the present invention proposes a hardware reinforcement scheme for the relay protection device to prevent abnormal storage of a single - core hardware. This includes circuit - structure reinforcement, such as adopting redundant design, adding filter circuits, etc., and system - level reinforcement measures, such as optimizing the heat - dissipation design and improving the power - supply stability. These measures are aimed at reducing the probability of chip failure and improving the reliability of the relay protection device. At the same time, the present invention also studies the risk - control measures for the protection program or sampled data and proposes a software reinforcement scheme for the relay protection device. By eliminating the influence of single - event upsets of CPU and FPGA and abnormal data in a single channel on the device, the present invention can ensure the stability of software operation and the accuracy of data. Finally, in order to verify the effectiveness of the software - hardware integrated reinforcement method, the present invention constructs an independently controllable verification system for the failure of the storage unit of the relay protection device. This system can simulate single - core hardware faults and memory - data anomalies to comprehensively test and evaluate the relay protection device. By comparing the performance before and after reinforcement, the present invention can verify the effectiveness of the reinforcement measures and provide a basis for subsequent improvements.
[0023] This detection process is a systematic and automated process for evaluating the failure risk and verifying the reinforcement effect of the core components of the relay protection device (including hardware such as CPU, FPGA, and software protection programs). The whole process ensures high performance and stability of the device under various working conditions through precise monitoring and data analysis. This detection process aims to comprehensively evaluate the failure risk and verify the reinforcement effect of the core hardware (such as CPU, FPGA) and software of the relay protection device through an automated program to ensure the stable operation of the device in a complex environment. The whole process is divided into four main stages: parameter detection, data processing, failure determination, and implementation of solution measures.
[0024] Hardware parameter detection: temperature detection, voltage and current monitoring, signal - integrity verification, and redundant - circuit - state detection.
[0025] Temperature Monitoring: High-precision temperature sensors are used to continuously monitor the operating temperatures of key chips such as CPUs and FPGAs, and the data is transmitted to the data analysis system in real time. By setting temperature thresholds, potential thermal runaway risks can be detected and warned in a timely manner.
[0026] Voltage and Current Monitoring: Voltage monitors and current sensors are installed to measure the power supply voltage and current fluctuations of each key circuit in real time. Parameters such as voltage stability and current peaks are analyzed to evaluate the power quality and load status.
[0027] Signal Integrity Verification: High-speed oscilloscopes or signal analyzers are used to accurately measure and analyze the waveforms of input and output signals. By comparing the standard signals with the measured signals, abnormal conditions such as signal distortion and interference are identified.
[0028] Redundant Circuit Status Detection: Dedicated status monitoring circuits are configured for the parts with redundant designs. The operating status of the standby circuits and the switching logic with the main circuits are tested regularly to ensure seamless switching at critical moments.
[0029] S102. Based on the hardware detection data and preset thresholds, perform hardware failure detection on the relay protection device to determine the hardware test results of the chips in the relay protection device.
[0030] As a possible implementation method, step S102 can be specifically implemented as steps S1021 - S1025.
[0031] S1021. Based on the hardware detection data and preset thresholds, perform threshold judgment to determine whether the hardware detection data is within the normal range.
[0032] S1022. If the hardware detection data is within the normal range, determine that the hardware test result is that the chip is normal.
[0033] S1023. If the hardware detection data is not within the normal range, based on the hardware detection data, perform signal integrity detection and redundant circuit status detection to determine the signal integrity of the hardware detection data and the redundant circuit status of the chips in the relay protection device.
[0034] S1024. If there are signal abnormalities in the hardware detection data, or the redundant circuits are operating abnormally, determine that the hardware test result is that the chip has failed.
[0035] S1025. If the signals of the hardware detection data are normal and the redundant circuits are operating normally, determine that the hardware test result is that the chip is normal.
[0036] S103. If the hardware test result is that the chip has failed, perform software failure testing on the chips in the relay protection device to obtain the software test results.
[0037] As a possible implementation, step S103 can be specifically implemented as steps S1031 - S1034.
[0038] S1031. Perform a program execution path tracking test on the chip in the relay protection device, record the chip program execution parameters, and the execution parameters include the number of calls, return values, and execution time.
[0039] S1032. Perform a sampling data verification on the chip in the relay protection device to determine the sampling data verification result.
[0040] S1033. Perform a single - event upset simulation test on the chip in the relay protection device to determine the response of single - event upset.
[0041] S1034. Based on the chip program execution parameters, the sampling data verification result, and the response of single - event upset, determine the software test result.
[0042] Software parameter detection includes program execution path tracking, sampling data verification, and single - event upset simulation test.
[0043] Program execution path tracking: Through debugging tools or code injection techniques, track the execution flow of the protection program. Record parameters such as the number of calls, return values, and execution time of key functions to evaluate the normality of the program operation.
[0044] Sampling data verification: Use various data verification methods such as CRC verification and parity verification to comprehensively check the sampling data. Ensure the integrity and accuracy of the sampling data to avoid misoperation caused by data errors.
[0045] Single - event upset simulation test: Use a dedicated single - event upset simulator or through radiation experiment means to simulate single - event upset events of chips such as CPUs and FPGAs. Monitor the response of the program when single - event upset occurs and evaluate its fault - tolerance ability.
[0046] S104. Based on the hardware test result and the software test result, determine the failure detection result of the chip in the relay protection device.
[0047] In the embodiments of the present application, the failure detection result includes the failure probability and the failure type.
[0048] As a possible implementation, step S104 can be specifically implemented as steps S1041 - S1044.
[0049] S1041. Generate a test vector based on the hardware test result and the software test result.
[0050] In some embodiments, the test vectors include the operating temperature, operating voltage, operating current, chip program execution parameters, sampling data verification results, and the response to single-event upsets.
[0051] S1042. Generate a failure detection result based on the test vectors and a preset test model.
[0052] In some embodiments, the preset test model is obtained by testing a chip of the same type as the chip in the relay protection device.
[0053] In the embodiments of the present invention, data collection and storage: The automated script periodically collects the detection parameters from each monitoring point and stores them in a high-performance database or a real-time processing system to ensure the integrity and traceability of the data. Anomaly identification algorithm: Use technical means such as statistical methods, machine learning algorithms, or expert systems to deeply analyze the collected data. By comparing historical data, preset thresholds, or model prediction results, identify potential abnormal data points. Trend analysis and prediction: Conduct trend analysis on time series data and use time series prediction models (such as ARIMA, LSTM, etc.) to predict the parameter change trend in the future for a period of time. Timely discover potential fault points and provide a basis for taking measures in advance.
[0054] Hardware failure determination criteria: Set clear failure thresholds for parameters such as temperature, voltage, and current. When any parameter continuously exceeds the threshold, it is determined as a hardware failure risk. Analyze the signal integrity detection results. If abnormal situations such as signal distortion, loss, or interference occur frequently, it is regarded as a potential hardware failure. Regularly check the working status and switching logic of the redundant circuit. If the standby circuit fails to work properly or the switching fails, it is determined that the hardware redundancy mechanism fails.
[0055] Software failure determination criteria: Compare the program execution path tracking result with the preset normal process. If it is found that the program execution deviates or the return value of the key function is abnormal, it is determined as a software error. When the number of sampling data verification failures reaches the preset threshold, it is considered that there is a problem with the sampling data, which may lead to software malfunction. In the single-event upset simulation test, if the program fails to respond correctly or the response time is too long, it is determined that the software fault tolerance ability is insufficient.
[0056] S105. Determine a reinforcement plan for the chip in the relay protection device based on the failure detection result.
[0057] As a possible implementation, step S105 can be specifically implemented as steps S1051 - S1053.
[0058] S1051. If the failure type is hardware failure and the failure probability is greater than the preset probability, determine that the reinforcement plan is hardware upgrade.
[0059] S1052: If the failure type is software failure and the failure probability is less than or equal to the preset probability, then determine the reinforcement plan as algorithm optimization.
[0060] S1053: If the failure type is software failure and the failure probability is greater than the preset probability, then determine the reinforcement plan as hardware upgrade and algorithm optimization.
[0061] The solutions after failure include: Automatic switching of redundant components: When the hardware failure risk is detected, the system automatically activates the redundant switching mechanism. According to the preset switching logic and priority order, the critical tasks are automatically switched to the standby circuit or chip for execution. During the switching process, the system automatically performs status synchronization and data verification to ensure the smoothness and stability of the switching process.
[0062] Software restart and recovery: When the software fails, the system automatically attempts to restart the protection program. During the restart process, the system conducts a comprehensive self-check and initialization operation to ensure that the program can resume normal operation. If the restart is ineffective or the problem persists, the system triggers a higher-level recovery operation (such as restoring default configuration, loading backup program, etc.).
[0063] Fault recording and reporting: The system automatically records the detailed information of each failure event, including timestamp, failure type, parameter values, handling measures and results, etc. And these information are stored in the local database or sent to the remote monitoring center through the network. After receiving the fault report, the monitoring center immediately conducts fault analysis and location. According to the analysis results, corresponding solutions are formulated and sent to the on-site equipment for execution.
[0064] Dynamic adjustment of reinforcement strategy: According to the statistical analysis results of the failure type and frequency, the system automatically evaluates the effectiveness of the current reinforcement measures. If it is found that a certain type of failure event occurs frequently and the existing reinforcement measures have poor effects, then the reinforcement strategy adjustment mechanism is triggered. When adjusting the reinforcement strategy, the system comprehensively considers various means such as hardware upgrade, software optimization, and algorithm improvement. By dynamically adjusting the reinforcement strategy to cope with the changing operating environment and failure risks.
[0065] This detection process and solution content realize the comprehensive monitoring, data processing, failure determination and implementation of solutions for the core hardware and software of the relay protection device through an automated program, effectively reducing the need for manual intervention, improving the fault response speed and solution efficiency, and providing a strong guarantee for the safe and stable operation of the relay protection device.
[0066] The present invention provides a method for detecting chip failures in a relay protection device. Aiming at the problems of incomplete and inaccurate chip detection, the chips of the relay protection device are first subjected to hardware detection and then software testing. Through the combination of hardware testing and software testing, comprehensive failure detection of the chips is carried out to evaluate the failure probability and failure type of the chips, improving the accuracy and comprehensiveness of chip failure detection in the relay protection device, and enhancing the safety and reliability of the relay protection device.
[0067] The present invention studies the failure mechanisms of core chips such as CPU and FPGA in the relay protection device, analyzes the hardware failure mechanism and the effectiveness of anti-error design in the existing protection software and hardware architecture, and evaluates the failure risk; according to the results of the failure mechanism research, a hardware reinforcement scheme for the relay protection device to prevent abnormal storage of a single core hardware is proposed, including measures such as circuit structure reinforcement and system-level reinforcement; research on risk control measures for protection programs or sampled data, and propose a software reinforcement scheme for the relay protection device to eliminate the influence of single-event upsets of CPU and FPGA and abnormal data in a single channel on the device; construct an autonomous and controllable failure verification system for the storage unit of the relay protection device, simulate single-core hardware failures and memory data anomalies, and test and evaluate the relay protection device to verify the effectiveness of the software and hardware integration reinforcement method. The system includes a failure mechanism research module, a hardware reinforcement module, a software reinforcement module, and a failure verification module, and the modules work together to achieve chip failure detection and reinforcement functions.
[0068] Through the software and hardware integration reinforcement technology, the present invention improves the ability of the storage unit of the core chips in the relay protection device to resist anomalies, reduces the device failure risk, and ensures the safe and stable operation of the power grid. At the same time, the failure verification system constructed by the present invention can simulate various chip failure scenarios, providing an effective means for the test and evaluation of the relay protection device.
[0069] Optionally, before step S104 in the method for detecting chip failures in the relay protection device provided by the embodiment of the present invention, steps S201-S204 are further included.
[0070] S201: Perform hardware failure testing on chips of the same type as the chips in the relay protection device to obtain first test data.
[0071] In some embodiments, the first test data includes hardware detection data when the chip is normal and hardware detection data when the chip is abnormal.
[0072] S202: Perform software failure testing on chips of the same type as the chips in the relay protection device to obtain second test data.
[0073] In some embodiments, the second test data includes software detection data when the chip is normal and software detection data when the chip is abnormal.
[0074] S203. Generate training samples based on the first test data and the second test data.
[0075] In some embodiments, the training samples take the hardware detection data and software detection data under different working conditions as inputs, and take the chip states under different working conditions as outputs.
[0076] S204. Based on the training samples, perform neural network training to obtain a preset test model.
[0077] As a possible implementation, step S204 can be specifically implemented as steps S2041 - S2043.
[0078] S2041. Based on the training samples, perform neural network training to obtain an initial model;
[0079] S2042. Use the initial model as the teacher model and a model with the same depth as the initial model as the student model to construct a teacher - student model;
[0080] S2043. Based on the training samples, perform knowledge distillation on the teacher - student model to obtain a preset test model.
[0081] In this way, the embodiments of the present invention can train the test data of the chip within the historical period through a neural network model, form a preset test model, and perform failure judgment to improve the failure detection efficiency.
[0082] Optionally, for the chip failure detection method provided by the embodiments of the present invention, before step S102, steps S301 - S303 are further included.
[0083] S301. Perform hardware failure tests on multiple chips to obtain the first test data of the multiple chips.
[0084] S302. Based on the first test data of the multiple chips, determine the failure thresholds of the multiple chips.
[0085] Among them, the failure threshold of each chip is the hardware detection data when the chip fails during the hardware failure test.
[0086] S303. Based on the failure thresholds of the multiple chips, determine a preset threshold.
[0087] In this way, the embodiments of the present invention can analyze the failure thresholds of multiple chips to determine a preset threshold, improve the accuracy of failure threshold judgment, and improve the accuracy of failure detection.
[0088] Exemplarily, embodiments of the present invention can conduct a detailed failure mechanism analysis on core chips such as CPUs and FPGAs. This includes experimental testing of the chip's performance under different working environments (such as high temperature, high humidity, electromagnetic interference, etc.), as well as a detailed study of the chip's internal structure and manufacturing process.
[0089] Taking the CPU as an example, by simulating a high-temperature environment, it is found that the performance of a certain CPU will significantly decline when the temperature exceeds 85°C, and even data errors may occur. At the same time, by observing the internal structure of the chip under a microscope, it is found that some batches of CPUs have minor manufacturing defects, which may be potential causes of failure.
[0090] When conducting in-depth failure mechanism research, the core objective of the present invention is to comprehensively and accurately reveal the failure causes and processes of key chips such as CPUs and FPGAs in relay protection devices. This process not only involves performance analysis of the chips under various complex working environments, but also includes in-depth research on the internal microstructure, material properties, and manufacturing process of the chips.
[0091] First, the present invention will conduct environmental adaptability tests on the target chips. This includes, but is not limited to, simulation tests of extreme environments such as high temperature, low temperature, high humidity, and electromagnetic interference. Through these tests, the performance of the chips under different working environments, such as operation speed, power consumption, and data stability, can be observed, so as to discover potential failure modes and triggering conditions.
[0092] Secondly, the present invention will use advanced microscopic observation and analysis techniques to conduct a detailed dissection of the chip's internal structure. This includes observing the surface and internal microstructure of the chip using equipment such as scanning electron microscopes (SEM) and transmission electron microscopes (TEM), and analyzing the material composition and crystal structure of the chip using equipment such as energy dispersive spectrometers (EDS) and X-ray diffractometers (XRD). Through these analyses, the microscopic mechanisms and physical processes of chip failure, such as material aging, interface failure, and microcrack propagation, can be revealed.
[0093] In addition, the present invention will also conduct in-depth research on the chip's manufacturing process. This includes analyzing the influence of process parameters, equipment accuracy, material purity, etc. in the chip manufacturing process on the chip quality. By comparing chips produced by different batches and different manufacturers, potential problems and improvement points in the manufacturing process can be found, thereby improving the reliability and durability of the chips.
[0094] Detailed example: Taking the CPU as an example, the present invention first conducted a series of environmental adaptability tests. In a simulated high-temperature environment (such as 85°C), it was found that the performance of a certain CPU significantly declined, specifically manifested as slower operation speed, increased power consumption, and rising data error rate. By disassembling and analyzing the chip, it was found that the transistor structure inside the CPU underwent minor deformation and failure under high-temperature conditions, resulting in hindered electron flow and performance degradation. Further, the present invention used a microscope to observe the internal structure of the CPU and found that the transistor arrangement in some areas was not neat, with manufacturing defects. By comparing CPUs from different batches and manufacturers, it was found that such manufacturing defects might be related to specific production processes and equipment precision. Therefore, to address this issue, the present invention proposed improvement measures such as optimizing the production process and improving equipment precision to reduce the failure risk of the CPU. In addition to the research on the failure mechanism of hardware, the present invention also focuses on risk control measures in software. For example, regarding the single-event upset problem that the CPU may encounter, the present invention studied how to detect and correct data errors through software algorithms to ensure the stable operation of the protection program. These software reinforcement measures, combined with the hardware reinforcement scheme, jointly constitute the comprehensive and effective relay protection device chip failure detection method and system proposed by the present invention.
[0095] Furthermore, to evaluate the chip failure risk, based on the research of the failure mechanism, the present invention evaluated the risk of chip failure. This includes analyzing the probability of chip failure, the degree of impact on the system, and the effectiveness of existing protection measures. For example: Based on the above high-temperature experiment and internal structure observation, it was evaluated that the risk of the CPU failing in a high-temperature environment was relatively high, and the existing heat dissipation measures were not sufficient to completely prevent failure. In addition, it was analyzed that once the CPU fails, it may cause the entire relay protection device to malfunction, and even trigger a power system fault. After deeply exploring the failure mechanism, the evaluation of the chip failure risk becomes a crucial step. This step not only involves predicting the possibility of chip failure but also includes evaluating the consequences of failure and considering the effectiveness of existing protection mechanisms. Through this step, the present invention can more comprehensively understand the potential threats of chip failure and provide a strong basis for formulating effective risk response strategies.
[0096] First of all, the present invention will use the results of the failure mechanism research, combined with the historical failure data and environmental conditions of the chip, and through statistical and probability methods, predict the probability of chip failure. For example, by comparing the chip failure data under different batches and different working environments, the key factors affecting the failure probability can be identified, and a failure probability prediction model can be established accordingly. At the same time, simulation software can also be used to conduct virtual tests on the chip, simulating various failure scenarios to verify the accuracy of the prediction model.
[0097] Secondly, the present invention will evaluate the impact degree of chip failure on the system. This includes the analysis of the role and function of the chip in the system, as well as the prediction of the possible consequences caused by the failure. For example, for a core chip such as a CPU, its failure may lead to the paralysis or malfunction of the entire system, thus triggering serious consequences. Therefore, special attention needs to be paid to the failure risk of such chips.
[0098] In addition, the present invention will also evaluate the effectiveness of existing protection measures. This includes analyzing the performance of existing software and hardware protection mechanisms in preventing chip failure, as well as evaluating the feasibility and reliability of these measures in practical applications. Through this step, the present invention can identify the deficiencies of existing protection measures and provide a basis for formulating targeted reinforcement solutions.
[0099] Detailed example: Taking a certain FPGA chip as an example, this chip plays an important role in the relay protection device and is responsible for processing a large amount of data and signals. Through the study of the failure mechanism, the present invention finds that this FPGA chip is prone to performance degradation and data errors in a high-temperature environment. Based on this discovery, the present invention uses a failure probability prediction model to predict the failure probability of this FPGA chip in a high-temperature environment. The results show that when the environmental temperature exceeds 70 °C, the failure probability of the FPGA chip will increase significantly. The present invention evaluates the impact degree of the failure of this FPGA chip on the system. By analyzing the function and role of this chip in the system, the present invention finds that its failure may lead to data processing delays or errors, thereby triggering misoperations or refusals of the protection device. Such consequences will pose a serious threat to the stable operation of the power system. Finally, the present invention evaluates the effectiveness of existing protection measures. The present invention finds that although the existing heat dissipation measures can reduce the working temperature of the FPGA chip to a certain extent, it is still difficult to completely prevent failure in extreme environments. In addition, there are certain limitations in the existing software protection mechanism in detecting and correcting data errors of the FPGA chip.
[0100] Based on the above evaluation results, the embodiments of the present invention propose targeted reinforcement solutions. In terms of hardware, the present invention strengthens the heat dissipation design and adopts more efficient heat dissipation materials and structures; in terms of software, the embodiments of the present invention optimize the error detection and correction algorithms and improve the processing ability of data errors of the FPGA chip. Through the implementation of these reinforcement measures, the embodiments of the present invention have successfully reduced the failure risk of the FPGA chip and improved the safety and reliability of the relay protection device.
[0101] Hardware and software reinforcement solutions. Hardware reinforcement: Based on the results of failure risk assessment, the present invention proposes targeted hardware reinforcement solutions. For example, for the problem of high-temperature failure of the CPU, a more efficient heat dissipation design can be adopted, such as adding heat sinks or fans; for manufacturing defect problems, chip suppliers with more stable quality can be selected or factory inspection can be strengthened. Software reinforcement: In addition to hardware reinforcement, the present invention also studies software reinforcement measures. For example, for the problem of single-event upsets of the CPU, the software algorithm can be optimized to make it have higher fault tolerance for single-event upsets; at the same time, data verification and error recovery mechanisms are added to ensure the accuracy and reliability of data. Based on the previous comprehensive assessment of chip failure risks, the present invention proposes comprehensive hardware and software reinforcement solutions for different types of failure risks. These solutions aim to enhance the robustness of the system from multiple levels, reduce the possibility of chip failure, and quickly restore the normal operation of the system when a failure occurs. Hardware reinforcement solutions mainly focus on improving the working environment of the chip, enhancing the stability of its physical structure, and improving the resistance to external interference. Heat dissipation optimization: For the situation where chip failure is easily caused by high-temperature environment, the present invention will adopt a more efficient heat dissipation design. For example, a larger area of heat sink can be added to the CPU, or a more efficient cooling fan can be used to ensure that the chip can maintain a lower temperature during operation. Material upgrade: Higher-quality, more heat-resistant, and more corrosion-resistant materials are selected for chip packaging and manufacturing to improve the durability and reliability of the chip. Strengthen testing and screening: Before the chip leaves the factory, strengthen the testing of its performance, stability, reliability, etc., and select chips with better performance and more stability for production. Software reinforcement solutions focus on improving the robustness of the system at the software level by optimizing software algorithms, adding error detection and recovery mechanisms, etc. Algorithm optimization: Optimize relevant software algorithms for specific chip failure problems. For example, for the problem of single-event upsets of the CPU, the present invention can design a more robust algorithm to make it have higher fault tolerance for such errors. Data verification and recovery: Add a verification mechanism during data transmission and storage to ensure the accuracy and integrity of data. At the same time, design a fast recovery mechanism to quickly repair or recover when data errors are detected. Redundant design: For critical functions or data, redundant design is adopted, that is, corresponding function modules or data are backed up. In this way, when a certain module or data has problems, the system can switch to the backup module or data to ensure the continuous operation of the system.
[0102] Detailed example: Taking an FPGA chip used in the aerospace field as an example, this chip needs to work stably under extreme environmental conditions. For this purpose, the present invention proposes the following software and hardware reinforcement solutions: In terms of hardware: The present invention selects high-quality, high-temperature resistant, and radiation-resistant materials for the packaging and manufacturing of the FPGA chip. A large-area heat sink is added to the FPGA chip, and an efficient cooling fan is adopted to ensure that the chip can maintain a relatively low operating temperature under extreme high-temperature environments. Before leaving the factory, the present invention conducts strict tests and screenings on the FPGA chip to ensure its stable and reliable performance. In terms of software: Aiming at the single-event upset problem that the FPGA chip may face, the present invention optimizes the relevant algorithms and logic designs to improve its fault tolerance for such errors. During the data transmission and storage process, the present invention adds a verification mechanism to ensure the accuracy and integrity of the data. At the same time, a fast recovery mechanism is designed to quickly repair or recover when data errors are detected. For key functional modules and data, the present invention adopts redundant design, that is, the corresponding functional modules and data are backed up. In this way, when a certain module or data has problems, the system can quickly switch to the backup module or data to ensure the continuous operation of the system. Through the implementation of the above software and hardware reinforcement solutions, the present invention has successfully improved the stability and reliability of the FPGA chip under extreme environments, providing a strong guarantee for applications in the aerospace field.
[0103] To verify the effectiveness of the reinforcement measures, the present invention constructs a failure verification system. This system can simulate various chip failure scenarios and comprehensively test the relay protection devices before and after reinforcement. For example: using the failure verification system, a failure scenario of the CPU at high temperature is simulated. First, the un-reinforced relay protection device is tested, and it is found that it indeed malfunctions at high temperature. Then, the reinforced device is tested, and it is found that it can work stably at high temperature without any malfunction. This result verifies the effectiveness of the reinforcement measures. Constructing the failure verification system is a key step in ensuring the effectiveness of the chip reinforcement measures. By simulating various chip failure scenarios, the present invention can comprehensively test the chip performance before and after reinforcement, so as to verify the actual effect of the reinforcement measures. Failure scenario simulation design: First, the present invention needs to design corresponding failure scenarios according to the known chip failure modes. These scenarios should cover various external factors that may cause chip failure, such as abnormal temperature, power fluctuation, electromagnetic interference, radiation effect, etc. Failure injection system construction: To simulate these failure scenarios, the present invention needs to construct a failure injection system. This system can simulate different failure conditions and apply them to the chips to be tested. Comparative testing before and after reinforcement: Using the failure injection system, the present invention tests the chips before and after reinforcement respectively. By comparing their performances under the same failure scenarios, the present invention can evaluate the effectiveness of the reinforcement measures. Data collection and analysis: During the testing process, the present invention needs to collect a large amount of data, including the performance parameters, error rate, response time, etc. of the chips under different failure scenarios. By analyzing these data, the present invention can deeply understand the impact of the reinforcement measures on the chip performance.
[0104] Detailed example: Taking the failure verification of the CPU in a high-temperature environment as an example, the present invention designs the following test plan: Design of failure scenarios: The present invention sets a high-temperature environment to simulate the extreme temperature conditions that the CPU may face in actual applications. Construction of a failure injection system: The present invention constructs an environmental chamber capable of precisely controlling the temperature and places the CPU to be tested therein. By adjusting the temperature of the environmental chamber, the present invention can simulate different high-temperature environments. Comparative tests before and after reinforcement: First, the present invention tests the un-reinforced CPU. During the process of gradually increasing the temperature of the environmental chamber, the present invention observes that the performance of the CPU begins to decline and finally malfunctions occur. Then, the present invention conducts the same test on the reinforced CPU. Even at higher temperatures, the reinforced CPU can still maintain a stable working state without malfunctions. Data collection and analysis: During the test, the present invention records the performance parameters of the CPU at different temperatures, including the operating speed, power consumption, error rate, etc. By comparing and analyzing these data, the present invention finds that the reinforcement measures significantly improve the stability and reliability of the CPU in a high-temperature environment. This detailed example demonstrates how to verify the effectiveness of chip reinforcement measures by constructing a failure verification system. By simulating real failure scenarios and conducting comparative tests, the present invention can more accurately evaluate the effect of the reinforcement measures and provide strong support for the optimization and improvement of the chip.
[0105] In summary, through in-depth research on the chip failure mechanism, assessment of failure risks, proposal of software and hardware reinforcement solutions, and construction of a failure verification system, the present invention provides a comprehensive and accurate method for chip failure detection of relay protection devices. This not only improves the safety and reliability of relay protection devices but also provides strong guarantee for the stable operation of the power system.
[0106] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0107] The following is an apparatus embodiment of the present invention. For the details not described in detail therein, reference may be made to the corresponding method embodiments above.
[0108] Figure 2 The structural schematic diagram of a chip failure detection apparatus for a relay protection device provided by an embodiment of the present invention is shown. The failure detection apparatus 400 includes a communication module 401 and a processing module 402.
[0109] The communication module 401 is used to obtain the hardware detection data of the chip in the relay protection device during the working process in the current period; the hardware detection data includes the working temperature, working voltage, and working current.
[0110] A processing module 402 is configured to perform hardware failure detection on the relay protection device based on the hardware detection data and a preset threshold, and determine the hardware test result of the chip in the relay protection device; if the hardware test result indicates chip failure, perform software failure testing on the chip in the relay protection device to obtain a software test result; determine the failure detection result of the chip in the relay protection device based on the hardware test result and the software test result; the failure detection result includes a failure probability and a failure type; determine a reinforcement plan for the chip in the relay protection device based on the failure detection result.
[0111] Figure 3 FIG. is a schematic structural diagram of an electronic device in a chip failure detection system for a relay protection device provided by an embodiment of the present invention. As Figure 3 shown, the electronic device 500 includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program 503, the steps in the above method embodiments are implemented, for example Figure 1 the steps S101 - S105 shown. Alternatively, when the processor 501 executes the computer program 503, the functions of each module / unit in the above device embodiments are implemented, for example, Figure 2 the functions of the communication module 401 and the processing module 402 shown.
[0112] Exemplarily, the computer program 503 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 503 in the electronic device 500. For example, the computer program 503 can be divided into Figure 2 the communication module 401 and the processing module 402 shown.
[0113] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A chip failure detection method for a relay protection device, characterized in that: include: Obtain hardware detection data of the chip in the relay protection device during the working process in the current period; The hardware detection data includes operating temperature, operating voltage and operating current; Based on the hardware detection data and a preset threshold, perform hardware failure detection on the relay protection device to determine a hardware test result of a chip in the relay protection device; If the hardware test result is a chip failure, a software failure test is performed on the chip in the relay protection device to obtain a software test result; Determining a failure detection result of a chip in the relay protection device based on the hardware test result and the software test result; The failure detection result includes failure probability and failure type; Based on the failure detection result, a reinforcement scheme for the chip in the relay protection device is determined.
2. The chip failure detection method for a relay protection device according to claim 1, characterized in that: The performing hardware failure detection on the relay protection device based on the hardware detection data and a preset threshold value to determine the hardware test result of the chip in the relay protection device includes: Based on the hardware detection data and the preset threshold, a threshold judgment is performed to determine whether the hardware detection data is within a normal range; If the hardware detection data is within a normal range, determining that the hardware test result is that the chip is normal; If the hardware detection data is not within a normal range, a signal integrity detection and a redundant circuit state detection are performed based on the hardware detection data to determine the signal integrity of the hardware detection data and the redundant circuit state of the chip in the relay protection device; If the hardware detection data has a signal anomaly, or the redundant circuit works abnormally, then the hardware test result is determined to be a chip failure; If the signal of the hardware detection data is normal and the redundant circuit works normally, it is determined that the hardware test result is that the chip is normal.
3. The chip failure detection method for a relay protection device according to claim 1, characterized in that: If the hardware test result is a chip failure, a software failure test is performed on the chip in the relay protection device to obtain a software test result, including: Performing a program execution path tracking test on the chip in the relay protection device and recording chip program execution parameters, wherein the execution parameters include the number of calls, return value and execution time; Performing sampling data verification on the chip in the relay protection device to determine the sampling data verification result; Performing a single-particle upset simulation test on the chip in the relay protection device to determine the response of the single-particle upset; The software test result is determined based on the chip program execution parameters, the sampling data verification result, and the response of the single event upset.
4. The chip failure detection method for a relay protection device according to claim 1, characterized in that: The determining, based on the hardware test result and the software test result, a failure detection result of the chip of the relay protection device comprises: Generate a test vector based on the hardware test results and the software test results; the test vector includes operating temperature, operating voltage, operating current, chip program execution parameters, sampling data verification results and single-particle upset response; The failure detection result is generated based on the test vector and a preset test model; the preset test model is obtained by testing a chip of the same type as the chip in the relay protection device.
5. The chip failure detection method for a relay protection device according to claim 4, characterized in that: Before determining the failure detection result of the chip of the relay protection device based on the hardware test result and the software test result, the method further includes: Performing a hardware failure test on a chip of the same type as the chip in the relay protection device to obtain first test data, wherein the first test data includes hardware detection data of a normal chip and hardware detection data of an abnormal chip; Performing a software failure test on a chip of the same type as the chip in the relay protection device to obtain second test data, wherein the second test data includes software detection data of a normal chip and software detection data of an abnormal chip; Generate training samples based on the first test data and the second test data, wherein the training samples take the hardware detection data and the software detection data of different working conditions as input and take the chip states of different working conditions as output; Based on the training samples, neural network training is performed to obtain the preset test model.
6. The chip failure detection method for a relay protection device according to claim 5, characterized in that: The performing of neural network training based on the training samples to obtain the preset test model includes: Based on the training samples, neural network training is performed to obtain an initial model; The initial model is used as a teacher model, and a model with the same depth as the initial model is used as a student model to construct a teacher-student model; Based on the training samples, knowledge distillation is performed on the teacher-student model to obtain the preset test model.
7. The chip failure detection method for a relay protection device according to claim 1, characterized in that: The step of determining a reinforcement scheme for the chip based on the failure detection result and reinforcing the chip includes: If the failure type is hardware failure and the failure probability is greater than a preset probability, the reinforcement solution is determined to be a hardware upgrade; If the failure type software fails and the failure probability is less than or equal to the preset probability, the reinforcement solution is determined to be algorithm optimization; If the failure type software fails and the failure probability is greater than the preset probability, the reinforcement plan is determined to be hardware upgrade and algorithm optimization.
8. The chip failure detection method for a relay protection device according to any one of claims 1 to 7, characterized in that: Before performing hardware failure detection on the relay protection device based on the hardware detection data and a preset threshold value and determining the hardware test result of the chip in the relay protection device, the method further includes: Performing hardware failure tests on multiple chips to obtain first test data of the multiple chips; Based on the first test data of the plurality of chips, determining the failure thresholds of the plurality of chips; wherein the failure threshold of each chip is hardware detection data of chip failure when the chip is subjected to hardware failure test; The preset threshold is determined based on the failure thresholds of the plurality of chips.
9. A chip failure detection device for a relay protection device, characterized in that: include: The communication module is specifically used to obtain the hardware detection data of the chip in the relay protection device during the working process in the current period; the hardware detection data includes the working temperature, working voltage and working current; A processing module, specifically used to perform hardware failure detection on the relay protection device based on the hardware detection data and a preset threshold value, and determine a hardware test result of a chip in the relay protection device; If the hardware test result is a chip failure, a software failure test is performed on the chip in the relay protection device to obtain a software test result; Determining a failure detection result of a chip in the relay protection device based on the hardware test result and the software test result; The failure detection result includes failure probability and failure type; Based on the failure detection result, a reinforcement scheme for the chip in the relay protection device is determined.
10. A chip failure detection system for a relay protection device, characterized in that: The chip failure detection system includes an electronic device, which includes a memory and a processor. The processor can call and execute a computer program stored in the memory, and implement the steps of the method described in any one of claims 1 to 8 when executing the computer program.