A Chip Adaptive Fault Handling System and Method
By designing an adaptive fault processing system in the chip, including configuration, signal acquisition, signal processing and output control units, the problem of existing chips relying on CPU to handle faults is solved, and chip-independent fault diagnosis and repair is realized, reducing the processing burden of the CPU and avoiding serious abnormalities in the server system.
Patent Information
- Application Number
- CN202510214448.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing chips need to rely on the CPU when handling faults, which leads to an increase in CPU processing burden, and cannot assist the chip in failing recovery when there is an abnormality in the CPU, which may cause serious abnormalities in the entire server system.
A chip adaptive fault processing system is designed, including a configuration unit, a signal acquisition unit, a signal processing unit and an output control unit. Through the connection and coordinated work between these units, the chip can perform fault diagnosis and fault repair without relying on the CPU.
It reduces the processing burden of the CPU, avoids server system failures caused by CPU exceptions, and realizes chip independent failure recovery capabilities.
Smart Images

Figure CN119718784B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip technology, and in particular, to a chip adaptive fault handling system and method. Background Art
[0002] With the advent of the AI era, AI servers have become increasingly common. In AI servers, the requirements for the transmission of various high-speed signals are getting higher and higher. Therefore, the demand for the number of chips in the server and the performance requirements for the chips are also getting higher and higher. With the wide application of chips, the need for technologies such as parameter statistics, fault diagnosis, and fault repair inside the chips is becoming increasingly urgent.
[0003] However, when dealing with faults, existing chips often rely on the CPU to assist in fault handling and fault recovery. For example, the retimer chip, which is widely used in servers, is currently mostly in its infancy, and there is no complete and perfect fault diagnosis and fault repair technology. When the retimer chip fails, it cannot actively and independently complete fault repair. Often, the CPU needs to generate a reset signal based on the interrupt signal when the chip fails, and then execute a reset on the faulty module in the chip. The existing processing method, on the one hand, increases the processing burden of the CPU, and on the other hand, if the CPU also has abnormalities, etc., it cannot assist the chip to achieve fault recovery, which may lead to a failure of the entire server system and further cause serious abnormalities. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a chip adaptive fault handling system and method to reduce the CPU processing burden and avoid serious abnormalities.
[0005] In a first aspect, the present invention provides a chip adaptive fault handling system, which is applied to a fault handling module in a chip including multiple modules. The system includes a configuration unit, a signal acquisition unit, a signal processing unit, and an output control unit. The configuration unit is respectively connected to the CPU in the server, the signal acquisition unit, and the signal processing unit. The signal acquisition unit is also connected to the signal processing unit, and the signal processing unit is also connected to the output control unit;
[0006] The configuration unit is used to configure mode information based on the information obtained from the CPU;
[0007] The signal acquisition unit is used to collect signals from a target module based on the mode information to obtain collected signals;
[0008] The signal processing unit is used to process the collected signals based on the mode information and obtain a corresponding reset signal when it is determined that an abnormality is triggered;
[0009] The output control unit is configured to output the reset signal to the target module to reset the target module.
[0010] In an alternative embodiment, the signal processing unit includes a trigger control subunit, an isolation processing subunit, an adaptive processing subunit, and an output control subunit;
[0011] The trigger control subunit is configured to determine whether the acquired signal triggers an anomaly based on the mode information;
[0012] The isolation processing subunit is configured to maintain or filter the anomaly based on the signal type of the acquired signal when the trigger control subunit determines that an anomaly is triggered;
[0013] The adaptive processing subunit is configured to determine whether to output a reset signal by combining the determination result of the trigger control subunit and the determination result of the isolation processing subunit;
[0014] The output control subunit is configured to output the reset signal to the output control unit when the adaptive processing subunit determines to output the reset signal.
[0015] In an alternative embodiment, the mode information includes an anomaly trigger mode, the anomaly trigger mode includes a timestamp trigger mode, and the timestamp trigger mode includes a trigger signal type, a trigger method, and a waiting duration;
[0016] The signal processing unit is configured to trigger an anomaly after the waiting duration is satisfied when it is determined that the acquired signal is the trigger signal type in the timestamp trigger mode and the acquired signal conforms to the trigger method in the timestamp trigger mode.
[0017] In an alternative embodiment, the system further includes a storage unit;
[0018] The signal acquisition unit is further configured to store the acquired signal in the storage unit;
[0019] The signal processing unit is further configured to, when it is determined that the acquired signal triggers the timestamp trigger mode, extract the signal segment within the waiting duration of the acquired signal from the storage unit and output it to the output control unit;
[0020] The output control unit is further configured to output the signal segment to other units of the chip for signal analysis.
[0021] Second aspect, the present invention provides a chip self-adaptive fault handling method, which is applied to the chip self-adaptive fault handling system described in any one of the above. The system includes a configuration unit, a signal acquisition unit, a signal processing unit, and an output control unit. The chip self-adaptive fault handling method includes:
[0022] The configuration unit configures mode information based on the information obtained from the CPU;
[0023] The signal acquisition unit acquires signals from the target module based on the mode information to obtain acquired signals;
[0024] The signal processing unit processes the acquired signals based on the mode information and obtains corresponding reset signals when it is determined that an exception is triggered;
[0025] The output control unit outputs the reset signals to the target module to achieve the reset of the target module.
[0026] The chip self-adaptive fault handling system and method provided by the embodiments of the present invention include a configuration unit, a signal acquisition unit, a signal processing unit, and an output control unit. The configuration unit is used to configure mode information based on the information obtained from the CPU. The signal acquisition unit is used to obtain the acquired signals of the target module based on the mode information. The signal processing unit is used to process the acquired signals based on the mode information and obtain reset signals when it is determined that an exception is triggered. The output control unit is used to output the reset signals to the target module to achieve the reset of the target module. In this solution, there is no need to rely on the CPU to execute the reset, reducing its processing burden, and the fault recovery can be successfully executed even when the CPU has an exception, avoiding the problem of serious exceptions in the entire server system. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a schematic structural diagram of the chip provided by the embodiment of the present invention;
[0029] Figure 2 It is a structural block diagram of the chip self-adaptive fault handling system provided by the embodiment of the present invention;
[0030] Figure 3 It is a structural block diagram of the configuration unit provided by the embodiment of the present invention;
[0031] Figure 4 Block diagram of the signal acquisition unit provided by an embodiment of the present invention;
[0032] Figure 5 Block diagram of the signal processing unit provided by an embodiment of the present invention;
[0033] Figure 6 Block diagram of the output control unit provided by an embodiment of the present invention;
[0034] Figure 7 Flowchart of the chip self - adaptive fault handling method provided by an embodiment of the present invention.
[0035] Icons: 10 - Fault handling module; 11 - Configuration unit; 111 - Interface sub - unit; 112 - Acquisition configuration sub - unit; 113 - Abnormality configuration sub - unit; 114 - Other configuration sub - unit; 12 - Signal acquisition unit; 121 - Signal acquisition selection sub - unit; 122 - Read - write control sub - unit; 123 - Signal processing control sub - unit; 13 - Signal processing unit; 131 - Trigger control sub - unit; 132 - Isolation processing sub - unit; 133 - Self - adaptive processing sub - unit; 134 - Output control sub - unit; 14 - Output control unit; 141 - Interface processing sub - unit; 142 - Reset output interface sub - unit; 143 - Interrupt output interface sub - unit; 144 - Backup output interface sub - unit; 15 - Storage unit. Detailed implementation manners
[0036] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described technical solutions are part of the technical solutions of the present invention, rather than all of the technical solutions. The components of the technical solutions of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0037] Therefore, the following detailed description of the technical solutions of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected technical solutions of the present invention. All other technical solutions obtained by those of ordinary skill in the art based on the technical solutions in the present invention without making creative efforts fall within the scope of protection of the present invention.
[0038] As described in the background art, when the chips in existing servers perform fault recovery, they often rely on the CPUs in the servers to assist in implementation. For example, for the retimer chips widely used in servers, most of the current retimer chips are in their infancy and do not have complete and perfect fault diagnosis and fault repair technologies. When a retimer chip fails, it cannot actively and independently complete fault repair, resulting in possible problems such as the failure of the entire server system.
[0039] To solve the above technical problems, please refer to Figure 1 , the present invention provides a chip adaptive fault processing system. The adaptive fault processing system is applied to a fault processing module 10 included in a chip in a server, and the chip can be, but is not limited to, a retimer chip. The chip includes multiple modules, including the fault processing module 10 and other modules such as a data processing module and a control module. The fault processing module 10 can be used to perform fault diagnosis and fault recovery processing on other modules in the chip.
[0040] Please refer to Figure 2 , the fault processing module 10 includes a configuration unit 11, a signal acquisition unit 12, a signal processing unit 13, and an output control unit 14. Among them, the configuration unit 11 is connected to the CPU, the signal acquisition unit 12, and the signal processing unit 13 in the server. In addition, the signal acquisition unit 12 is also connected to the signal processing unit 13, and the signal processing unit 13 is also connected to the output control unit 14.
[0041] In this embodiment, each unit in the fault processing module 10 can be connected through metal wiring, polysilicon wiring, etc.
[0042] The functions of each functional unit in the chip adaptive fault processing system provided in this embodiment are realized as follows:
[0043] The configuration unit 11 is used to configure mode information based on the information obtained from the CPU.
[0044] The signal acquisition unit 12 is used to collect signals from the target module based on the mode information to obtain the collected signals.
[0045] The signal processing unit 13 is used to process the collected signals based on the mode information and obtain the corresponding reset signal when it is determined that an exception is triggered;
[0046] The output control unit 14 is used to output the reset signal to the target module to reset the target module.
[0047] In this embodiment, both the signal acquisition unit 12 and the signal processing unit 13 implement signal acquisition and signal processing respectively based on the mode information pre-configured by the CPU. Among them, the target module is any other module in the chip except the fault processing module 10, such as Figure 1 the data processing module, the control module, etc. in
[0048] During implementation, the acquisition of the required signals can be achieved based on the relevant configurations under actual requirements, and the processing of signals can be achieved based on the configurations related to the processing under actual requirements, so as to meet the relevant requirements of acquisition and processing. Moreover, flexible information configuration can be supported, thereby meeting diverse acquisition and processing methods.
[0049] In addition, in this embodiment, since the signal processing unit 13 can process the acquired signals based on the pre-configured mode information and obtain a reset signal when it is determined that an abnormality is triggered, and the reset signal is output to the target module through the output control unit 14 to achieve reset. In the existing method, it is necessary to rely on the CPU to reset the module to be reset, which has the problem of a large CPU processing burden. Moreover, in the existing method, when the CPU fails, other modules cannot be reset, resulting in the inability to achieve fault recovery. Therefore, this solution can avoid the problem of a large CPU processing burden existing in the existing method, and can avoid the problem that when it is necessary to rely on the CPU for reset, the CPU cannot be reset when it fails, resulting in a serious abnormality in the entire server system.
[0050] As an optional implementation method, as shown in Figure 3 , the configuration unit 11 includes an interface sub-unit 111, an acquisition configuration sub-unit 112, and an exception configuration sub-unit 113. Among them, the interface sub-unit 111 is respectively connected to the acquisition configuration sub-unit 112 and the exception configuration sub-unit 113.
[0051] The interface sub-unit 111 can be an Advanced High-performance Bus (AHB) sub-unit. The interface sub-unit 111 can be connected to the CPU in the server to receive the relevant information configured by the CPU, and send the received information to the acquisition configuration sub-unit 112 and the exception configuration sub-unit 113. This information includes the range information of signal acquisition, the definition of signal abnormality, the information related to signal storage, the information related to signal reporting, etc.
[0052] The acquisition configuration subunit 112 can be used to configure the acquisition mode based on the obtained information. Based on this, in this embodiment, the mode information includes multiple acquisition modes, and each acquisition mode corresponds to an acquisition port and a signal type of the chip. And each signal type under each acquisition port corresponds to each module in the chip. Thus, when collecting a signal of a certain signal type for a certain acquisition port, the target module corresponding to the collected signal can be determined.
[0053] The acquisition mode configured by the acquisition configuration subunit 112 can be provided to the signal processing unit 13 to perform operations related to signal acquisition.
[0054] It can be seen that each acquisition mode includes multiple types of information, such as acquisition ports, signal types, etc. And the acquisition port may also include multiple port information, and the signal type may also include multiple types of information. That is, the amount of information data included in each acquisition mode is huge.
[0055] For the sake of the speed and smoothness of processing during the implementation of fault handling, the configuration unit 11, specifically the acquisition configuration subunit 112 in the configuration unit 11, can also be used to uniformly configure the port information of the acquisition ports corresponding to each acquisition mode and the signal information corresponding to the signal types as a mode template.
[0056] Thus, through the configured mode template, when performing signal acquisition processing, the relevant mode template can be directly called, and signal acquisition can be performed based on the information in the mode template, without having to extract multiple information and then combine the multiple information, and then perform signal acquisition based on the combined information. The unified configured mode template can greatly improve the speed and smoothness of the signal acquisition process.
[0057] In this embodiment, the composition form of the configured acquisition mode can be: mode number, acquisition port, signal type. Exemplarily, the configured acquisition modes can include the following several:
[0058] Mode 0: Collect signals related to rx of port a of the chip;
[0059] Mode 1: Collect signals related to rx of port b of the chip;
[0060] Mode 2: Collect signals related to EQ of port x (the port configured for collecting EQ signals) of the chip;
[0061] Mode 3: Collect signals related to deskew of port a of the chip;
[0062] Mode 4: Collect signals related to deskew of port b of the chip;
[0063] Mode 5: Collect other related signals of other ports of the chip.
[0064] In addition, the anomaly configuration subunit 113 can be used to configure the anomaly trigger mode based on the obtained information. Based on this, in this embodiment, the mode information further includes multiple anomaly trigger modes, and each anomaly trigger mode includes a trigger method and a trigger signal type.
[0065] The anomaly trigger mode configured by the anomaly configuration subunit 113 can be provided to the signal processing unit 13 to perform signal processing and judgment operations such as whether to trigger an anomaly.
[0066] Each anomaly trigger mode includes a trigger method and a trigger signal type. Similarly, in order to improve the speed and smoothness of the signal processing process, the anomaly configuration subunit 113 can configure the trigger method and the trigger signal type in each anomaly trigger mode as a mode template. In this way, when performing signal processing, the mode templates corresponding to each anomaly trigger mode can be directly retrieved and anomaly judgment can be performed, which can greatly improve the processing efficiency.
[0067] In this embodiment, the trigger methods in each anomaly trigger mode include trigger events, the number of trigger events, and combination methods. Among them, the trigger events include rising edge events, falling edge events, and preset level width events, and the combination methods include multiple methods respectively composed of at least one trigger event.
[0068] That is to say, the trigger event can be, for example, to trigger an anomaly when a preset number of rising edge events are detected, or to trigger an anomaly when a preset number of falling edge events are detected, or to trigger an anomaly when a preset number of preset level widths are detected, or to trigger an anomaly when at least two combined events among the rising edge event, falling edge event, and preset level width event of a certain signal are detected, or to trigger an anomaly when two certain signals satisfy the set combination method of the rising edge event, falling edge event, and preset level width event, and so on.
[0069] In addition, in order to obtain the complete anomaly information of the anomaly acquisition signal, the anomaly trigger mode includes a timestamp trigger mode, and the timestamp trigger mode includes a trigger signal type, a trigger method, and a waiting duration. That is, the timestamp trigger mode adds a waiting duration compared with the above various anomaly trigger modes, and the meaning represented by the waiting duration is that the signal segment of the acquisition signal within the waiting duration can be acquired. In this way, in the scenario of performing chip testing, the complete anomaly information of the acquisition signal in the case of an anomaly can be obtained, and then used for detailed analysis and other processing, which is convenient for analyzing the cause of the anomaly and determining the corresponding solution for the anomaly.
[0070] Based on this, in this embodiment, exemplarily, the multiple anomaly trigger modes configured can be as follows:
[0071] Trigger mode 0: Rising edge trigger, number of triggers, trigger signal type;
[0072] Trigger mode 1: Falling edge trigger, number of triggers, trigger signal type;
[0073] Trigger mode 2: Preset level width trigger, number of triggers, trigger signal type;
[0074] Trigger mode 3: Mixed trigger (mix of rising edge, falling edge, preset level width, etc. of a single signal), number of triggers, trigger signal type;
[0075] Trigger mode 4: Pattern trigger (mix of rising edge, falling edge, preset level width, etc. of two or more signals), number of triggers, trigger signal type;
[0076] Trigger mode 5: Timestamp trigger, number of triggers, trigger signal type, waiting duration.
[0077] In this embodiment, in addition to the above-mentioned acquisition configuration sub-unit 112 and exception configuration sub-unit 113, the configuration unit 11 further includes other configuration sub-units 114, and the other configuration sub-units 114 are used for other configurations of relevant information such as signal backup storage and signal reporting.
[0078] Based on this, in this embodiment, the configuration information further includes storage method information and off-chip storage address. Among them, the storage method information is used to provide to the signal acquisition unit 12 to store the acquired signal into the on-chip storage unit 15, and the off-chip storage address is used to provide to the output control unit 14 to store the signal to be backed up and stored into the off-chip storage unit.
[0079] In this embodiment, when the fault processing module 10 determines that the target module is abnormal, it can also generate an interrupt signal and report it to the CPU. After receiving the interrupt signal, the CPU can configure a reset signal based on the interrupt signal, and then the CPU can reset the target module based on the reset signal.
[0080] Therefore, the configuration information may further include signal reporting information, and the signal reporting information includes information such as the type of interrupt signal to be reported to the CPU and the interrupt number carried during reporting.
[0081] The signal reporting information is mainly used to provide to the output control unit 14 to perform the signal reporting operation.
[0082] In a possible implementation manner, please refer to Figure 4, in this embodiment, the signal acquisition unit 12 includes a signal acquisition selection subunit 121, a read / write control subunit 122, and a signal processing control subunit 123. Among them, the signal acquisition selection subunit 121 is connected to the read / write control subunit 122 and the signal processing control subunit 123.
[0083] The signal acquisition selection subunit 121 is used to acquire signals based on the acquisition mode obtained from the configuration unit 11. As can be seen from the above, the mode information includes multiple acquisition modes, and each acquisition mode corresponds to an acquisition port and a signal type of the chip.
[0084] Based on this, the signal acquisition unit 12 (specifically, the signal acquisition selection subunit 121) is used to determine the target acquisition port and the target signal type based on the acquisition mode in the mode information, and acquire the signal corresponding to the target signal type based on the target acquisition port to obtain the acquired signal.
[0085] The ports and signal types indicated by the acquisition mode include port a and rx signal in mode 0, port b and rx signal in mode 1, EQ signal of port x (based on configuration) in mode 2, port a and deskew signal in mode 3, port b and deskew signal in mode 4, and other ports and other signals in mode 5.
[0086] Based on this, the signal acquisition selection subunit 121 can select the signal acquisition source based on the specific port information and signal type information in the configuration information and perform signal acquisition.
[0087] Refer to in combination Figure 2 , in this embodiment, the fault processing module 10 further includes a storage unit 15, and the storage unit 15 can be an SRAM unit. As can be seen from the above, the mode information further includes storage method information, and the signal acquisition unit 12 (specifically, the read / write control subunit 122) can be used to store the acquired signal in the storage unit 15 based on the storage method information.
[0088] Among them, the storage method information may include, for example, the processing method and storage method of the acquired signal. Among them, the processing method includes, for example, the packing method and filtering method of the acquired signal. The storage method includes, for example, the byte order method and the byte information corresponding to each storage bit in the storage unit 15.
[0089] Byte order refers to the byte arrangement order of multi-byte data in storage unit 15, mainly including big-endian mode and little-endian mode. In big-endian mode, the highest-order byte of a data is stored at the lowest memory address, while the lowest-order byte is stored at the highest memory address. That is, the byte order of the data is consistent with the growth pattern of the memory address. In little-endian mode, the lowest-order byte of a data is stored at the lowest memory address, while the highest-order byte is stored at the highest memory address. Contrary to big-endian mode, the byte order of the data is opposite to the growth direction of the memory address.
[0090] Different computer system architectures may adopt different byte orders. Therefore, the specific byte order can be determined based on the configuration.
[0091] In addition, the signal processing control subunit 123 is mainly used to connect with the signal processing unit 13, and read relevant signals from the storage unit 15 based on the indication information of the signal processing unit 13 for other relevant analysis and processing.
[0092] In this embodiment, the signal processing unit 13 is the core unit of the exception handling system. As Figure 5 shown, the signal processing unit 13 includes a trigger control subunit 131, an isolation processing subunit 132, an adaptive processing subunit 133, and an output control subunit 134. Among them, the trigger control subunit 131 is respectively connected to the isolation processing subunit 132, the adaptive processing subunit 133, and the output control subunit 134. In addition, the adaptive processing subunit 133 is also connected to the isolation processing subunit 132 and the output control subunit 134.
[0093] The trigger control subunit 131 is used to judge whether the collected signal triggers an exception based on the mode information. The isolation processing subunit 132 is used to judge whether to maintain or filter the exception based on the signal type of the collected signal when the trigger control subunit 131 determines that an exception is triggered. The adaptive processing subunit 133 is used to judge whether to output a reset signal by combining the judgment result of the trigger control subunit 131 and the judgment result of the isolation processing subunit 132. The output control subunit 134 is used to output the reset signal to the output control unit 14 when the adaptive processing subunit 133 determines to output the reset signal.
[0094] Specifically, the trigger control subunit 131 is connected to the configuration unit 11 and the signal acquisition unit 12, and is used to receive the collected signal sent by the signal acquisition unit 12, obtain the mode information configured by the configuration unit 11, and judge whether the collected signal triggers an exception based on the exception trigger mode included in the mode information.
[0095] Based on the above description of the anomaly trigger mode, the anomaly trigger mode includes the trigger method and the trigger signal type. Therefore, the signal processing unit 13 (specifically, the trigger control subunit 131) can be used to determine the target trigger signal type corresponding to the acquired signal, and judge whether the acquired signal triggers an anomaly based on the target trigger method corresponding to the target trigger signal type.
[0096] For example, in the configured anomaly trigger mode, the anomaly trigger for signal a is to trigger an anomaly after detecting 3 rising edge events of signal a. When the acquired signal is signal a, that is, the target trigger signal type corresponding to the acquired signal is signal a, the corresponding target trigger method can be obtained as detecting 3 rising edge events. Therefore, if 3 rising edge events of the acquired signal are detected, it can be determined that an anomaly is triggered.
[0097] As can be seen from the above, the configuration unit 11 can configure multiple anomaly trigger modes, including trigger mode 0 - trigger mode 5, which are rising edge trigger, falling edge trigger, preset level width trigger, mixed trigger, Pattern trigger, and timestamp trigger respectively.
[0098] The composition form of each anomaly trigger mode can be unified as: trigger event, number of trigger events, combination method, trigger signal type.
[0099] Based on this, when the signal processing unit 13 (specifically, the trigger control subunit 131) performs anomaly judgment, it can implement anomaly judgment in one or more of the following ways based on the configured anomaly trigger mode:
[0100] When the anomaly trigger mode is rising edge trigger, the trigger event is a rising edge event, the number of trigger events is A times, the combination method is a single signal (i.e., analyze each signal separately), and the trigger signal type is signal a. If the trigger control subunit 131 detects that the acquired signal is signal a, then if it detects that signal a reaches A rising edge events, it can be determined that an anomaly is triggered.
[0101] When the anomaly trigger mode is falling edge trigger, the trigger event is a falling edge event, the number of trigger events is B times, the combination method is a single signal, and the trigger signal type is signal b. If the trigger control subunit 131 detects that the acquired signal is signal b, then if it detects that signal b reaches B falling edge events, it can be determined that an anomaly is triggered.
[0102] When the anomaly trigger mode is preset level width trigger, the trigger event is a preset level width event, the number of trigger events is C times, the combination method is a single signal, and the trigger signal type is signal c. If the trigger control subunit 131 detects that the acquired signal is signal c, then if it detects that signal c reaches C preset level width events, it can be determined that an anomaly is triggered.
[0103] When the abnormal trigger mode is hybrid trigger, the trigger event can be any one or combination of a falling edge event, a rising edge event, and a preset level width event based on the configuration. The number of trigger events is D times, the combination method is a single signal, and the trigger signal type is signal d. If the trigger control subunit 131 detects that the acquisition signal is signal d, and if it detects any one or combination of the configured trigger events that signal d reaches D times, an abnormal trigger can be determined.
[0104] When the abnormal trigger mode is Pattern trigger, the trigger event can be any one or combination of a falling edge event, a rising edge event, and a preset level width event (such as a rising edge event and a falling edge event). The number of trigger events is E times and F times (exemplary), the combination method is a combination of two signals (exemplary), and the trigger signal types are signal e and signal f. If the trigger control subunit 131 detects that the acquisition signals are signal e and signal f, and if it detects that after signal e reaches the rising edge event E times, signal f appears F times at the falling edge event, an abnormal trigger can be determined.
[0105] In addition, for the timestamp trigger mode, the control processing unit (specifically, the trigger control subunit 131) is used to trigger an abnormality after waiting for a waiting duration when it is determined that the acquisition signal is the trigger signal type in the timestamp trigger mode and the acquisition signal conforms to the trigger method in the timestamp trigger mode.
[0106] That is to say, compared with the above several abnormal trigger modes, the timestamp trigger mode adds a waiting duration. The purpose of adding this waiting duration is to collect the complete abnormal signal information for certain types of signals, so as to be used for other analyses, such as fault cause analysis and fault location, etc., and thus contribute to the thorough analysis and recovery of faults.
[0107] Based on this, when the abnormal trigger mode is the timestamp trigger mode, the trigger event is any one or combination of a falling edge event, a rising edge event, and a preset level width event. The number of trigger events is G times, the combination method is a single signal, the trigger signal type is signal g, and the waiting duration is t milliseconds. If the trigger control subunit 131 detects that the acquisition signal is signal g, after detecting that signal g reaches the configured trigger event G times, and after waiting for t milliseconds, an abnormal trigger can be determined.
[0108] As can be seen from the above, after the signal acquisition unit 12 obtains the acquisition signal, it stores the acquisition signal in the storage unit 15. The signal processing unit 13 can also be used to extract the signal segment within the waiting duration of the acquisition signal from the storage unit 15 and output it to the output control unit 14 when it is determined that the acquisition signal triggers the timestamp trigger mode. The output control unit 14 can also be used to output the signal segment to other modules of the chip for signal analysis.
[0109] Considering that after an abnormal trigger of the acquired signal, the acquisition of the acquired signal may be interrupted, thus making it impossible to obtain the complete abnormal signal information of the acquired signal. Therefore, in this embodiment, in the timestamp trigger mode, when it is determined that the acquired signal is abnormal, the abnormal trigger is performed after a waiting duration. In this way, the signal segment within the waiting duration of the acquired signal after the abnormal determination can be obtained, which can be provided to other modules for analyzing the signal segment to analyze the cause of the fault, locate the fault, find a solution to the fault, etc.
[0110] The above is the detection and judgment process of whether the acquired signal triggers an abnormality by the trigger control subunit 131 in the signal processing unit 13 based on the abnormal trigger mode. Considering that during the operation of the chip, frequent abnormal triggers may occur based on the detection and judgment of the trigger control subunit 131, and then interrupt signals are frequently uploaded to the CPU or other modules are frequently reset, etc., which may cause the CPU load to be too heavy and the processing flow to be frequently interrupted.
[0111] Therefore, in this embodiment, the isolation processing subunit 132 is used to alleviate this problem. The isolation processing subunit 132 can determine whether the abnormality needs to be filtered based on the signal type of the acquired signal when the trigger control subunit 131 determines that an abnormal trigger has occurred. The filtering here can mean eliminating the abnormality or downgrading the abnormality. For example, if the originally triggered abnormality is to upload an interrupt signal to the CPU, it can be downgraded to only record the abnormality without uploading an interrupt signal to the CPU.
[0112] Among them, when determining whether to filter the abnormality based on the signal type of the acquired signal, the signal type of the acquired signal can be compared with the pre-configured signal type. If the signal type of the acquired signal can match the pre-configured signal type, it can be determined that the abnormality can be filtered. If the signal type of the acquired signal cannot match the pre-configured signal type, it is determined to maintain the abnormality.
[0113] Among them, the pre-configured signal types can be some signal types that have little impact on the system operation and do not cause serious errors, such as CRC signals, etc.
[0114] In this way, for some signals that have little impact on the system operation, even if the signal triggers an abnormality, after the judgment and processing by the isolation processing subunit 132, the abnormality can be filtered, thus avoiding frequent abnormal triggers of such signals, which may lead to an overly heavy CPU processing load or frequent interruption of the processing flow.
[0115] On this basis, the adaptive processing subunit 133 in the signal processing unit 13 can determine whether to output a reset signal to the output control subunit 134 in combination with the determination result of the trigger control subunit 131 and the determination result of the isolation processing subunit 132.
[0116] Specifically, if the trigger control subunit 131 triggers an abnormality in outputting a reset signal and the isolation processing subunit 132 determines to maintain the abnormality, the adaptive processing subunit 133 can determine to output a reset signal.
[0117] If the trigger control subunit 131 triggers an abnormality in outputting a reset signal and the isolation processing subunit 132 determines to filter out the abnormality, the adaptive processing subunit 133 can determine that there is no need to output a reset signal.
[0118] Similarly, if the trigger control subunit 131 triggers an abnormality in outputting an interrupt signal and the isolation processing subunit 132 determines to maintain the abnormality, the adaptive processing subunit 133 can determine to output the interrupt signal.
[0119] If the trigger control subunit 131 triggers an abnormality in outputting an interrupt signal and the isolation processing subunit 132 determines to filter out the abnormality, the adaptive processing subunit 133 can determine that there is no need to output the interrupt signal.
[0120] In this embodiment, by determining the final output through the adaptive processing subunit 133 in combination with the determination results of the trigger control subunit 131 and the isolation processing subunit 132, the reset signal and the interrupt signal can be autonomously determined. On the one hand, it can ensure the accurate processing of important signals and avoid further errors caused by abnormal important signals. On the other hand, it can reduce unnecessary processing burdens and also achieve the purpose of accelerating the fault processing efficiency to a certain extent.
[0121] On the basis of the above, in this embodiment, the output control unit 14 can output the reset signal to the target module to achieve the reset of the target module. The output control unit 14 can also output the interrupt signal to the CPU to perform an interrupt operation through the CPU.
[0122] Please refer to Figure 6 , in this embodiment, the output control unit 14 includes an interface processing subunit 141, a reset output interface subunit 142, an interrupt output interface subunit 143, and a backup output interface subunit 144. Among them, the interface processing subunit 141 is respectively connected to the reset output interface subunit 142, the interrupt output interface subunit 143, and the backup output interface subunit 144.
[0123] The interface processing subunit 141 is configured to determine to send the signal to the reset output interface subunit 142, the interrupt output interface subunit 143, or the backup output interface subunit 144 after receiving the signal.
[0124] For example, if the interface processing subunit 141 receives a reset signal, it determines to send the reset signal to the reset output interface subunit 142, so as to send the reset signal to the target module through the reset output interface subunit 142 to reset the target module.
[0125] In this embodiment, the reset output interface subunit 142 can directly output the reset signal to the target module to achieve the reset of the target module. This process makes autonomous decisions without the participation of the CPU, reducing the load pressure on the CPU and realizing a faster and more flexible fault handling method based on the chip.
[0126] If the interface processing subunit 141 receives an interrupt signal, it determines to send the interrupt signal to the interrupt output interface subunit 143, so as to send the interrupt signal to the CPU through the interrupt output interface subunit 143 to perform an interrupt operation on the target module through the CPU.
[0127] In addition, the output control unit 14 can also back up the signals that need to be backed up to the off-chip Flash, that is, to the non-volatile storage medium outside the chip.
[0128] In this embodiment, the mode information pre-configured by the configuration unit 11 further includes an off-chip storage address, which mainly refers to the starting address where the off-chip Flash records information.
[0129] Based on this, the output control unit 14 can be used to store the signals that need to be backed up into the off-chip Flash based on the off-chip storage address. For example, the signal segment within the waiting duration extracted from the above timestamp abnormal mode is stored in the off-chip Flash.
[0130] Specifically, after the interface processing subunit 141 in the output control unit 14 obtains the signals that need to be backed up, it determines to send the signals that need to be backed up to the backup output interface subunit 144, so as to output the signals that need to be backed up to the off-chip Flash through the backup output interface subunit 144 to achieve backup. In this way, the signals that need to be backed up are saved in the off-chip Flash, and the recovery of the abnormal scene can be realized subsequently.
[0131] The chip adaptive fault handling system provided in this embodiment can flexibly collect the signals of other modules inside the chip according to the configuration of the CPU, so as to facilitate the fault monitoring and recovery of debugging simulation, operation monitoring, etc. at different stages of the chip.
[0132] This solution supports multiple abnormal triggering modes. Among them, in the timestamp triggering mode, when it is determined that the acquisition signal is abnormal, the abnormality is triggered after the waiting duration has elapsed. In this way, the signal information within the waiting duration of the acquisition signal can be retained, facilitating further analysis of the abnormal information of the acquisition signal.
[0133] In addition, in this solution, an isolation processing subunit 132 is added to isolate the generated abnormality. On the one hand, it can ensure the timely detection of faults in important signals and the recovery of faults. On the other hand, it can avoid the frequent triggering of abnormalities in signals that have little impact on the system operation, thereby reducing the processing load of the CPU.
[0134] Furthermore, in this solution, in combination with the adaptive processing subunit 133, it can autonomously decide whether to generate a reset signal, an interrupt signal, etc., avoiding the time delay caused by the CPU's decision-making, thereby accelerating the efficiency of fault recovery.
[0135] Based on the same inventive concept, please refer to Figure 7 , an embodiment of the present invention further provides a chip adaptive fault handling method, which can be applied to the chip adaptive fault handling system in any of the implementation manners in the above embodiments. The method includes the following steps:
[0136] S11, the configuration unit 11 performs mode information configuration based on the information obtained from the CPU.
[0137] S12, the signal acquisition unit 12 acquires signals from the target module based on the mode information to obtain acquisition signals.
[0138] S13, the signal processing unit 13 processes the acquisition signals based on the mode information and obtains corresponding reset signals when it is determined that an abnormality is triggered.
[0139] S14, the output control unit 14 outputs the reset signal to the target module to achieve the reset of the target module.
[0140] In the chip adaptive fault handling method provided in this embodiment, the chip can flexibly implement signal acquisition based on the configuration of the CPU, perform abnormal detection processing on the acquisition signals, and then generate a reset signal when an abnormality is triggered. Based on the reset signal, the reset of the target module within the chip can be achieved. In this solution, it is not necessary to completely rely on the CPU to execute reset operations, etc. The chip can make autonomous decisions and execute reset operations, reducing the processing burden of the CPU, and can also perform fault recovery on the CPU when the CPU has an abnormality, avoiding the situation where the entire server system has a serious abnormality.
[0141] As a possible implementation, the signal processing unit 13 includes a trigger control subunit 131, an isolation processing subunit 132, an adaptive processing subunit 133, and an output control subunit 134. The signal processing unit 13 can process the acquired signal in the following manner:
[0142] The trigger control subunit 131 determines whether the acquired signal triggers an abnormality based on the mode information;
[0143] When the trigger control subunit 131 determines that an abnormality is triggered, the isolation processing subunit 132 determines whether to maintain or filter the abnormality based on the signal type of the acquired signal;
[0144] The adaptive processing subunit 133 determines whether to output a reset signal by combining the determination result of the trigger control subunit 131 and the determination result of the isolation processing subunit 132;
[0145] When the adaptive processing subunit 133 determines to output a reset signal, the output control subunit 134 outputs the reset signal to the output control unit 14.
[0146] As a possible implementation, the abnormal trigger mode includes a timestamp trigger mode. The timestamp trigger mode includes a trigger signal type, a trigger method, and a waiting duration. The signal processing unit 13 can determine whether to trigger an abnormality in the following manner:
[0147] When the signal processing unit 13 determines that the acquired signal is the trigger signal type in the timestamp trigger mode and the acquired signal conforms to the trigger method in the timestamp trigger mode, an abnormality is triggered after the waiting duration is satisfied.
[0148] As a possible implementation, the chip adaptive fault handling method may further include the following steps:
[0149] The signal acquisition unit 12 stores the acquired signal in the storage unit 15;
[0150] When the signal processing unit 13 determines that the acquired signal triggers the timestamp trigger mode, it extracts the signal segment within the waiting duration of the acquired signal from the storage unit 15 and outputs it to the output control unit 14;
[0151] The output control unit 14 outputs the signal segment to other modules of the chip for signal analysis.
[0152] As a possible implementation, the mode information includes storage method information and an off-chip storage address. The chip adaptive fault handling method further includes the following steps:
[0153] The signal acquisition unit 12 stores the acquired signal in the storage unit 15 based on the storage method information;
[0154] The output control unit 14 stores the signals to be backed up into the off-chip Flash based on the off-chip storage address.
[0155] As a possible implementation, the mode information includes multiple abnormal trigger modes, and each abnormal trigger mode includes a trigger method and a trigger signal type. The signal processing unit 13 determines whether the acquired signal triggers an abnormality in the following manner:
[0156] The signal processing unit 13 determines the target trigger signal type corresponding to the acquired signal, and determines whether the acquired signal triggers an abnormality based on the target trigger method corresponding to the target trigger signal type.
[0157] As a possible implementation, the trigger method includes a trigger event, the number of trigger events, and a combination method;
[0158] The trigger event includes a rising edge event, a falling edge event, and a preset level width event, and the combination method includes multiple methods composed of at least one trigger event.
[0159] As a possible implementation, the mode information includes multiple acquisition modes, and each acquisition mode corresponds to an acquisition port and a signal type of the chip. The signal acquisition unit 12 obtains the acquired signal in the following manner:
[0160] The signal acquisition unit 12 determines the target acquisition port and the target signal type based on the acquisition mode in the mode information, and acquires the signal corresponding to the target signal type based on the target acquisition port to obtain the acquired signal.
[0161] As a possible implementation, the configuration unit 11 is used to uniformly configure the port information of the acquisition port corresponding to each acquisition mode and the signal information corresponding to the signal type as a mode template.
[0162] The chip adaptive fault handling method provided in this embodiment is applied to the chip adaptive fault handling system in any one of the above implementation manners, and has the same, corresponding or similar technical features and beneficial effects as the above chip adaptive fault handling system. Therefore, for the details not described in this embodiment, reference may be made to the relevant descriptions of the above chip adaptive fault handling system, and this embodiment will not be elaborated herein.
[0163] In summary, the chip adaptive fault handling system and method provided by the embodiments of the present invention include a configuration unit 11, a signal acquisition unit 12, a signal processing unit 13, and an output control unit 14. The configuration unit 11 is used to configure mode information based on the information obtained from the CPU. The signal acquisition unit 12 is used to obtain the acquisition signal of the target module based on the mode information. The signal processing unit 13 is used to process the acquisition signal based on the mode information and obtain a reset signal when it is determined that an exception is triggered. The output control unit 14 is used to output the reset signal to the target module to implement reset. In this solution, the chip can execute the processing of the acquisition signal based on the configuration of the CPU, and thus can generate a reset signal to execute reset when an exception is triggered, without relying on the CPU to execute reset, reducing the processing burden of the CPU, and can also perform fault recovery on the CPU when the CPU has an exception, avoiding the problem of serious exceptions in the entire server system.
[0164] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0165] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A chip adaptive fault handling system, characterized in that: A fault processing module applied to a chip including multiple modules, the system includes a configuration unit, a signal acquisition unit, a signal processing unit and an output control unit, the configuration unit is respectively connected to a CPU in a server, the signal acquisition unit, and the signal processing unit, the signal acquisition unit is also connected to the signal processing unit, and the signal processing unit is also connected to the output control unit; The configuration unit is used to configure mode information based on information obtained from the CPU; The signal acquisition unit is used to acquire signals from the target module based on the mode information to obtain acquisition signals; The signal processing unit is used to process the collected signal based on the mode information, and obtain a corresponding reset signal when determining that a trigger abnormality occurs; The output control unit is used to output the reset signal to the target module to achieve the reset of the target module; The signal processing unit includes a trigger control subunit, an isolation processing subunit, an adaptive processing subunit and an output control subunit; The trigger control subunit is used to determine whether the acquisition signal triggers an abnormality based on the mode information; The isolation processing subunit is used for determining whether to maintain or filter out the abnormality based on the signal type of the acquisition signal when the trigger control subunit determines that the abnormality is triggered; The adaptive processing subunit is used to determine whether to output a reset signal in combination with the determination result of the trigger control subunit and the determination result of the isolation processing subunit; The output control subunit is used for outputting the reset signal to the output control unit when the adaptive processing subunit determines to output the reset signal.
2. The chip adaptive fault processing system according to claim 1, characterized in that: The mode information includes an abnormal trigger mode, the abnormal trigger mode includes a timestamp trigger mode, and the timestamp trigger mode includes a trigger signal type, a trigger mode, and a waiting time; The signal processing unit is used for triggering an exception after the waiting time is satisfied when it is determined that the acquisition signal is a trigger signal type in the timestamp trigger mode and the acquisition signal conforms to the trigger mode in the timestamp trigger mode.
3. The chip adaptive fault processing system according to claim 2, characterized in that: The system also includes a storage unit; The signal acquisition unit is also used to store the acquired signal into the storage unit; The signal processing unit is further configured to extract a signal segment within the waiting time of the acquisition signal from the storage unit and output the signal segment to the output control unit when determining that the acquisition signal triggers the timestamp trigger mode; The output control unit is further used to output the signal segment to other modules of the chip used for signal analysis.
4. The chip adaptive fault processing system according to claim 1, characterized in that: The system further comprises a storage unit, wherein the mode information comprises storage mode information and an off-chip storage address; The signal acquisition unit is used to store the acquired signal into the storage unit based on the storage mode information; The output control unit is further used to store the required backup signal into the off-chip Flash based on the off-chip storage address.
5. The chip adaptive fault processing system according to claim 1, characterized in that: The mode information includes a plurality of abnormal trigger modes, each of which includes a triggering method and a triggering signal type; The signal processing unit is used to determine the target trigger signal type corresponding to the acquisition signal, and judge whether the acquisition signal triggers an abnormality based on the target trigger mode corresponding to the target trigger signal type.
6. The chip adaptive fault processing system according to claim 5, characterized in that: The triggering method includes triggering events, the number of triggering events and the combination method; The trigger events include rising edge events, falling edge events and preset level width events, and the combination methods include multiple methods each consisting of at least one trigger event.
7. The chip adaptive fault processing system according to claim 1, characterized in that: The mode information includes a plurality of acquisition modes, each of which corresponds to an acquisition port and a signal type of the chip; The signal acquisition unit is used to determine a target acquisition port and a target signal type based on the acquisition mode in the mode information, and acquire a signal corresponding to the target signal type based on the target acquisition port to obtain an acquisition signal.
8. The chip adaptive fault processing system according to claim 7, characterized in that: The configuration unit is used to uniformly configure the port information of the corresponding collection ports in each of the collection modes and the signal information corresponding to the signal type into a mode template.
9. A chip adaptive fault processing method, characterized in that: The chip adaptive fault processing system applied to any one of claims 1 to 8 comprises a configuration unit, a signal acquisition unit, a signal processing unit and an output control unit, and the chip adaptive fault processing method comprises: The configuration unit performs mode information configuration based on information obtained from the CPU; The signal acquisition unit acquires signals from the target module based on the mode information to obtain acquisition signals; The signal processing unit processes the collected signal based on the mode information, and obtains a corresponding reset signal when determining that a trigger abnormality occurs; The output control unit outputs the reset signal to the target module to achieve the reset of the target module.
Citation Information
Patent Citations
System and method for assisting CPU to drive chips
CN101131657A
Fault management system for vehicle-specification-level chip function safety
CN110955571A