Quick exchange chip resetting method and system

By identifying chip anomalies through hardware monitoring circuits and anomaly templates, dynamically determining the reset granularity and performing hierarchical resets, the problem of low chip reset efficiency in existing technologies is solved, enabling fast and stable fault handling and system recovery.

CN120973590APending Publication Date: 2025-11-18SOUTHERN POWER GRID DIGITAL GRID RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510831768.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing chip reset technology suffers from low efficiency, false positives and false negatives, and insufficient reset range in anomaly identification and reset granularity control, leading to delays in fault location and system instability.

Method used

The system uses hardware monitoring circuits to sample abnormal signals in real time, and combines hardware hash circuits with preset abnormal templates to identify and locate abnormal types, dynamically determine the reset granularity, perform hierarchical resets, and switch redundant paths to reconstruct the system after a reset failure.

Benefits of technology

It enables precise anomaly location and rapid response, ensures priority recovery of critical modules, improves system fault handling efficiency and stability, reduces energy consumption, and enhances system reliability and fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973590A_ABST
    Figure CN120973590A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid exchange chip resetting method and system, and belongs to the technical field of chips. The method comprises the steps of abnormal trigger type identification, reset granularity dynamic determination, operation state freezing, hierarchical reset execution, reset validity verification and redundant reset system reconstruction, rapid abnormal positioning is achieved through hardware monitoring and feature matching, a dynamic reset strategy generation mechanism is combined, the reset granularity is accurately controlled according to the abnormal severity degree, and the system reliability is improved. Excessive reset or insufficient reset is avoided, the system resource utilization rate and the fault processing efficiency are improved, a state freezing technology is operated to maintain the state of a key register, progressive pulse injection and priority control of hierarchical reset execution are matched, it is ensured that a key module is preferentially recovered, the energy consumption is reduced while the reset effect is ensured, and the system performance is improved. Reset effectiveness multi-dimensional verification is combined with redundancy reset and a system reconstruction mechanism, a standby repair scheme is dynamically triggered for different fault types, and the reliability, fault-tolerant capability and maintainability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of chips, in particular to a fast exchange chip reset method and system. BACKGROUND

[0002] With the deep penetration of embedded systems in key fields such as power grids and industrial control, the stability and fault recovery efficiency of chip operation have become technical bottlenecks. The existing chip reset technology often has the following problems:

[0003] Abnormal identification depends on software scanning mechanism, which is difficult to quickly locate the specific function module or CPU core exception in complex hardware environment, resulting in delay of fault location, especially in the multi-signal concurrent scene of power grid Internet of Things terminal equipment, the traditional method is easy to misjudge or miss; the reset granularity control is extensive, often using full system reset or fixed level reset strategy, which cannot be dynamically adjusted according to the severity of the exception, which may cause state loss in non-fault area due to excessive reset, and may cause repeated exceptions due to insufficient reset range; there is lack of hardware level state freezing and redundancy reconstruction mechanism, the key register state is easy to lose during reset process, and when facing stubborn faults of more than three times of reset failure, the hardware level migration of function cannot be realized, resulting in long-time interruption of system. SUMMARY

[0004] The purpose of the present application is to provide a fast exchange chip reset method and system to solve the problems raised in the background art.

[0005] To achieve the above purpose, the present application provides the following technical scheme: a fast exchange chip reset method, comprising the following steps:

[0006] Abnormal trigger type identification: in response to the abnormal signal generated in the process of chip operation, the abnormal characteristic parameters are obtained and the abnormal type is judged, the specific function module or CPU core where the abnormality occurs is determined, and the abnormal type identification is generated;

[0007] Dynamic determination of reset granularity: according to the abnormal type identification, the preset reset strategy library is automatically matched to determine the corresponding reset granularity, and the reset instruction and reset sequence are generated;

[0008] Freezing of running state: before executing the reset instruction, the current running state of the target reset region is frozen, and the key register state is maintained at the instantaneous value before reset;

[0009] Hierarchical reset execution: according to the determined reset granularity and reset sequence, reset pulse injection is performed on the abnormal module or CPU core, and the clock signal and power signal fluctuation of the reset region are detected in real time;

[0010] Reset effectiveness verification: after each reset operation, a preset check instruction set test path is sent to the reset area, the response signal returned by the reset area is received and analyzed, and the reset result is determined;

[0011] Redundant reset system reconstruction: when the hierarchical reset fails three times, the standby reset path is switched to, a hot restart operation is performed, and the hardware state reconstruction logic is called to dynamically remap the configuration registers of the reset area, and temporarily migrate the functions of the abnormal module to the redundant hardware unit.

[0012] Further, the abnormal trigger type identification specifically includes:

[0013] In response to the abnormal signal generated during the operation of the chip, the hardware monitoring circuit samples the power fluctuation signal, the clock offset signal, the bus error signal and the CPU core abnormal interrupt signal in real time; the sampling signal is converted into a feature parameter group containing voltage fluctuation amplitude, clock frequency deviation value, bus error code and interrupt vector address;

[0014] According to the preset abnormal feature library, the feature parameter group is matched in stages;

[0015] The bus error code and the interrupt vector address are quickly compared by the hardware hash circuit to locate the functional module or the CPU core where the abnormality occurs;

[0016] The voltage fluctuation amplitude and the clock frequency deviation value are analyzed in time and frequency domains, and the similarity with the hardware abnormal template stored in the feature library is calculated to determine the abnormal type, and the hardware abnormal template includes power transient disturbance template, clock crystal oscillator offset template and bus timing error template;

[0017] An abnormality degree threshold is set, a location identifier containing the abnormal location is generated according to the matching result, and the severity level is determined according to the deviation degree of the feature parameter group;

[0018] If the feature parameter deviation degree threshold is ≤10%, it is determined to be a slight abnormality;

[0019] If the deviation degree threshold is between 10% and 50%, it is determined to be a serious abnormality;

[0020] If the deviation degree threshold is > 50% or involves a key module, it is determined to be a fatal abnormality, and the key module includes a master CPU core and a high-speed bus controller;

[0021] Finally, an abnormal type identifier containing the location identifier and the severity level is generated.

[0022] Further, the setting of the abnormality degree threshold specifically includes:

[0023] During normal operation of the chip, the voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling point are recorded;

[0024] The voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling point are normalized by using the maximum allowed voltage dynamic amplitude and clock frequency deviation value, to obtain normalized voltage dynamic amplitude and clock frequency deviation value;

[0025] The entropy increase deviation index corresponding to each sampling point is obtained by using the normalized voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling node;

[0026] The entropy increase deviation index average value and the standard deviation of the entropy increase deviation index corresponding to all sampling points are obtained according to the entropy increase deviation index of each sampling point;

[0027] The abnormality degree threshold value corresponding to each hardware abnormality template is set by using the entropy increase deviation index average value and the standard deviation of the entropy increase deviation index corresponding to all sampling points.

[0028] Further, the entropy increase deviation index corresponding to each sampling point is obtained by using the normalized voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling node, and specifically includes:

[0029] The power spectrum density P v (f) of the voltage fluctuation signal corresponding to all sampling points is retrieved;

[0030] The defined frequency band range corresponding to the power transient disturbance template, the clock crystal oscillator offset template and the bus timing error template is retrieved from the database;

[0031] The defined frequency band range corresponding to the power transient disturbance template, the clock crystal oscillator offset template and the bus timing error template is retrieved from the database;

[0031] The frequency energy ratio E r is obtained by using the defined frequency band range corresponding to the power transient disturbance template, the clock crystal oscillator offset template and the bus timing error template in combination with the power spectrum density P v (f) of the voltage fluctuation signal corresponding to all sampling points;

[0032] The entropy increase deviation index corresponding to each sampling point is obtained by using the frequency energy ratio in combination with the normalized voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling node.

[0033] Further, the reset granularity is dynamically determined, and specifically includes:

[0034] The abnormal position identifier and the severity level in the abnormal type identifier are extracted by the hardware analysis circuit, and a strategy retrieval address code is generated;

[0035] The address code driven hardware retrieval circuit retrieves a preset reset strategy library in parallel, the reset strategy library is stored in the form of a hardware register group, and contains a reset granularity mapping table corresponding to an exception severity and an exception position;

[0036] Among them, the slight exception corresponds to the local reset strategy; the severe exception corresponds to the hierarchical reset strategy; and the fatal exception corresponds to the whole system hot restart strategy;

[0037] According to the retrieved reset strategy, a corresponding reset instruction is generated through a hardware state machine;

[0038] If it is a local reset strategy of a single CPU core exception, a reset instruction containing only the address of the target CPU core is generated, and is sent to the reset control register of the core through a dedicated hardware channel;

[0039] If it is a hierarchical reset strategy across modules, a reset instruction set containing the addresses of related functional modules is generated, and a granularity control signal is generated according to a preset module importance priority, the granularity control signal is used to identify the influence range and execution order of the reset operation;

[0040] When the hierarchical reset strategy is triggered, the module addresses in the reset instruction set are sorted through a hardware priority arbitration circuit, a high-priority reset pulse is generated for a key module; a medium-priority reset pulse is generated for a non-key peripheral module; and a low-priority reset pulse is generated for an auxiliary functional module;

[0041] The sorted reset sequence is output in the form of a pulse queue controlled by a hardware timing circuit.

[0042] Further, the running state is frozen, specifically including:

[0043] A corresponding area selection signal is generated through a hardware address decoding circuit, and a freeze enable signal and a clock gating signal are generated by driving a freeze control logic circuit with the area selection signal;

[0044] Among them, the freeze enable signal is used to activate the register latching mechanism of the target area; and the clock gating signal is used to cut off the clock input of the target area, and prevent the state update of the timing circuit;

[0045] The hardware latching circuit of the key register in the target area is triggered by the freeze enable signal, the instantaneous state value before reset is latched to a backup storage unit, and the connection between the target area and the system bus is disconnected through a bus isolation circuit;

[0046] After the latching operation is completed, the latched value of the key register is read through a state monitoring circuit, and a hardware comparison is performed with the instantaneous value before reset;

[0047] If the comparison is consistent, a freeze completion flag is generated, which is synchronized with the reset pulse generation circuit through the hardware timing control circuit; if the comparison is inconsistent, the latch flow is retriggered.

[0048] Further, the hierarchical reset execution specifically includes:

[0049] An initial reset pulse is injected to the target abnormal module or CPU core through a hardware pulse generator, and the hardware monitoring circuit is activated to sample the clock signal and power signal of the reset region in real time;

[0050] The hardware monitoring circuit acquires signal fluctuation data at a sampling frequency of 10 MHz and performs hardware comparison with a preset normal signal threshold range to generate an abnormality elimination state identifier, which includes elimination and non-elimination.

[0051] If the abnormality is not eliminated after the first reset, the hardware timing controller increases the reset pulse width at a preset interval time, and sequentially performs secondary and tertiary reset operations; each time the reset is performed, the hardware monitoring circuit synchronously updates the signal monitoring threshold, and cuts off the clock synchronization link between the target region and the non-reset region during the pulse injection period.

[0052] When the reset sequence contains multiple modules, the hardware priority arbitration circuit generates a pulse queue according to the importance priority of the modules.

[0053] Further, the reset effectiveness verification specifically includes:

[0054] The hardware test channel control circuit sends a preset verification instruction set to the dedicated test register group of the reset region;

[0055] The verification instruction set contains function module read-write instructions, register state query instructions, and bus timing verification instructions, which are sent in sequence in the form of a hardware queue, and the response receiving circuit is activated at the same time, and the response data frame returned by the reset region is captured in real time through a differential signal interface, the data frame contains instruction execution status code, register return value, and bus timing feedback signal;

[0056] The response data frame is decoded by the hardware analysis circuit to extract the binary feature code of the instruction execution status code, the register return value, and the clock period deviation value of the bus timing feedback;

[0057] The binary feature code of the instruction execution status code, the register return value, and the clock period deviation value of the bus timing feedback are compared with a preset normal state threshold library in parallel, which contains instruction execution status code threshold, register feature code hash value, and bus clock deviation threshold.

[0058] If all parameters are within the threshold range, a reset success identifier is generated; if any parameter exceeds the threshold, a reset failure identifier is generated, and the hardware code of the abnormal parameter is attached;

[0059] When receiving the reset failure identifier, an activation signal is sent to the backup reset mechanism control module through the hardware interrupt controller, which drives the backup reset path selection circuit to dynamically select the backup mechanism according to the abnormal parameter hardware code;

[0060] If the abnormal parameter involves clock deviation, switch to the redundant clock source calibration path; if the abnormal parameter involves register error, trigger the hardware error correction code recalculation mechanism; if the abnormal parameter involves instruction execution error, call the pre-stored fault recovery microcode to reconstruct the instruction stream;

[0061] Write the reset failure information into the hardware log register group, record the clock cycle number, power voltage instantaneous value and reset pulse width parameters at the time of failure.

[0062] Further, the redundant reset system reconstruction specifically includes:

[0063] Cut off the main reset path through the hardware path switching circuit, and activate the hot restart control logic of the built-in redundant clock source and power management module in the chip;

[0064] The hot restart control logic sends a hot restart instruction to the main control CPU core, and at the same time maintains the system power voltage at 80%-105% of the normal working range through a hardware signal;

[0065] Call the hardware state reconstruction logic, dynamically remap the configuration register group of the reset area through the address decoding circuit, extract the functional configuration parameters of the abnormal module, locate the available redundant hardware unit through the circuit, generate a remapping instruction set, write the configuration parameters of the abnormal module into the registers of the redundant hardware unit, and update the address mapping relationship of the module in the system routing table;

[0066] Switch the input and output signals of the abnormal module to the corresponding interfaces of the redundant hardware unit through the bus arbitration circuit;

[0067] Perform functional verification on the redundant hardware unit, send a preset reconstruction verification instruction set, and the reconstruction verification instruction set includes register read-write verification, interrupt response test and bus timing verification;

[0068] Receive the response signal returned by the redundant unit and compare it with the preset normal state threshold;

[0069] If the verification passes, a reconstruction completion flag is generated and the module state bit of the system state register is updated; if the verification fails, a full-system cold restart process is triggered, and the timestamp of the reconstruction operation, the abnormal module address and the redundant unit address are written into the hardware log register group.

[0070] Further, a fast exchange chip reset system comprises:

[0071] An abnormality monitoring and identification module is configured to sample abnormal signals in real time and convert them into feature parameters, locate the abnormal position through hardware hashing and template matching, and generate an abnormal type identification containing the position and severity according to the deviation threshold.

[0072] A strategy dynamic generation module is configured to analyze the abnormal type identification and retrieve a reset strategy library, generate a local reset, a hierarchical reset or a hot restart instruction, and form a reset sequence according to the module importance priority.

[0073] A freezing control module is configured to generate a freezing signal through address decoding, latch the key register state of the target area and isolate the bus, and compare the latched value with the instantaneous value to ensure the effectiveness of the state freezing.

[0074] A reset execution module is configured to inject a microsecond-level pulse according to the reset sequence and monitor the execution effect, and increase the pulse width when the abnormality is not eliminated, and control the multi-module reset timing according to the priority.

[0075] An effectiveness verification module is configured to send a verification instruction set, analyze the response data and compare them with the threshold library, generate a reset success or failure identification, and trigger the corresponding backup mechanism according to the abnormal code when the reset fails.

[0076] A system reconstruction module is configured to activate the redundant clock source to perform a hot restart after hierarchical reset fails, dynamically remap the configuration register to the redundant hardware unit, verify the reconstruction state and record the log.

[0077] Compared with the prior art, the present application has the following advantages:

[0078] 1. The abnormal trigger type identification and reset strategy dynamic generation of the present application can realize accurate positioning and type determination of abnormal position through real-time sampling of abnormal signals such as power supply and clock by a hardware monitoring circuit and hierarchical matching of a hardware hashing circuit and a preset abnormal template, and then generate corresponding reset instructions according to the abnormal type by dynamically matching a reset strategy library, so as to complete abnormal feature extraction and strategy formulation in a very short time, avoid excessive response of non-critical abnormalities, ensure that core functional units can be quickly captured, improve system resource utilization and fault handling efficiency, and make the system more efficiently cope with various abnormal situations.

[0079] 2. The operating state freezing and hierarchical reset execution of the application generates a freeze signal to latch the target region key register state and isolate the bus through hardware address decoding, simultaneously injects pulses according to the reset sequence and monitors in real time, dynamically adjusts the pulse width and reset sequence according to abnormal conditions, maintains the accuracy of the key register state, prevents reset interference from spreading, ensures that critical modules are restored first, realizes precise control of the reset range and progressive abnormality cleaning, reduces energy consumption while ensuring reset effectiveness, improves the timing reliability of the reset operation and system stability, and makes the reset process more stable, efficient and energy-saving.

[0080] 3. The reset effectiveness verification and redundant reset system reconstruction of the application verifies the reset effect through a multi-dimensional verification instruction set, triggers a backup mechanism for different abnormal parameters, activates the backup path and performs system reconstruction after hierarchical reset fails, quickly determines the reset effectiveness, provides customized repair solutions for different fault types, realizes function migration using redundant hardware units, ensures that the system can maintain high functional continuity and availability when facing complex abnormalities or reset failures, improves the maintainability and fault tolerance of the system, and makes the system more reliable and stable when facing various faults. BRIEF DESCRIPTION OF DRAWINGS

[0081] Figure 1 The figure is a flowchart of the fast exchange chip reset method of the application. DETAILED DESCRIPTION

[0082] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0083] Please refer to Figure 1 The application provides the following technical solutions:

[0084] A fast exchange chip reset method, comprising the following steps:

[0085] Abnormal trigger type identification: in response to abnormal signals generated during chip operation, abnormal characteristic parameters are obtained and abnormal types are determined to determine the specific functional module or CPU core where the abnormality occurs, and an abnormal type identifier is generated;

[0086] Dynamic reset granularity determination: according to the abnormal type identifier, the preset reset strategy library is automatically matched to determine the corresponding reset granularity, and reset instructions and reset sequences are generated;

[0087] Running state freezing: freeze the current running state of the target reset area before executing the reset instruction, and maintain the key register state at the instantaneous value before the reset;

[0088] Hierarchical reset execution: perform reset pulse injection on the abnormal module or CPU core according to the determined reset granularity and reset sequence, and real-time detect the clock signal and power signal fluctuation of the reset area;

[0089] Reset effectiveness verification: after each reset operation, send a preset verification instruction set test path to the reset area, receive and analyze the response signal returned by the reset area, and determine the reset result;

[0090] Redundant reset system reconstruction: when the hierarchical reset fails three times, switch to the standby reset path, perform a hot restart operation and call the hardware state reconstruction logic to dynamically remap the configuration registers of the reset area, and temporarily migrate the functions of the abnormal module to the redundant hardware unit.

[0091] Abnormal trigger type identification, specifically including:

[0092] In response to the abnormal signals generated during the operation of the chip, the hardware monitoring circuit samples the power fluctuation signal, clock offset signal, bus error signal and CPU core abnormal interrupt signal in real time; convert the sampling signal into a feature parameter group containing voltage fluctuation amplitude, clock frequency deviation value, bus error code and interrupt vector address;

[0093] According to the preset abnormal feature library, the feature parameter group is matched in stages;

[0094] The bus error code and interrupt vector address are quickly compared by the hardware hash circuit to locate the functional module or CPU core where the abnormality occurs;

[0095] The voltage fluctuation amplitude and clock frequency deviation value are analyzed in time and frequency domains, and the similarity with the hardware abnormality template stored in the feature library is calculated to determine the abnormal type, and the hardware abnormality template includes power transient disturbance template, clock crystal oscillator offset template and bus timing error template;

[0096] Set the abnormality degree threshold, generate a location identifier containing the abnormal location according to the matching result, and grade the severity according to the deviation degree of the feature parameter group;

[0097] If the feature parameter deviation degree threshold is ≤10%, it is determined to be a slight abnormality;

[0098] If the deviation degree threshold is between 10% and 50%, it is determined to be a serious abnormality;

[0099] If the deviation threshold is greater than 50% or involves a critical module, it is determined as a fatal exception, and the critical module includes a master CPU core and a high-speed bus controller;

[0100] Finally, an exception type identifier containing a location identifier and a severity level is generated.

[0101] In the above embodiment, through real-time sampling of signals such as power fluctuations and clock drifts by the hardware monitoring circuit and conversion of feature parameters, combined with the hierarchical matching mechanism of the hardware hash circuit and the preset exception template, accurate positioning and type determination of the exception location can be achieved. The extraction and comparison of exception features can be completed within nanoseconds, which improves efficiency compared to traditional software scanning methods. At the same time, the three-level exception grading strategy based on the deviation threshold (mild / serious / fatal) can avoid excessive response to non-critical exceptions, reduce unnecessary reset operations, and improve system resource utilization. In addition, the priority recognition mechanism of the critical module (master CPU core, high-speed bus controller) can ensure that the exception of the core functional unit is quickly captured, providing data support for accurate formulation of subsequent reset strategies, thereby shortening the average fault positioning time of the system.

[0102] Specifically, the setting of the exception degree threshold value specifically includes:

[0103] When the chip is normally running, record the voltage dynamic amplitude and clock frequency deviation value corresponding to each collection point;

[0104] Normalize the voltage dynamic amplitude and clock frequency deviation value corresponding to each collection point using the maximum allowed voltage dynamic amplitude and clock frequency deviation value, to obtain the normalized voltage dynamic amplitude and clock frequency deviation value;

[0105] Obtain the entropy increase deviation index corresponding to each sampling point using the normalized voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling node;

[0106] Obtain the entropy increase deviation index average value and the standard deviation of the entropy increase deviation index corresponding to all sampling points according to the entropy increase deviation index of each sampling point;

[0107] Set the exception degree threshold value corresponding to each hardware exception template using the entropy increase deviation index average value and the standard deviation of the entropy increase deviation index corresponding to all sampling points.

[0108] Wherein, the exception degree threshold value corresponding to each hardware exception template is obtained by the following formula:

[0109] G = λ * μ hist + 2 * σ hist

[0110] Wherein, G represents the exception degree threshold value corresponding to each hardware exception template; μ histrepresents the average value of the entropy deviation index corresponding to all sampling points; λ represents a data scaling coefficient, and the value range is 0.1-1.8; σ hist represents the standard deviation of the entropy deviation index corresponding to all sampling points.

[0111] The technical effect of the above technical solution is that the technical solution sets around the hardware exception degree threshold value, builds a complete, dynamic and intelligent system from data acquisition to threshold generation, and realizes the deep optimization of hardware exception management. In the normal running phase of the chip, the basic data such as voltage fluctuation amplitude and clock frequency deviation value are collected in detail first, and the maximum allowed value is normalized to eliminate the interference of different parameter magnitude differences on subsequent analysis, so that the data is in a unified analysis dimension. Then the entropy deviation index is introduced, and with the help of this index which can reflect the change of system disorder degree, the subtle trend of voltage, clock and other parameters deviating from the normal state is accurately captured, and the implicit fluctuation of hardware operation is converted into a quantifiable numerical feature. Subsequently, based on the entropy deviation index of all sampling points, the average value and the standard deviation are calculated, and these two statistics respectively outline the concentration trend and the dispersion degree of the entropy deviation in the normal running of the chip, and then combined with the flexible adjustable data scaling coefficient λ, the formula G = λ * μ hist + 2 * σ hist The abnormality degree threshold value of each hardware exception template is constructed. This process breaks the experience dependence of traditional threshold setting, makes the threshold deeply fit the "health baseline" of the actual running of the chip, not only can dynamically adapt to the normal fluctuation of the chip under different loads, environment (such as temperature change, long-term aging), but also can be targeted at different hardware exception templates such as power transient interference and clock crystal oscillator offset, according to the differences of their physical propagation law and the influence on the system, customize differentiated threshold, lay a solid foundation for subsequent accurate identification of exception type and reasonable judgment of exception degree, improve the scientificity and intelligence of chip hardware exception management from the bottom logic, and escort the stable operation of the chip.

[0112] Compared with the abnormal degree threshold acquisition method in the prior art, the technical scheme realizes multi-dimensional breakthrough in rationality and system matching. In terms of threshold acquisition rationality, the traditional technology mainly depends on experience value or chip nominal parameter to set the threshold, which has obvious shortcomings: on the one hand, the experience value is independent of the actual working condition of the chip, and the voltage fluctuation and clock deviation baseline of the chip under different application scenarios (such as high load operation of industrial-grade chips and daily use of consumer-grade chips) and different environmental conditions (temperature and humidity changes) are significantly different. Fixed experience threshold is easy to misjudge normal fluctuation as abnormal or miss real abnormality; on the other hand, the hardware abnormality type characteristics are not distinguished, the physical mechanisms of power transient disturbance and clock crystal oscillator shift are different, and the unified threshold cannot accurately adapt to the "normal-abnormal" boundary. The present scheme starts from the actual running data of the chip, takes the voltage, clock and other parameters during normal operation as the basis, normalizes and entropy-increases the modeling, so that the threshold closely fits the real working condition of the chip, and the threshold is set for different abnormal templates, so that the threshold setting considers the individual working condition difference of the chip and adapts to the physical characteristics of the abnormal type, greatly improving the rationality. In terms of system matching, the traditional threshold is static, and once it is set, it will not change for a long time, and it is difficult to cope with the normal baseline drift caused by chip aging and environmental gradual change, and a single threshold is used to process all abnormal types, which cannot distinguish the damage degree of abnormality to the system (such as the influence of main CPU core abnormality and ordinary module abnormality). The technical scheme has the advantages of dynamic updating and hierarchical matching: by continuously collecting normal running data, the mean and standard deviation of the entropy-increasing deviation index can be automatically updated, the threshold can be dynamically adjusted, and the system state can be adapted in long-term operation; the adjustability of the data scaling coefficient λ can flexibly control the threshold sensitivity according to the system demand, balance the system safety and running efficiency, and change the threshold from a "stereotyped static value" to an "intelligent baseline that evolves with the system and adapts to the demand", which improves the matching degree of the chip system, significantly reduces the risk of abnormal false alarm and omission in high reliability scenarios (such as power grid system), and enhances the system fault tolerance and stability.

[0113] Specifically, the entropy-increasing deviation index corresponding to each sampling point is obtained by using the normalized voltage dynamic amplitude and clock frequency deviation value corresponding to each sampling node, specifically including:

[0114] The power spectrum density P of the voltage fluctuation signal corresponding to each sampling point is obtained by using the voltage fluctuation signal corresponding to each sampling point. v (f);

[0115] The defined frequency band range corresponding to the power transient disturbance template, the clock crystal oscillator shift template and the bus timing error template is retrieved from the database; wherein the upper limit value of the defined frequency band range corresponding to the power transient disturbance template, the clock crystal oscillator shift template and the bus timing error template is f max; the lower limit value of the defined frequency band range corresponding to the power supply transient interference template, the clock crystal oscillator offset template and the bus timing error template is f min ;

[0116] The power spectrum density P of the voltage fluctuation signal corresponding to all sampling points is obtained by using the defined frequency band range corresponding to the power supply transient interference template, the clock crystal oscillator offset template and the bus timing error template v (f) obtaining a frequency domain energy ratio E r ;

[0117] The frequency domain energy ratio is obtained by the following formula:

[0118]

[0119] Wherein, E r represents the frequency domain energy ratio; f max represents the upper limit value of the defined frequency band range corresponding to the power supply transient interference template, the clock crystal oscillator offset template and the bus timing error template; f min represents the lower limit value of the defined frequency band range corresponding to the power supply transient interference template, the clock crystal oscillator offset template and the bus timing error template; specifically, represents the voltage fluctuation signal energy in the specific frequency band [f min ,f max ]. represents the total energy of the voltage fluctuation signal (full frequency domain integration);

[0120] The entropy increase deviation index corresponding to each sampling point is obtained by using the frequency domain energy ratio combined with the normalized voltage amplitude and the clock frequency deviation value corresponding to each sampling node.

[0121] The entropy increase deviation index corresponding to each sampling point is obtained by the following formula:

[0122]

[0123] Wherein, D hist represents the entropy increase deviation index corresponding to each sampling point; d v and d f respectively represent the normalized voltage amplitude and the clock frequency deviation value; r represents a preset frequency domain energy adjustment factor, and the value range of the preset frequency domain energy adjustment factor is 0.5-1.5. Specifically, log 10 (1+d v ) and log 10 (1+d f ) When the parameters approach normal (d v →0, d f →0), the logarithmic term tends to log10 (1) = 0, indicating that the system is ordered; when the parameter deviation increases (such as voltage mutation, clock drift), d v / d f rise, the logarithmic term grows nonlinearly, simulating the entropy increase characteristic that "the system disorder degree increases due to parameter abnormality". When the chip is normally running, the voltage and clock are stable (low disorder degree), and when the parameter fluctuates abnormally, the stability is broken (high disorder degree), and the logarithmic function makes "small deviation weak response and large deviation strong response", which conforms to the evolution law of hardware failure from "gradual change to mutation". At the same time, and E r in the formula represents the energy of the abnormal frequency band (such as the high frequency band corresponding to power supply interference, and the specific frequency point corresponding to clock offset) calculated by power spectrum integration, and the proportion in the total signal energy; r is a frequency energy adjustment factor (0.5-1.5), which controls the influence weight of voltage fluctuation (d v ) and clock deviation (d f ) on the frequency energy abnormality:

[0124] When Er is high (abnormal frequency band energy dominates):

[0125] If r < 1 (emphasis on voltage-related abnormalities such as power supply interference), amplify the weight of log 10 (1+d v ), so that the contribution of voltage fluctuation to entropy increase is more significant; if r > 1 (emphasis on clock-related abnormalities such as crystal oscillator offset), decreases due to high Er, but through the adjustment of the denominator |1-r|, the weight of clock deviation (d f ) is actually adapted to the frequency domain characteristics. The above formula dynamically allocates the contribution of voltage and clock parameters to "entropy increase deviation" according to the energy distribution of abnormal frequency band, so that the entropy increase index is more consistent with the physical source of hardware abnormality (such as power supply problems preferentially associated with voltage fluctuation, and clock problems preferentially associated with frequency deviation).

[0126] The technical effects of the above technical solution are: the technical solution constructs a complete process from signal power spectrum analysis to multi-parameter fusion modeling around the entropy increase deviation index calculation of chip hardware abnormality detection, and realizes accurate quantification of the potential trend of hardware abnormality. First, by collecting the power spectrum density of voltage fluctuation signal, combining the exclusive frequency band range of different hardware abnormality templates (power supply transient interference, clock crystal oscillator offset, etc.), and using integral operation to obtain the frequency energy ratio E r , the energy proportion of the abnormal related frequency band is accurately extracted, and the characteristic distribution of hardware abnormality in the frequency domain is described. Subsequently, the normalized voltage amplitude, clock frequency deviation value and frequency energy ratio are fused, and the entropy increase deviation index D histThe time domain fluctuations of voltage and clock parameters are deeply integrated with frequency domain energy characteristics, and subtle abnormal deviations of hardware operation parameters are converted into quantifiable entropy increase indicators. This process realizes accurate description of chip hardware anomalies from "signal collection-frequency domain feature extraction-multi-domain parameter fusion modeling", provides high sensitivity and high differentiation quantitative basis for subsequent abnormal degree threshold setting and abnormal type identification, improves the intelligentization and precision level of chip hardware anomaly detection from the bottom logic, helps to capture potential hardware failure risks in advance, and ensures stable operation of the chip.

[0127] Compared with the hardware anomaly feature extraction and quantization method in the prior art, the technical scheme has obvious advantages in accuracy, comprehensiveness and adaptability to hardware characteristics. In the abnormal feature extraction dimension, the traditional technology focuses on single time domain or frequency domain parameters (such as only monitoring voltage fluctuation amplitude and clock frequency deviation absolute value), but there are obvious limitations. Although time domain parameters can reflect instantaneous anomalies, they are easily affected by random interference and are difficult to distinguish between real anomalies and noise. If the frequency domain analysis is not targeted to the exclusive frequency band of the hardware anomaly template, the extracted energy characteristics lack physical meaning and cannot accurately correspond to specific anomalies such as power transient interference and clock crystal oscillator offset. The present scheme calculates the frequency energy ratio E r by power spectral density integration, deeply combines the defined frequency band range of the hardware anomaly template, and makes the extracted frequency energy characteristics strongly related to the physical mechanism of the anomaly (such as specific high-frequency energy anomaly corresponding to power transient interference), so that the frequency domain analysis is upgraded from "non-discriminatory energy statistics" to "precise abnormal frequency band mining", greatly improving the physical explainability of abnormal features. In the aspect of multi-parameter fusion modeling, the traditional method often simply concatenates time domain and frequency domain parameters without considering the coupling relationship between parameters and the difference in contribution to anomalies. The present technical scheme introduces the entropy increase theory to fuse voltage, clock normalized parameters and frequency energy ratio through logarithmic operation, balances the influence weight of frequency energy on entropy increase deviation through adjustment factor r, and constructs D histThe hardware operation in time domain can be comprehensively reflected, the abnormal energy distribution in frequency domain is solved, the multi-dimensional abnormal characteristics are organically integrated, and the problem of single parameter "one-sided description of abnormality" is solved. In addition, the existing technology often lacks deep adaptation with the actual running state of the chip for the quantitative index of hardware anomaly. Based on the normalization processing of the chip normal running data and the exclusive frequency band setting of the hardware anomaly template, the entropy increase deviation index accurately fits the chip hardware characteristics. Whether it is the transient fluctuation of power interference or the gradual shift of clock crystal oscillator, it can be captured and distinguished with high sensitivity through the index. In practical application, it can effectively reduce the risk of abnormal missed detection (such as potential fault omission caused by ignoring frequency energy abnormality in traditional method) and false detection (such as misjudgment of noise as abnormality by single time domain parameter), and provide more reliable quantitative support for early discovery and early diagnosis of chip hardware anomaly. Especially in the scene of high reliability requirement of power grid control, it can significantly improve the fault prediction and health management level of chip system, and ensure the stable operation of equipment.

[0128] The reset granularity is dynamically determined, and specifically includes:

[0129] The abnormal position identifier and severity level in the abnormal type identifier are extracted by the hardware analysis circuit to generate a strategy retrieval address code;

[0130] The strategy retrieval address code is used to drive the hardware retrieval circuit to perform parallel retrieval on the preset reset strategy library, and the reset strategy library is stored in the form of a hardware register group, including a reset granularity mapping table corresponding to the abnormal severity and abnormal position;

[0131] Among them, the slight abnormality corresponds to a local reset strategy; the serious abnormality corresponds to a hierarchical reset strategy; and the fatal abnormality corresponds to a full-system hot restart strategy;

[0132] According to the retrieved reset strategy, a corresponding reset instruction is generated by a hardware state machine;

[0133] If it is a local reset strategy of a single CPU core abnormality, a reset instruction containing only the address of the target CPU core is generated, and is sent to the reset control register of the core through a dedicated hardware channel;

[0134] If it is a hierarchical reset strategy across modules, a reset instruction set containing the addresses of related functional modules is generated, and a granularity control signal is generated according to the preset module importance priority (master CPU core > high-speed bus controller > general-purpose peripheral module), and the granularity control signal is used to identify the influence range and execution order of the reset operation;

[0135] When the hierarchical reset strategy is triggered, the module addresses in the reset instruction set are sorted by a hardware priority arbitration circuit, a high-priority reset pulse is generated for the key module, a medium-priority reset pulse is generated for the non-key peripheral module, and a low-priority reset pulse is generated for the auxiliary functional module.

[0136] The sorted reset sequence is output in the form of a pulse queue controlled by a hardware timing circuit.

[0137] In the above embodiment, by using the parallel search mechanism of the hardware analysis circuit and the policy library, a matching reset policy can be dynamically generated according to the abnormal type identification, the granularity mapping table stored in the hardware register group supports microsecond-level search response, the memory query mode delay is reduced, and the local reset, hierarchical reset and hot restart strategies matched for different abnormal degrees (slight / serious / fatal) can realize accurate control of the reset range, for example, only 10 μs is needed for local reset when a single CPU core is abnormal, which can reduce the energy consumption by 85% compared with the full system reset, the module importance priority sorting (master CPU core > high-speed bus controller > general-purpose peripheral module) combined with the hardware priority arbitration circuit can ensure that the reset pulse of the key module is injected first, which improves the core function recovery speed, and at the same time, the hardware timing circuit controlled pulse queue output ensures that the execution error of the reset sequence is not more than 2 ns, which significantly improves the timing reliability of the reset operation.

[0138] The running state is frozen, specifically including:

[0139] The corresponding region selection signal is generated by the hardware address decoding circuit, and the freeze control logic circuit is driven by the region selection signal to generate a freeze enable signal and a clock gating signal;

[0140] The freeze enable signal is used to activate the register latching mechanism of the target region; the clock gating signal is used to cut off the clock input of the target region to prevent the state update of the timing circuit;

[0141] The hardware latching circuit of the key register in the target region is triggered by the freeze enable signal to latch the instantaneous state value before reset to the backup storage unit, and the connection between the target region and the system bus is disconnected through the bus isolation circuit;

[0142] After the latching operation is completed, the latched value of the key register is read by the state monitoring circuit, and a hardware comparison is performed with the instantaneous value before reset;

[0143] If the comparison is consistent, a freeze completion flag is generated, which is synchronized with the reset pulse generation circuit through the hardware timing control circuit; if the comparison is inconsistent, the latching process is triggered again.

[0144] In the above embodiment, the freeze enable signal and clock gate signal generated by the hardware address decoding and freeze control logic can realize the instantaneous latching of the target region register state and the bus isolation, can maintain the precision of the key register state within ±1 clock cycle, ensures that the program can seamlessly continue after reset, avoids the context loss problem caused by traditional reset, improves the task recovery efficiency by 90%, the bus isolation circuit disconnects the connection between the target region and the system bus, which can prevent the spread of interference signals during reset, reduces the running interference rate of the non-reset region, at the same time, the hardware comparison mechanism of the latched value and the instantaneous value can ensure the effectiveness of the state freezing, and the three times of re-latching flow design improves the freezing success rate, thereby improving the functional consistency after system reset.

[0145] The hierarchical reset execution specifically includes:

[0146] An initial reset pulse is injected to the target abnormal module or CPU core through a hardware pulse generator, and a hardware monitoring circuit is activated to sample the clock signal and power signal of the reset region in real time;

[0147] The hardware monitoring circuit acquires signal fluctuation data at a sampling frequency of 10 MHz, and compares the data with a preset normal signal threshold range to generate an abnormality elimination state identifier, which includes elimination and non-elimination.

[0148] If the abnormality is not eliminated after the first reset, the reset pulse width is increased by a preset interval time Δt (T2=T1+Δt, T3=T2+Δt) through a hardware timing controller, and secondary and tertiary reset operations are sequentially performed; each time the reset is performed, the hardware monitoring circuit synchronously updates the signal monitoring threshold, and cuts off the clock synchronization link between the target region and the non-reset region during the pulse injection period;

[0149] When the reset sequence contains multiple modules, a pulse queue is generated by a hardware priority arbitration circuit according to the importance priority of the modules;

[0150] A reset pulse with a width T1 is injected to the high-priority module, and its clock / power signal is continuously monitored until it is stable;

[0151] If the high-priority module is successfully reset, a pulse (pulse width T1) is injected to the medium-priority module in sequence, and if it fails, the pulse width increment strategy is executed;

[0152] The reset pulse width of the low-priority module defaults to T1, but can be dynamically adjusted to T2 / T3 after the high-priority module is successfully reset.

[0153] The reset pulse injection interval of each module is controlled by a hardware delay circuit to ensure that the next module reset is performed after the previous module reset is stable.

[0154] In the above embodiment, the hardware pulse generator cooperates with the real-time monitoring circuit with a 10MHz sampling frequency to achieve accurate injection of the reset pulse and effect feedback. The combination of the initial pulse width T1 and the incremental strategy (T2=T1+Δt, T3=T2+Δt) can provide a progressive reset scheme for different degrees of stubbornness of the abnormality, improve the success rate of clearing complex abnormalities, and cut off the clock synchronization link to avoid reset interference diffusion, control the clock jitter of the non-reset area within ±3%, and control the priority pulse queue of the multi-module reset, such as executing the operation of the medium and low priority modules after the high priority module reset is successful, to ensure the priority recovery of the key function of the system and improve the overall reset efficiency by 50%. The module reset interval controlled by the hardware delay circuit can ensure that the subsequent operation is performed after the previous module reset is stable, reduce the failure rate of cascading reset to below 1%, and significantly improve the reliability of the reset process.

[0155] The reset effectiveness verification specifically includes:

[0156] The hardware test path control circuit sends a preset verification instruction set to the special test register group of the reset area;

[0157] The verification instruction set includes function module read-write instructions, register state query instructions, and bus timing verification instructions. The verification instructions are sent in sequence in the form of a hardware queue, and a response receiving circuit is activated at the same time. The response data frame returned by the reset area is captured in real time through a differential signal interface. The data frame includes an instruction execution status code, a register return value, and a bus timing feedback signal;

[0158] The response data frame is decoded by a hardware analysis circuit to extract the binary feature code of the instruction execution status code, the register return value, and the clock period deviation value of the bus timing feedback;

[0159] The binary feature code of the instruction execution status code, the register return value, and the clock period deviation value of the bus timing feedback are compared with a preset normal state threshold library in parallel by hardware. The normal state threshold library includes an instruction execution status code threshold, a register feature code hash value, and a bus clock deviation threshold;

[0160] If all parameters are within the threshold range, a reset success identifier is generated. If any parameter exceeds the threshold, a reset failure identifier is generated, and the hardware code of the abnormal parameter is attached;

[0161] When the reset failure identifier is received, a hardware interrupt controller sends an activation signal to a backup reset mechanism control module. The activation signal drives a backup reset path selection circuit to dynamically select a backup mechanism according to the hardware code of the abnormal parameter;

[0162] If the abnormal parameter relates to clock deviation, switch to a redundant clock source calibration path; if the abnormal parameter relates to register error, trigger a hardware error correction code recalculation mechanism; if the abnormal parameter relates to instruction execution error, call pre-stored fault recovery microcode to reconstruct the instruction stream;

[0163] Write the reset failure information into the hardware log register group, record the clock cycle number, power supply voltage instantaneous value and reset pulse width parameters at the failure moment.

[0164] In the above embodiment, the parallel analysis mechanism of the verification instruction set and the response data sent through the hardware test path can complete the determination of the validity of the reset within 20us. The multi-dimensional threshold comparison of the instruction execution state code, the register feature code and the bus clock deviation can effectively improve the accuracy compared with single index verification. The dynamic selection of the backup mechanism based on abnormal parameter coding (clock deviation / register error / instruction error), such as redundant clock source calibration and hardware ECC recalculation, can provide customized repair solutions for different fault types, improve the effective response rate of the backup mechanism, and the real-time recording of the failure information by the hardware log register group provides complete timing data for subsequent fault analysis.

[0165] The redundant reset system reconstruction specifically includes:

[0166] Cut off the main reset path through the hardware path switching circuit, and activate the built-in redundant clock source and the power management module of the chip.

[0167] The hot restart control logic sends a hot restart instruction to the main control CPU core, and at the same time maintains the system power supply voltage at 80%-105% of the normal working range through a hardware signal;

[0168] Call the hardware state reconstruction logic, dynamically remap the configuration register group of the reset area through the address decoding circuit, extract the functional configuration parameters of the abnormal module, locate the available redundant hardware unit through the circuit, generate a remapping instruction set, write the configuration parameters of the abnormal module into the registers of the redundant hardware unit, and update the address mapping relationship of the module in the system routing table;

[0169] Switch the input and output signals of the abnormal module to the corresponding interfaces of the redundant hardware unit through the bus arbitration circuit;

[0170] Perform functional verification on the redundant hardware unit, send a preset reconstruction verification instruction set, and the reconstruction verification instruction set includes register read-write verification, interrupt response test and bus timing verification;

[0171] Receive the response signal returned by the redundant unit and compare it with the preset normal state threshold;

[0172] If the verification is passed, a reconstruction completion flag is generated and the module state bit of the system state register is updated; if the verification fails, a full-system cold restart process is triggered, and the timestamp of the reconstruction operation, the abnormal module address and the redundant unit address are written into the hardware log register group.

[0173] In the above embodiment, the hardware path switching circuit and the hot restart control logic of the redundant clock source can activate the backup reset path within 100 microseconds after the hierarchical reset fails, maintain the system voltage to ensure that the register state is not lost during the hot restart process, reduce power consumption and improve recovery speed compared to cold restart, the dynamic remapping of the configuration register and the function migration mechanism of the redundant hardware unit can automatically switch to the backup resource when the abnormal module fails, maintain the continuity of system function, the reconstruction verification instruction set and the threshold comparison mechanism can ensure the functional consistency of the redundant unit, the reconstruction completion flag generated after the verification is passed can make the system recover to normal operation within 50 milliseconds, and the full-system cold restart as a bottom strategy can control the system recovery time under extreme failure within 200 milliseconds, greatly improving the fault tolerance and availability of the system.

[0174] A fast exchange chip reset system, comprising:

[0175] An abnormality monitoring and identification module configured to sample abnormal signals in real time and convert them into feature parameters, locate the abnormal position through hardware hashing and template matching, and generate an abnormal type identification containing the position and severity according to the deviation threshold;

[0176] A strategy dynamic generation module configured to analyze the abnormal type identification and retrieve the reset strategy library, generate local reset, hierarchical reset or hot restart instructions, and form a reset sequence according to the module importance priority;

[0177] A freeze control module configured to generate a freeze signal through address decoding, latch the key register state of the target area and isolate the bus, and compare the latched value with the instantaneous value to ensure the effectiveness of the state freezing;

[0178] A reset execution module configured to inject microsecond-level pulses according to the reset sequence and monitor the execution effect, increment the pulse width when the abnormality is not eliminated, and control the multi-module reset timing according to the priority;

[0179] An effectiveness verification module configured to send a verification instruction set, analyze the response data and compare them with the threshold library, generate a reset success / failure identification, and trigger the corresponding backup mechanism according to the abnormal code when the verification fails;

[0180] A system reconstruction module configured to activate the redundant clock source to perform hot restart after hierarchical reset fails, dynamically remap the configuration register to the redundant hardware unit, verify the reconstruction state and record the log.

[0181] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, should be covered within the protection scope of the present application.

Claims

1. A method for fast exchange chip reset, the method comprising: The method comprises the following steps: Abnormal trigger type identification: in response to an abnormal signal generated during chip operation, the characteristic parameters of the abnormal signal are obtained, and the type of the abnormal signal is determined. The average value of the entropy increase deviation index and the standard deviation of the entropy increase deviation index corresponding to the sampling point of each characteristic parameter are obtained, and the abnormal degree threshold value corresponding to each hardware abnormal template is set. The specific functional module or CPU core where the abnormality occurs is determined, and an abnormal type identifier is generated; Dynamic determination of reset granularity: according to the abnormal type identifier, the preset reset strategy library is automatically matched to determine the corresponding reset granularity, and a reset instruction and a reset sequence are generated; Freezing of the running state: before the reset instruction is executed, the current running state of the target reset region is frozen, and the state of the key register is maintained at the instantaneous value before the reset; Hierarchical reset execution: according to the determined reset granularity and reset sequence, a reset pulse is injected to the abnormal module or CPU core, and the clock signal and power signal fluctuations of the reset region are detected in real time; Reset effectiveness verification: after each reset operation, a preset verification instruction set test path is sent to the reset region, the response signal returned by the reset region is received and analyzed, and the reset result is determined; Redundant reset system reconstruction: when the hierarchical reset fails three times, the standby reset path is switched to, a hot restart operation is performed, and the hardware state reconstruction logic is called to dynamically remap the configuration register of the reset region and temporarily migrate the function of the abnormal module to the redundant hardware unit.

2. The method of Claim 1, wherein, The abnormal trigger type identification specifically comprises: In response to an abnormal signal generated during chip operation, the hardware monitoring circuit is used to sample the power fluctuation signal, the clock offset signal, the bus error signal and the CPU core abnormal interrupt signal in real time. The sampling signal is converted into a characteristic parameter group containing voltage fluctuation amplitude, clock frequency deviation value, bus error code and interrupt vector address; According to the preset abnormal characteristic library, the characteristic parameter group is matched hierarchically; The bus error code and the interrupt vector address are quickly compared by the hardware hash circuit to locate the functional module or CPU core where the abnormality occurs; The voltage fluctuation amplitude and the clock frequency deviation value are analyzed in time domain and frequency domain, and the similarity with the hardware abnormal template stored in the characteristic library is calculated to determine the type of the abnormality. The hardware abnormal template includes a power transient disturbance template, a clock crystal oscillator offset template and a bus timing error template; An abnormal degree threshold value is set, a location identifier containing the location of the abnormality is generated according to the matching result, and the severity level is determined according to the deviation degree of the characteristic parameter group; If the deviation of the characteristic parameter from the abnormal degree threshold value is less than or equal to 10%, it is determined to be a slight abnormality; If the deviation of the abnormal degree threshold value is between 10% and 50%, it is determined to be a serious abnormality; If the deviation of the abnormal degree threshold value is greater than 50% or involves a key module, it is determined to be a fatal abnormality. The key module includes a master CPU core and a high-speed bus controller; Finally, an abnormal type identifier containing the location identifier and the severity level is generated.

3. A method for fast switching chip reset as recited in claim 2, wherein, The setting of the abnormal degree threshold value specifically comprises: During normal chip operation, the voltage amplitude and the clock frequency deviation value corresponding to each sampling point are recorded; The voltage amplitude and the clock frequency deviation value corresponding to each sampling point are normalized by using the maximum allowed voltage amplitude and the clock frequency deviation value, to obtain the normalized voltage amplitude and the clock frequency deviation value; The entropy increase deviation index corresponding to each sampling point is obtained by using the normalized voltage amplitude and the clock frequency deviation value corresponding to each sampling node. The average value and the standard deviation of the entropy increase deviation index corresponding to all sampling points are obtained according to the entropy increase deviation index of each sampling point. The abnormality degree threshold value corresponding to each hardware abnormality template is set by using the average value and the standard deviation of the entropy increase deviation index corresponding to all sampling points.

4. The method of Claim 3, wherein, The entropy increase deviation index corresponding to each sampling point is obtained by using the normalized voltage amplitude and the clock frequency deviation value corresponding to each sampling node, and specifically includes: retrieve power spectrum density P of all sampling point corresponding voltage fluctuation signal v (f); The defined frequency band range corresponding to the power transient disturbance template, the clock crystal oscillator offset template and the bus timing error template is called from the database; The defined frequency band range corresponding to the power supply transient disturbance template, the clock crystal oscillator offset template and the bus timing error template is combined with the power spectrum density P of the voltage fluctuation signal corresponding to all sampling points v (f) obtaining a frequency domain energy ratio E r ; The entropy increase deviation index corresponding to each sampling point is obtained by using the frequency energy ratio and the normalized voltage amplitude and the clock frequency deviation value corresponding to each sampling node.

5. The method of Claim 1, wherein The reset granularity is dynamically determined, and specifically includes: The abnormal position identifier and the severity level in the abnormal type identifier are extracted by the hardware analysis circuit to generate a strategy retrieval address code; The preset reset strategy library is searched in parallel by the hardware retrieval circuit driven by the strategy retrieval address code, the reset strategy library is stored in the form of a hardware register group, and contains a reset granularity mapping table corresponding to the abnormal severity and the abnormal position; The local reset strategy corresponds to a slight abnormality, the hierarchical reset strategy corresponds to a serious abnormality, and the full-system hot restart strategy corresponds to a fatal abnormality; According to the retrieved reset strategy, a corresponding reset instruction is generated by a hardware state machine; If the local reset strategy of a single CPU core abnormality is generated, a reset instruction containing only the address of the target CPU core is generated, and is sent to the reset control register of the core through a dedicated hardware channel; If the hierarchical reset strategy across modules is generated, a reset instruction set containing the addresses of related functional modules is generated, and a granularity control signal is generated according to a preset module importance priority, the granularity control signal is used to identify the influence range and the execution order of the reset operation; When the hierarchical reset strategy is triggered, the module addresses in the reset instruction set are sorted by the hardware priority arbitration circuit, a high-priority reset pulse is generated for the key module, a medium-priority reset pulse is generated for the non-key peripheral module, and a low-priority reset pulse is generated for the auxiliary functional module; The sorted reset sequence is output in the form of a pulse queue controlled by a hardware timing circuit.

6. The method of Claim 1, wherein The running state is frozen, and specifically includes: A corresponding area selection signal is generated by a hardware address decoding circuit, and a freeze enable signal and a clock gating signal are generated by driving a freeze control logic circuit by using the area selection signal; The freeze enable signal is used to activate the register latching mechanism of the target area, and the clock gating signal is used to cut off the clock input of the target area to prevent the state update of the timing circuit. The hardware latching circuit of the key register in the target area is triggered by the freeze enable signal to latch the instantaneous state value before reset to the spare storage unit, and the connection between the target area and the system bus is disconnected by the bus isolation circuit. After the latching operation is completed, the latched value of the key register is read through the status monitoring circuit and compared with the instantaneous value before the reset in hardware. If the comparison matches, a freeze completion flag is generated, which is synchronized with the reset pulse generation circuit via a hardware timing control circuit; if the comparison does not match, the latching process is retried.

7. The method of Claim 1, wherein The hierarchical reset execution specifically includes: An initial reset pulse is injected into the target faulty module or CPU core using a hardware pulse generator, while the hardware monitoring circuit is activated to sample the clock and power signals of the reset area in real time. The hardware monitoring circuit acquires signal fluctuation data at a sampling frequency of 10MHz and performs hardware comparison with a preset normal signal threshold range to generate an anomaly elimination status indicator, which includes eliminated and not eliminated. If the abnormality is not eliminated after the first reset, the hardware timing controller increments the reset pulse width at preset intervals to perform the second and third reset operations in sequence. During each reset, the hardware monitoring circuit synchronously updates the signal monitoring threshold and cuts off the clock synchronization link between the target area and the non-reset area during the pulse injection period. When the reset sequence contains multiple modules, a pulse queue is generated according to the importance priority of the modules through a hardware priority arbitration circuit.

8. A method for fast switching chip reset as described in claim 1, wherein, The reset validity verification specifically includes: The hardware test path control circuit sends a preset set of verification instructions to the dedicated test register group in the reset area. The verification instruction set includes functional module read / write instructions, register status query instructions, and bus timing verification instructions. The verification instructions are sent sequentially in the form of a hardware queue. At the same time, the response receiving circuit is activated, and the response data frame returned by the reset area is captured in real time through the differential signal interface. The data frame includes the instruction execution status code, the register return value, and the bus timing feedback signal. The response data frame is decoded by a hardware parsing circuit to extract the instruction execution status code, the binary feature code of the register return value, and the clock cycle deviation value of the bus timing feedback. The instruction execution status code, the binary signature of the register return value, and the clock cycle deviation value of the bus timing feedback are compared in parallel with a preset normal state threshold library. The normal state threshold library includes instruction execution status code thresholds, register signature hash values, and bus clock deviation thresholds. If all parameters are within the threshold range, a reset success flag is generated; if any parameter exceeds the threshold, a reset failure flag is generated, and the hardware code of the abnormal parameter is appended. When a reset failure flag is received, an activation signal is sent to the backup reset mechanism control module through the hardware interrupt controller. The activation signal drives the backup reset path selection circuit to dynamically select the backup mechanism according to the hardware encoding of the abnormal parameters. If the abnormal parameter involves clock deviation, switch to the redundant clock source calibration path; if the abnormal parameter involves register error, trigger the hardware error correction code recalculation mechanism; if the abnormal parameter involves instruction execution error, call the pre-stored fault recovery microcode to reconstruct the instruction stream. The reset failure information is written to the hardware log register group, recording the clock cycle number, instantaneous power supply voltage value, and reset pulse width parameter at the time of failure.

9. The method of claim 1, wherein the fast switching chip reset is performed by a reset controller. The redundant reset system reconfiguration specifically includes: The main reset path is cut off by a hardware path switching circuit, and the hot restart control logic of the chip's built-in redundant clock source and power management module is activated. The hot restart control logic sends a hot restart command to the main control CPU core, and at the same time maintains the system power supply voltage at 80%-105% of the normal operating range through hardware signals; The hardware state reconstruction logic is invoked, and the configuration register group of the reset area is dynamically remapped through the address decoding circuit. The functional configuration parameters of the abnormal module are extracted, and the search circuit locates the available redundant hardware unit. A remapping instruction set is generated, the configuration parameters of the abnormal module are written into the register of the redundant hardware unit, and the address mapping relationship of the module in the system routing table is updated. The input and output signals of the fault module are switched to the corresponding interface of the redundant hardware unit through the bus arbitration circuit. Perform functional verification on redundant hardware units and send a preset set of reconfiguration verification instructions, which includes register read / write verification, interrupt response testing and bus timing verification. Receive the response signal returned by the redundant unit and compare it with the preset normal state threshold; If the verification passes, a refactoring completion flag is generated and the module status bit in the system status register is updated; if the verification fails, a system-wide cold restart process is triggered, and the timestamp of the refactoring operation, the address of the abnormal module, and the address of the redundant unit are written to the hardware log register group.

10. A fast switching chip reset system applied to the fast switching chip reset method of claim 1, characterized in that, include: The anomaly monitoring and identification module is configured to sample abnormal signals in real time and convert them into feature parameters. It locates the anomaly location by matching hardware hashes with templates and generates an anomaly type identifier containing location and severity according to the degree of deviation from the threshold. The strategy dynamic generation module is configured to parse the exception type identifier and retrieve the reset strategy library to generate local reset, hierarchical reset or hot restart instructions, which are sorted according to the importance priority of the modules to form a reset sequence. The freeze control module is configured to generate a freeze signal through address decoding, latch the state of key registers in the target area and isolate the bus, and compare the latched value with the instantaneous value to ensure the effectiveness of the state freeze. The reset execution module is configured to inject microsecond-level pulses according to the reset sequence and monitor the execution effect. When the abnormality is not eliminated, the pulse width is increased, and the reset timing of multiple modules is controlled according to priority. The validity verification module is configured to send a set of verification instructions, parse the response data and compare it with the threshold library, generate a reset success or failure flag, and trigger the corresponding backup mechanism according to the exception code when it fails. The system reconfiguration module is configured to activate a redundant clock source to perform a hot restart after a failed tiered reset, dynamically remap configuration registers to redundant hardware units, verify the reconfiguration status, and record logs.

Citation Information

Cited By

  • Resetting method and system for reading abnormity of SD (Secure Digital) card

    CN121542093A

  • A method and system for resetting SD card read errors

    CN121542093B

  • Generation method of signal detector of clock and reset generator

    CN121560143A

  • Chip fast reset method and programmable distributed network chip

    CN122457570A