Clock line state-based SPI (Serial Peripheral Interface) bus active hang-up prevention method and device

By monitoring the clock line status in real time through the SPI controller and utilizing the interrupt mechanism, the SPI bus can be proactively protected from hanging up, thus resolving the problem of system paralysis caused by the clock line being pulled to death and improving the system's robustness and fault response speed.

CN120804009AActive Publication Date: 2025-10-17XIAMEN UNISOC TECH CO LTD

Patent Information

Application Number
CN202511311759.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

In the prior art, when the SPI clock line is hung by a slave device, it is usually discovered passively after a transmission failure. This cannot accurately distinguish between clock problems and other problems, resulting in low efficiency and long fault detection time.

Method used

The SPI controller monitors the clock line status in real time and compares it with the preset idle level. The internal counter and interrupt mechanism are used to perform active diagnosis when the clock line is pulled dead, and the interrupt request signal is used to notify the processor to recover from the fault.

Benefits of technology

It has achieved a transition from passive discovery after transmission failure to active diagnosis when the bus is abnormal, shortening the fault detection time, improving the robustness and stability of the system, and ensuring the accuracy of fault root cause location and the speed of fault response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804009A_ABST
    Figure CN120804009A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of electronic equipment communication, and discloses an SPI bus active hang-up prevention method and device based on a clock line state, and the method comprises the steps: presetting a first level signal of an SPI clock line in an idle state, and presetting a hang-up prevention counting threshold; the SPI controller monitors a second level signal of the clock line in real time; when the condition that the second level signal is inconsistent with the first level signal continues, an internal counter is started for counting; when the count value of the counter reaches an anti-hanging counting threshold value, the SPI controller judges that the clock line is hung and sets an internal hanging state flag; if the SPI is enabled to hang the interrupt function through the software configuration, the SPI controller sends an interrupt request signal to the processor; and the processor responds to the interrupt request signal and executes a corresponding fault recovery operation. The state of the clock line is monitored in real time through the SPI controller and compared with the preset idle level, and the recognition process can be started at the moment that the clock line is pulled dead.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of electronic equipment communication, and particularly relates to a SPI bus active anti-hanging method and device based on clock line state. BACKGROUND

[0002] SPI (Serial Peripheral Interface) is a high-speed, full-duplex, synchronous serial communication bus, which is widely used in short-distance communication between microcontrollers (MCU), system-on-chip (SOC) and various peripheral devices (such as memories, sensors, display modules, radio frequency chips, etc.) due to its simple protocol and high communication efficiency.

[0003] A typical SPI system is composed of a master device (Master) and one or more slave devices (Slave), which are connected through four basic signal lines. Respectively: CLK (Clock, clock) is a clock signal generated by the master device, which is used to synchronize data transmission; MOSI (Master Output Slave Input, master output / slave input) is the master data output and slave data input line; MISO (Master Input Slave Output, master input / slave output) is the master data input and slave data output line; SS / CS (Slave Select / Chip Select, chip select signal sent by the master device) is used to select a specific slave device for communication.

[0004] When a certain slave device connected to the SPI bus has a hardware failure, it may force the clock line to a fixed level, i.e. the clock line is pulled dead by the device, which is a common and serious problem. This problem directly leads to the SPI host being unable to generate a correct clock, thereby interrupting communication on the entire SPI bus.

[0005] In the prior art, when the SPI clock line is hung dead by the slave device, the transmission failure phenomenon after the occurrence of the situation is generally used for processing, i.e. by detecting whether the data transmission is timed out or failed to indirectly infer the problem. This method is not accurate and cannot distinguish whether the transmission failure is caused by clock problem or other problems; and it is inefficient and needs to wait for the entire transmission process to timeout. SUMMARY

[0006] The purpose of the present application is to monitor the clock line state in real time through the internal hardware of the SPI controller, and compare it with the preset idle level, which can start the identification process at the moment when the clock line is pulled dead, realizing the fundamental change from the "passive discovery after transmission failure" in the prior art to the "active diagnosis when the bus is abnormal", greatly shortening the fault detection time.

[0007] In a first aspect, an embodiment of the present application provides a SPI bus active anti-hang-up method based on clock line state, applied to a system on chip (SOC), the SOC serving as an SPI master device, comprising an SPI controller and a processor, and communicating with a slave device through an SPI interface, the method comprising: The processor configures a register of the SPI controller through software, and presets a first level signal of the SPI clock line in an idle state and a preset anti-hang-up count threshold value; The SPI controller monitors an actual level of the clock line in real time as a second level signal; When a condition that the second level signal is inconsistent with the first level signal is sustained, the SPI controller enables an internal counter to count; When a count value of the counter reaches the anti-hang-up count threshold value, the SPI controller determines that the clock line is hung up, and sets an internal hang-up state flag; If an SPI hang-up interrupt function is enabled through software configuration, the SPI controller sends an interrupt request signal to the processor; The processor executes a corresponding fault recovery operation in response to the interrupt request signal.

[0008] Optionally, the preset anti-hang-up count threshold value is determined by the following formula: Anti-hang-up count threshold value = time threshold value / system clock period; Wherein, the time threshold value is a minimum time width preset for filtering interference glitches on the clock line; and the system clock period is a period of a clock signal driving the SPI controller.

[0009] Optionally, the step of enabling the SPI hang-up interrupt function through software configuration comprises: The processor configures a specific bit of an interrupt enable register of the SPI controller through a write operation to enable the SPI hang-up interrupt function.

[0010] Optionally, the processor executing the corresponding fault recovery operation comprises the following steps: A reset signal is generated for the faulty slave device; wherein, the reset signal is generated by inverting a level of a chip select signal of the SPI bus, or by generating a reset pulse through a general input / output interface connected to a reset pin of the faulty slave device.

[0011] The reset signal is sent to the faulty slave device, so that the faulty slave device is reset and released from a pull-down state of the SPI clock line.

[0012] Optionally, after sending the reset signal to the faulty slave device, the method further includes: The processor reads a hang status flag in the SPI controller to check whether the clock line has returned to an idle state; If the deadlock status flag indicates that the clock line has returned to normal, continue subsequent data transmission; If the hang state flag indicates that the abnormality still exists, the fault recovery operation is re-executed.

[0013] Optionally, re-performing the fault recovery operation includes: Before reaching the preset maximum number of retries, the operation of generating and sending a reset signal is executed cyclically; After the maximum number of retries is reached, if the hang status flag still indicates that an abnormality exists, the processor generates and records an error log message; the error log message includes at least one of the following: a fault timestamp; the SPI controller number or bus identifier where the hang occurred; the clock line state level at the time of the hang; a preset idle state level; the value of the hang status flag; the number of retries of the reset operation; and identification information of the faulty slave device.

[0014] Optionally, the SPI controller can generate and send the interrupt request signal to the processor only after the SPI deadlock interrupt function is configured to be enabled through software; The SPI controller's deadlock status flag is configured to be cleared in one of the following ways: automatically cleared by hardware logic after the processor reads the status flag, or manually cleared by the processor through a software write operation.

[0015] In a second aspect, an embodiment of the present invention provides an SPI bus active anti-hang-up device based on the clock line state, which is applied to a system-on-chip (SOC). The device is an SPI controller, comprising: A configuration module is used to receive configuration information of the processor through a register interface to preset a first level signal and an anti-hanging count threshold when the SPI clock line is in an idle state; A monitoring module, configured to monitor the actual level of the clock line in real time as a second level signal; a counting and judging module, configured to enable an internal counter to count when the second level signal is inconsistent with the first level signal; and to determine that the clock line is hung when the count value of the counter reaches the anti-hang count threshold; A status flag module, configured to set an internal dead-end status flag when the counting and judging module determines that the clock line is dead-end; an interrupt control module, configured to send an interrupt request signal to the processor when the SPI hang-up interrupt function is enabled and the status flag module is set; The processor is configured to perform a corresponding fault recovery operation in response to the interrupt request signal.

[0016] In a third aspect, an electronic device is provided, comprising: The SPI bus active hang-up prevention device based on clock line state as described in the second aspect; The one or more processors are coupled to the device and configured to execute instructions to operate the device and perform a corresponding fault recovery operation in response to an interrupt request signal generated by the device.

[0017] In a fourth aspect, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method of the first aspect.

[0018] The technical solution provided by the embodiments of the present application can monitor the clock line state in real time through the SPI controller, compare it with the preset idle level, and start the identification process at the moment when the clock line is pulled up, thereby realizing a fundamental change from the passive discovery after transmission failure in the prior art to active diagnosis when the bus is abnormal, greatly shortening the fault detection time. Moreover, the judgment mechanism based on the configurable count threshold is adopted, and the hang-up is determined only when the abnormal level lasts for more than a certain time (determined by the hang-up prevention count threshold), thereby effectively filtering the transient glitches and noise interference on the bus, avoiding false alarms, and significantly improving the robustness and stability of the system.

[0019] Moreover, the present application can clearly locate the fault source to the specific hardware problem of "the clock line being pulled up", rather than the general "communication failure". This provides an accurate basis for subsequent execution of targeted recovery operations (such as resetting a specific slave device), and avoids the incorrect operations that may be taken due to the vague fault location in the prior art. Moreover, through the hardware linkage of interrupt enable and interrupt request signal, the interrupt request is automatically and timely sent to the processor on the premise of confirming the fault and interrupt enable. This allows the processor to quickly interrupt the current task and process the SPI bus exception, thereby greatly optimizing the fault response speed and meeting the requirements of real-time systems. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 A timing diagram is provided for the embodiments of the present application; Figure 2 A flowchart of an SPI bus active hang-up prevention method based on clock line state is provided for the embodiments of the present application; Figure 3A flow chart of one specific example provided for the embodiment of the present application; Figure 4 A schematic diagram of a clock line state-based SPI bus active anti-hanging device provided for the embodiment of the present application. DETAILED DESCRIPTION

[0021] The present application will be described in detail below through examples.

[0022] SPI (Serial Peripheral Interface) is a high-speed, full-duplex, synchronous serial communication bus, which is widely used in short-distance communication between microcontrollers (MCU), system-on-chip (SOC) and various peripheral devices (such as memory, sensors, display modules, RF chips, etc.) due to its simple protocol and high communication efficiency.

[0023] A typical SPI system consists of a master device (Master) and one or more slave devices (Slave), connected through four basic signal lines. Respectively: CLK (Clock, clock) is a clock signal generated by the master device, used to synchronize data transmission; MOSI (Master Output Slave Input, master device data output / slave device data input) is the master device data output and slave device data input line; MISO (Master Input Slave Output, master device input / slave device output) is the master device data input and slave device data output line; SS / CS (Slave Select / Chip Select, master device chip select signal) is used to select a specific slave device to communicate.

[0024] When a hardware failure occurs in a slave device connected to the SPI bus, it may force the clock line to a fixed level, i.e. the clock line is pulled dead by the device, which is a common and serious problem. This problem directly leads to the SPI host being unable to generate a correct clock, thus interrupting communication on the entire SPI bus.

[0025] In the prior art, when the SPI clock line is hung by the slave device, the transmission failure phenomenon after the occurrence of the situation is generally used for processing, i.e. the problem is indirectly inferred by detecting whether the data transmission is timed out or failed. This method is not accurate and cannot distinguish whether the transmission failure is caused by clock problem or other problems; and it is inefficient and needs to wait for the entire transmission process to timeout.

[0026] In order to solve the above-mentioned technical problems existing in the prior art, an embodiment of the present invention provides an SPI bus active anti-hanging method based on the clock line state. The embodiment of the present invention is applicable to various electronic devices that use the SPI interface for data transmission. For example, the electronic devices may include: smartphones, tablet computers, laptops, and wearable devices, etc. The embodiment of the present invention does not specifically limit the electronic devices.

[0027] The core concept of this embodiment of the present invention is to monitor the SPI clock line voltage in real time through an SPI controller (also known as an SPI IP) to see if it matches a preset idle state voltage, which can be either a high level (1) or a low level (0). This allows the SPI bus to be determined to be hung by a slave device. Upon confirming that the SPI bus is hung, the SPI controller issues an interrupt to the CPU, which then performs further operations such as resetting the abnormal device, thereby restoring the SPI bus.

[0028] The key protection points of the embodiments of the present invention are as follows: First, the SPI controller includes but is not limited to monitoring the bus, handling related logic when the bus hangs, and ensuring the correctness of judgment and reporting of related interrupts.

[0029] Secondly, after the SPI controller generates an interrupt, the processor coordinates the further software processing of the interrupt, including but not limited to: how to reset the faulty slave device and re-enable the bus to restore its normal transmission function after the corresponding interrupt.

[0030] For the sake of clarity, the overall technical solution of the embodiment of the present invention will be described below in conjunction with the timing diagram of the embodiment of the present invention.

[0031] like Figure 1 , which is a schematic diagram of a timing diagram provided by an embodiment of the present invention.

[0032] from Figure 1 As can be seen, assuming the default idle level of the CLK (clock line) is high, after several normal data transmissions, if the CLK signal is abnormally pulled low due to a peripheral failure, the SPI controller will initiate internal processing: at this point, the hang_cnt counter within the SPI controller will begin counting according to the value configured during software initialization. Here, hang_cnt can be understood as the hang count value. When the count value reaches the preset anti-hang count threshold, the SPI controller's SPI_hang (SPI hang) status signal is immediately asserted, indicating that the SPI controller has determined that the CLK signal line is in a hang state. This status can be read from the SPI controller's internal status register.

[0033] If the hang interrupt function is enabled by software (i.e. the SPI_int_en signal is set) in the SPI controller initialization phase, after the SPI_hang signal is valid, the SPI_int_req interrupt request signal will be pulled high after a certain time delay, and then an interrupt request is sent to the SoC to trigger the subsequent exception handling process. Among them, SPI_int_en is an interrupt enable signal, and SPI_int_req is an interrupt request signal.

[0034] Then the SPI controller generates the interrupt request information and reports it to the processor for subsequent processing. The processor responds to the interrupt request signal and performs corresponding fault recovery operations.

[0035] The beneficial effects brought by the embodiments of the present application include at least the following aspects: 1. Active and rapid fault diagnosis is achieved. By monitoring the clock line state in real time through the internal hardware of the SPI controller and comparing it with the preset idle level, the identification process can start at the moment when the clock line is pulled down, realizing a fundamental change from the "passive discovery after transmission failure" in the prior art to "active diagnosis when the bus is abnormal", greatly shortening the fault detection time.

[0036] 2. The system reliability and anti-interference ability are improved. The judgment mechanism based on a configurable count threshold is adopted. Only when the abnormal level lasts for more than a certain time (determined by the Hang_cnt threshold) is the hang determined, effectively filtering the transient glitches and noise interference on the bus, avoiding false alarms, and significantly improving the robustness and stability of the system.

[0037] 3. Precise fault location is achieved. The present application can clearly locate the fault source to the specific hardware problem of "clock line pulled down", rather than the general "communication failure". This provides accurate basis for subsequent execution of targeted recovery operations (such as resetting specific slave devices), avoiding the wrong operations that may be taken due to the ambiguity of fault location in the prior art.

[0038] 4. An efficient interrupt-driven response mechanism is provided. Through the hardware linkage of SPI_int_en (interrupt enable) and SPI_int_req (interrupt request) signals, under the premise of confirming the fault and enabling the interrupt, an interrupt request is automatically and timely sent to the processor. This allows the processor to quickly interrupt the current task and process the SPI bus exception, greatly optimizing the fault response speed and meeting the requirements of real-time systems.

[0039] 5. Optimized the division of labor between software and hardware, and system efficiency. The most time-consuming status monitoring and judgment logic is performed in parallel by the SPI controller hardware, with only the final interrupt result and status information submitted to the CPU software for processing. This design reduces the burden of frequent CPU polling of status registers, freeing up CPU computing power to more efficiently handle core application tasks, improving overall system efficiency.

[0040] In summary, the embodiments of the present invention provide an efficient, reliable, and real-time SPI bus fault detection and interrupt response mechanism, which fundamentally solves the problem of system paralysis caused by clock line hanging, and greatly enhances the availability and maintainability of the system.

[0041] After describing the overall technical solution of the embodiment of the present invention in conjunction with the timing diagram, the following will describe in detail an SPI bus active anti-hangup method based on the clock line state provided by the embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention provides an SPI bus active anti-hanging method based on the clock line state, which is applied to a system-on-chip (SOC). The SOC acts as an SPI master device, includes an SPI controller and a processor, and communicates with a slave device through an SPI interface. The method may include the following steps: S210: The processor configures the register of the SPI controller through software, presets the first level signal of the SPI clock line when it is in an idle state, and presets an anti-hang count threshold.

[0042] Specifically, this step is the initialization phase of the entire anti-hang mechanism and is completed by the software driver. The processor accesses the internal configuration register group of the SPI controller through the system bus.

[0043] The first level signal in the idle state can be determined by the clock polarity parameter in the SPI protocol. The software can preset the first level signal to a high or low level based on the characteristics of the communication slave device. This preset value will serve as a reference value for the subsequent hardware to determine whether the clock line status is abnormal.

[0044] The anti-hang count threshold (i.e., hang_cnt in the above embodiment) is a key parameter whose value determines the strength of the system's anti-interference capability. It is not set arbitrarily.

[0045] As an implementation of the embodiment of the present invention, the preset anti-hanging count threshold may be determined by the following formula: Anti-hang count threshold = time threshold / system clock cycle.

[0046] The time threshold is a preset minimum time width for filtering interference glitches on the clock line; and the system clock cycle is the cycle of the clock signal that drives the SPI controller.

[0047] wherein the time threshold is a time value pre-set by the software according to the electromagnetic environment the system is in and the reliability requirement, which must be greater than the width of the maximum noise glitch that can occur on the SPI clock line (e.g. 500 ns or 1 μs). The system clock period is the period of the main clock that drives the SPI controller to work. By configuring this threshold, it is ensured that only the abnormal level that lasts more than the time threshold will be determined as a fault, effectively filtering out the transient interference.

[0048] S220, the SPI controller monitors the actual level of the clock line in real time as a second level signal.

[0049] Specifically, this step is performed by a special hardware circuit inside the SPI controller, independent of the processor, realizing continuous and parallel monitoring. The actual physical level on the clock line is sampled by the digital logic of the SPI controller core through an input buffer, and for the sake of clear scheme description, the result of this real-time sampling can be called the second level signal. This monitoring is a hardware behavior, without software intervention, ensuring the real-time and high efficiency of the monitoring.

[0050] S230, when the condition that the second level signal is inconsistent with the first level signal lasts, the SPI controller enables an internal counter to count.

[0051] Specifically, the SPI controller internally contains a digital comparator, which compares the second level signal monitored in real time with the first level signal pre-set by the software. When it is detected that the two are inconsistent (for example, the pre-set is an idle high level, but the actual measurement is pulled low), an internal enable signal is activated, starting a dedicated hardware counter to begin counting. The condition lasting here is crucial: as long as the inconsistent condition exists, the counter will add 1 every system clock period; as soon as the inconsistent condition disappears (such as the interference glitch passes, the level returns to normal), the counter will be immediately reset to zero. This design ensures that only persistent abnormalities will accumulate count values, and transient noise will be ignored.

[0052] S240, when the count value of the counter reaches the anti-hang count threshold, the SPI controller determines that the clock line is hung, and sets the internal hung state flag.

[0053] Specifically, the counter compares its real-time count value with the pre-set anti-hang count threshold. When the count value is equal to or exceeds the anti-hang count threshold, it means that the duration of the abnormal level has exceeded the safety window set by the software, and the judgment logic of the SPI controller will determine that the clock line has entered the hung state. Subsequently, a hung state flag bit in a specific state register inside the SPI controller (i.e. Figure 1The SPI_hang in the embodiment will be automatically set to be valid (e.g. pulled high to 1) by hardware. The status bit can be queried by software at any time, which provides a clear hardware status indication for diagnosis.

[0054] S250, if the SPI hang interrupt function is enabled by software configuration, the SPI controller sends an interrupt request signal to the processor.

[0055] Specifically, this step is the key of automatic notification of software by hardware. Two conditions need to be met at the same time for the generation of the interrupt: 1. The hang event occurs. That is, the SPI_hang flag bit has been set, i.e. pulled high.

[0056] 2. The interrupt function has been enabled. In the initialization phase, the software has set the "hang interrupt enable bit (e.g. SPI_int_en)" in the interrupt enable register of the SPI controller to 1, which opens the switch of the interrupt output.

[0057] Only when the above two conditions are met, the interrupt logic of the SPI controller will generate an interrupt request signal (SPI_int_req) and send it to the central interrupt controller in the system-level chip. Finally, the interrupt controller will report the SPI abnormal interrupt to the processor core CPU.

[0058] As an implementation manner of the embodiment of the application, the step of enabling the SPI hang interrupt function by software configuration can include: The processor configures a specific bit of the interrupt enable register of the SPI controller through a write operation to enable the SPI hang interrupt function. Specifically, when the software writes the SPI_int_en bit (e.g. bit 2) to 1, it means that the hang interrupt function is enabled. At this time, the interrupt generation logic path in the SPI controller is completely opened. After that, once the hardware detects the clock hang event and sets the internal status flag, the interrupt control logic will immediately generate a valid interrupt request signal (SPI_int_req) and send it to the processor. Conversely, if the SPI_int_en bit is written to 0 or remains 0, even if the hang is detected, the interrupt request path will be blocked and the controller will not report the interrupt to the processor.

[0059] S260, the processor executes corresponding fault recovery operations in response to the interrupt request signal.

[0060] Specifically, after the processor receives the interrupt request signal, it will pause the current task and execute the corresponding fault recovery operation. Specifically, it can include the following steps: first, query the status register. Read the status register of the SPI controller to confirm whether it is a hang interrupt by checking the SPI_hang flag bit. Second, execute the corresponding fault recovery operation.

[0061] The technical scheme provided by the embodiment of the application can monitor the clock line state in real time through the SPI controller, compare with the preset idle level, start to identify the flow at the moment when the clock line is pulled to death, realize the fundamental change from the prior art of passive discovery after transmission failure to active diagnosis when the bus is abnormal, greatly shorten the fault detection time. Moreover, the judgment mechanism based on the configurable counting threshold is adopted, and the hanging is determined only when the abnormal level lasts more than a certain time (determined by the anti-hanging counting threshold), the transient glitches and noise interference on the bus are effectively filtered, false alarms are avoided, and the robustness and stability of the system are significantly improved.

[0062] Moreover, the application can definitely locate the fault source to the specific hardware problem of "the clock line is pulled to death", instead of the general "communication failure". This provides accurate basis for subsequent execution of targeted recovery operation (such as resetting the specific slave device), avoids the wrong operation that may be taken due to the fuzzy fault location in the prior art. And through the hardware linkage of the interrupt enable and the interrupt request signal, the interrupt request is automatically and timely sent to the processor on the premise of confirming the fault and the interrupt enable. This allows the processor to quickly interrupt the current task and process the SPI bus exception, greatly optimizes the fault response speed, and meets the requirements of real-time system.

[0063] As an implementation manner of the embodiment of the application, the processor executes the corresponding fault recovery operation, which can include the following steps, steps a1 and a2 respectively: Step a1, a reset signal for the fault slave device is generated.

[0064] The reset signal is generated in the following manner: the level of the chip select signal of the SPI bus is flipped; or a reset pulse is generated through the general input and output interface connected to the reset pin of the fault slave device.

[0065] Step a2, the reset signal is sent to the fault slave device, so that the fault slave device is reset and the pulled-to-death state of the SPI clock line is released.

[0066] Specifically, after confirming the hanging interrupt, the software executes the pre-defined recovery flow. The most common operation is to generate and send a reset signal to the fault slave device. Specifically, a series of level flips can be performed through the control of the SPI chip select signal (SS / CS), or a reset pulse can be output to the reset pin of the slave device through the configuration of a general purpose input / output (GPIO) pin.

[0067] The above fault recovery flow can bring the following significant beneficial effects: 1. The diversity and flexibility of the recovery means are realized. A variety of optional reset signal generation paths (such as chip select signal operation or dedicated GPIO control) are provided, so that the system designer can select the most suitable and effective reset mode according to the hardware characteristics and reset logic of the target slave device. This flexibility ensures that the anti-hang-up solution can be widely applied to various types of SPI slave devices, enhancing its universality and practical value.

[0068] 2. An efficient and targeted fault isolation and recovery mechanism is provided. By accurately sending a reset signal to a specific slave device that has been identified as the source of the fault, rather than resetting all devices on the system or bus, precise fault isolation and minimum range recovery are achieved. This greatly shortens the time for the system to recover from a fault, avoids unnecessary context loss or service interruption, and significantly improves the availability and service continuity of the system.

[0069] 3. The existing hardware resources are fully utilized to achieve low-cost reliable reset. The above two specific means do not require additional hardware costs. Using the chip select signal (SS / CS) for level inversion, the control signal required by the SPI protocol itself is cleverly reused, which is an extremely simple and efficient software reset method. And by generating a reset pulse through a general-purpose input / output pin (GPIO), a separate and reliable hardware reset path is provided by utilizing the widely available general-purpose resources on the SoC. Both of these methods achieve the highest reliability at the lowest cost.

[0070] In summary, the pre-defined recovery process provides a flexible, accurate, low-cost and efficient fault recovery mechanism, which is an indispensable part of the entire SPI anti-hang-up solution, ensuring that the system can be quickly and automatically repaired after encountering serious bus faults, thereby achieving the design goal of high robustness.

[0071] As an implementation manner of the embodiment of the application, after the reset signal is sent to the faulty slave device, the method can further include the following steps, steps b1 to b3: Step b1, the processor reads the hang-up state flag in the SPI controller to check whether the clock line has recovered to an idle state.

[0072] Step b2, if the hang-up state flag indicates that the clock line has returned to normal, continue the subsequent data transmission.

[0073] Step b3, if the hang-up state flag indicates that the abnormality still exists, re-execute the fault recovery operation.

[0074] Specifically, after sending the reset signal to the faulty slave device, the software reads the SPI_hang flag again to check whether the clock line has returned to normal. If it has, the interrupt is exited and normal communication continues. If it has not, the software can trigger multiple reset attempts.

[0075] The present implementation can bring the following beneficial effects: The verification and retry steps described in this paragraph bring the following significant beneficial effects: 1. A reliable fault handling closed loop is formed, ensuring the effectiveness of the recovery operation. By actively reading back the hardware status flag after performing the recovery operation for verification, this step establishes a "perform-verify-confirm" negative feedback closed loop. This ensures that the recovery operation (such as resetting the slave device) effectively resolves the clock line hang-up state, avoiding blind subsequent operations in the case of no real recovery, thereby fundamentally ensuring the reliability and effectiveness of the fault handling process.

[0076] 2. Automatic healing and rapid recovery of the system is achieved, minimizing business interruption. When the verification is successful (the status flag returns to normal), the system can seamlessly continue the subsequent data transmission automatically. This greatly shortens the duration of business interruption, meeting the stringent requirements of high continuity and real-time application scenarios.

[0077] 3. A fault tolerance level is provided, enhancing the system's ability to handle stubborn faults. When the first recovery operation fails (the status flag is still abnormal), the mechanism of re-executing the fault recovery operation provides valuable fault tolerance capability for the system. For non-permanent stubborn faults (such as devices that require multiple resets to recover), this mechanism increases the probability of eventual success through multiple attempts, thereby improving the system's resilience and robustness in non-ideal environments.

[0078] In summary, the verification and retry steps add crucial reliability guarantees and intelligent fault tolerance capabilities to the entire anti-hang mechanism, enabling it to evolve from a simple fault detector to an intelligent system that can autonomously decide, attempt multiple times, and ultimately ensure the bus returns to health. This is a key design for improving the long-term stability of the SPI bus.

[0079] Based on the above implementation, as an implementation of an embodiment of the present application, re-executing the fault recovery operation can include the following steps, steps c1 and c2 respectively: Step c1: cyclically execute the operation of generating and sending a reset signal before reaching the preset maximum number of retries.

[0080] Specifically, this step is the core embodiment of system fault-tolerant design. A maximum retry number (e.g. 3 or 5) is predefined in the software driver, which is set based on the system's evaluation of the response time to reset the slave device and the tolerance to recovery delay.

[0081] When the first recovery operation is performed, and the check finds that the hang state is still not cleared, the interrupt service program does not give up immediately, but the control flow jumps again to perform the operation of generating and sending the reset signal. This loop structure provides a second or even multiple recovery opportunities for non-fatal failures caused by signal glitches, device response delays, etc.

[0082] However, the loop operation is not infinite, and the termination condition is to reach the maximum retry number. This is a key safety mechanism to prevent the system from falling into a dead loop due to permanent physical damage to a slave device. Once the retry count reaches the maximum retry number, the loop is immediately terminated, and the program flow is transferred to step c2. This ensures the sustainability of system resources and the stability of core functions.

[0083] Step c2. After reaching the maximum retry number, if the hang state flag still indicates that the exception exists, the processor generates and records an error log information.

[0084] Among them, the error log information includes at least one of the following: fault timestamp; SPI controller number or bus identifier where the hang occurs; clock line state level at the time of hang; preset idle state level; value of hang state flag; retry number of reset operation; and identification information of the faulty slave device.

[0085] Specifically, this step is the final link of fault management, and its purpose is to provide the most detailed fault information for system maintenance personnel or upper-level diagnostic systems when automatic recovery measures are completely ineffective.

[0086] The premise of performing this step is that the hang state still exists after reaching the maximum retry number. This clearly determines that the current fault is a stubborn hardware fault that cannot be solved by software reset, for example, the slave device may have suffered irreversible physical damage. The processor will generate an error log record containing rich context. The various types of information it contains have high diagnostic value.

[0087] For fault timestamp, it accurately records the time of fault occurrence, which is used for system log analysis, problem tracing and multi-device fault association.

[0088] SPI controller number or bus identifier where the hang occurs. In a complex SOC with multiple SPI controllers, the specific bus where the fault occurs is accurately located.

[0089] Clock line state level at hang-up time and preset idle state level: record actual level (e.g. pulled low) at the time of exception and preset idle level (e.g. should be high), assist in determining fault type, e.g. short to ground or short to power.

[0090] Value of hang-up state flag and retry times of reset operation. Provide process evidence of SPI controller internal state and software recovery effort.

[0091] Identification information of the faulty slave device. This is the most critical information, directly indicating which specific device has failed, and is the core basis for hardware replacement or isolation.

[0092] After recording the error log information, the software can perform the final operation. For example: permanently disable the faulty slave device (e.g. mark it as a bad piece to avoid subsequent communication triggering hang-up again), and report the error information to the upper management unit or operating system through the system management bus, etc., which may eventually trigger a system-level alarm (e.g. turn on the fault light, send a network alarm, etc.), notifying the operation and maintenance personnel to intervene.

[0093] The implementation mode at least includes a complete and intelligent fault handling final process, and the beneficial effects brought by the implementation mode at least include the following: 1. Enhanced system fault tolerance. Through the limited retry mechanism, non-persistent faults are effectively handled, and the self-recovery ability and resilience of the system in harsh environments are improved.

[0094] 2. Avoid system death. By setting a retry upper limit, the system resources are prevented from being exhausted indefinitely when encountering permanent faults, ensuring the stable operation of core functions.

[0095] 3. Precise fault diagnosis and positioning is achieved. The detailed error log recorded provides precise remote fault diagnosis capability for technical personnel, greatly shortening the average repair time.

[0096] 4. Support for predictive maintenance. The system can long-term statistics the number of faults of a specific slave device. If a device frequently triggers this process and eventually records an error, it can be warned in advance that it will fail soon, thereby realizing predictive maintenance.

[0097] On the basis of the above embodiment, as an implementation mode of the embodiment of the application, the SPI controller can only generate and send an interrupt request signal to the processor after the SPI hang-up interrupt function is configured to be enabled by software.

[0098] And the hang-up state flag of the SPI controller is configured in one of the following clearing modes: automatically cleared by hardware logic after the processor reads the state flag, or manually cleared by the processor through software write operation.

[0099] Specifically, before the system starts and the software initialization is completed, even if there is an abnormal state of the bus, no uncontrollable interrupt request will be generated due to the fact that the interrupt is not enabled, so that the stability of the system starting process is ensured. In addition, the software can dynamically enable or disable the hang interrupt function according to different stages and modes of system operation, so that the fine and flexibility of interrupt management are realized. In addition, the separation of event occurrence and interrupt reporting makes the design of the hardware IP more clear and modular.

[0100] In addition, in the embodiment, two optional strategies for clearing the hang state flag are provided, and one of them is usually configured by parameterization or a register during IP design: The first mode is automatic clearing. This is a read-to-clear mode. When the processor reads the register containing the state flag through the bus, the hardware logic will automatically and immediately clear the flag bit. This mode is simple to operate, and the software does not need additional clearing operation, which reduces the programming complexity and the possibility of errors; and can effectively prevent the interrupt storm caused by the continuous validity of the state bit, that is, the same interrupt request is repeatedly triggered, so as to excessively occupy the CPU resources.

[0101] The second mode is manual clearing. In this mode, the state flag will always remain set until the processor explicitly performs a specific write operation (for example, writing 1 or 0 to the state flag bit) to clear it. In this mode, the flag bit will always remain set until it is explicitly cleared by the software, which provides great convenience for debugging and diagnosis. Software developers can read the register at a convenient time to accurately understand the historical fault information without losing information due to early reading. And the software can fully control the clearing time to ensure that all related processing flows (such as log recording and state reporting) are completed before the state bit is cleared.

[0102] In order to clearly describe the scheme, the overall technical scheme of the embodiment of the application will be described below in combination with a specific example. As shown in the figure, Figure 3 During the system initialization stage, the software driver needs to initialize and configure the SPI controller. In addition to communication parameters such as baud rate, clock phase and polarity, transmission bit width, etc., the configuration particularly needs to preset the following key parameters for the hang prevention mechanism described in the application: the default level of the clock line in the idle state; the hang prevention count threshold (hang_cnt); the hang interrupt enable bit (SPI_int_en).

[0103] When the clock line (CLK) is hung up due to slave device failure during data transmission of the SPI bus, the processing flow of the system completely depends on the configuration of the hang interrupt enable during the initialization stage.

[0104] 1. If the hang-up interrupt is not enabled, it indicates that the software does not pay attention to such failures. Even if the clock line is hung up, the SPI controller will not generate an interrupt, and the software will not respond to such events, and the bus communication will continue to be interrupted.

[0105] 2. If the hang-up interrupt is enabled (which is the typical configuration), the SPI controller will immediately send an interrupt request to the processor (CPU) after detecting and confirming the hang-up condition. After receiving this interrupt, the software will start a predefined fault recovery process. Common reset methods include: specific level inversion operation through the control of the chip select signal (SS / CS); or, sending a reset pulse to the reset pin of the fault slave device through the configuration of the general-purpose input-output pin (GPIO).

[0106] After the fault recovery operation is completed, the software needs to read the hang-up status flag of the SPI controller for verification. If the status flag indicates that the clock line has returned to the idle default state, it is determined that the fault has been resolved, and normal data transmission can continue. If the status flag still shows an abnormality, the software should start a retry mechanism to perform the reset operation in a loop. If the fault still exists after reaching the maximum number of retries, detailed error log information is generated and recorded, and can be reported to the upper system.

[0107] It should be noted that the hang-up interrupt function is usually enabled during the SPI initialization process. Therefore, if the clock line is hung up before the initialization is completed, the interrupt request needs to be reported to the processor after the initialization process is completed and the interrupt enable bit is valid.

[0108] In addition, the hang-up status flag can be configured in one of two clearing modes: First, automatic clearing. The hardware logic automatically clears the status flag after the processor reads it, which can effectively prevent repeated triggering of interrupts due to the persistent validity of the status bit, and avoid long-term occupation of CPU resources.

[0109] Second, manual clearing. The processor needs to explicitly clear it through a write operation, which is beneficial for the software to retain status information for diagnosis during the debugging stage.

[0110] In a second aspect, an embodiment of the present application provides an SPI bus active hang-up prevention device 40 based on the state of the clock line, which is applied in a system-on-chip (SOC), and the device is an SPI controller, as shown in Figure 4 includes: A configuration module 410 is configured to receive configuration information of a processor through a register interface, to preset a first level signal of an SPI clock line in an idle state and a hang-up prevention count threshold; A monitoring module 420 is configured to monitor an actual level of the clock line in real time as a second level signal. The counting and judging module 430 is configured to enable an internal counter to count when the condition that the second level signal is inconsistent with the first level signal lasts, and determine that the clock line is dead when the count value of the counter reaches the anti-hanging dead count threshold. The state flag module 440 is configured to set an internal dead state flag when the counting and judging module determines that the clock line is dead. The interrupt control module 450 is configured to send an interrupt request signal to the processor when the SPI dead interrupt function is enabled and the state flag module is set. The processor is configured to perform a corresponding fault recovery operation in response to the interrupt request signal.

[0111] In a third aspect, an embodiment of the present application provides an electronic device, comprising: The SPI bus active anti-hanging dead device based on the clock line state as described in the second aspect; One or more processors coupled to the device are configured to execute instructions to operate the device and perform a corresponding fault recovery operation in response to an interrupt request signal generated by the device.

[0112] In a fourth aspect, a computer readable storage medium is provided, and when instructions in the computer readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method of the first aspect.

[0113] Although the above has shown and described the embodiments of the present application, the above embodiments are exemplary, and should not be understood as limiting the present application. Those skilled in the art can make changes, modifications, replacements and variations to the above embodiments without departing from the principles and spirit of the present application, and the changes, modifications, replacements and variations are within the scope of the present application.

Claims

1. A method for actively preventing SPI bus from hanging up based on clock line status, characterized in that: Applied to a system-on-chip (SOC), the SOC acts as an SPI master device, includes an SPI controller and a processor, and communicates with a slave device via an SPI interface. The method includes: The processor configures the register of the SPI controller through software, presets the first level signal of the SPI clock line when it is in an idle state, and presets the anti-hanging count threshold; The SPI controller monitors the actual level of the clock line in real time as a second level signal; When the condition that the second level signal is inconsistent with the first level signal continues, the SPI controller enables an internal counter to count; When the count value of the counter reaches the anti-hang count threshold, the SPI controller determines that the clock line is hung and sets an internal hang state flag; If the SPI deadlock interrupt function has been enabled through software configuration, the SPI controller sends an interrupt request signal to the processor; The processor executes a corresponding fault recovery operation in response to the interrupt request signal.

2. The method according to claim 1, characterized in that The preset anti-hanging count threshold is determined by the following formula: Anti-hang count threshold = time threshold / system clock cycle; The time threshold is a preset minimum time width for filtering interference glitches on the clock line; and the system clock cycle is a cycle of a clock signal driving the SPI controller.

3. The method according to claim 1, characterized in that The step of enabling the SPI deadlock interrupt function through software configuration includes: The processor configures a specific bit of the interrupt enable register of the SPI controller through a write operation to enable the SPI deadlock interrupt function.

4. The method according to claim 1, wherein The processor performs corresponding fault recovery operations, including the following steps: Generate a reset signal for the faulty slave device; wherein the reset signal is generated by: controlling the chip select signal of the SPI bus to perform level flipping; or generating a reset pulse by controlling a general input / output interface connected to a reset pin of the faulty slave device; The reset signal is sent to the faulty slave device to reset the faulty slave device and release the dead state of the SPI clock line.

5. The method according to claim 4, characterized in that After sending the reset signal to the faulty slave device, the method further includes: The processor reads a hang status flag in the SPI controller to check whether the clock line has returned to an idle state; If the deadlock status flag indicates that the clock line has returned to normal, continue subsequent data transmission; If the hang state flag indicates that the abnormality still exists, the fault recovery operation is re-executed.

6. The method according to claim 5, characterized in that The re-execution of the fault recovery operation includes: Before reaching the preset maximum number of retries, the operation of generating and sending a reset signal is executed cyclically; After the maximum number of retries is reached, if the hang status flag still indicates that an abnormality exists, the processor generates and records an error log message; the error log message includes at least one of the following: a fault timestamp; the SPI controller number or bus identifier where the hang occurred; the clock line state level at the time of the hang; a preset idle state level; the value of the hang status flag; the number of retries of the reset operation; and identification information of the faulty slave device.

7. The method according to any one of claims 1 to 6, characterized in that: The SPI controller can generate and send the interrupt request signal to the processor only after the SPI deadlock interrupt function is configured to be enabled through software; The SPI controller's deadlock status flag is configured to be cleared in one of the following ways: automatically cleared by hardware logic after the processor reads the status flag, or manually cleared by the processor through a software write operation.

8. An SPI bus active anti-hang-up device based on clock line status, characterized in that: Applied to a system-on-chip (SOC), the device is an SPI controller, comprising: A configuration module is used to receive configuration information of the processor through a register interface to preset a first level signal and an anti-hanging count threshold when the SPI clock line is in an idle state; A monitoring module, configured to monitor the actual level of the clock line in real time as a second level signal; a counting and judging module, configured to enable an internal counter to count when the second level signal is inconsistent with the first level signal; and to determine that the clock line is hung when the count value of the counter reaches the anti-hang count threshold; A status flag module, configured to set an internal dead-end status flag when the counting and judging module determines that the clock line is dead-end; An interrupt control module, configured to send an interrupt request signal to the processor when the SPI deadlock interrupt function is enabled and the status flag module is set; The processor is configured to execute a corresponding fault recovery operation in response to the interrupt request signal.

9. An electronic device, characterized in that: include: The SPI bus active anti-hang-up device based on the clock line state as claimed in claim 8; and one or more processors, coupled to the device, configured to execute instructions to operate the device and perform corresponding fault recovery operations in response to an interrupt request signal generated by the device.

10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Serial peripheral interface (SPI) anomaly detection method and SPI anomaly detection device

    CN102841303A

  • Optical module-based fault processing method, device and optical module

    CN103795459A

  • Error signaling on individual output pins of battery monitoring system

    CN116775535A

  • Data communication method and device, electronic device and storage medium

    CN117667813A

  • Communication method, SPI host, SPI slave and SPI communication system

    CN119597689A

Cited By

  • Hardware self-recovery system and method and artificial intelligence processor

    CN121901026A

  • Hardware self-recovery system, method and artificial intelligence processor

    CN121901026B