Chip operation state monitoring and self-recovery method and system

By using a mutual control and monitoring mechanism between the main control chip and the monitoring chip, the self-healing function of the charging pile system is realized, which solves the problem of system-wide failure caused by single-point failure, reduces maintenance costs, and supports remote monitoring and data collection.

CN115729782BActive Publication Date: 2025-12-12XIAN WANMA SMART NEW ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211514518.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-12-12
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

In existing charging pile systems, any single point of failure can lead to a system-wide failure. The self-healing ability is weak, requiring human intervention for recovery. Furthermore, remote monitoring and data collection are impossible when the chip is disconnected from the network, increasing labor costs.

Method used

The system employs a main control chip and a monitoring chip that act as both controllers and monitors the system. It achieves self-healing through monitoring signals, reset commands, and network firmware reflashing. The monitoring chip automatically restarts or reverts to the previous firmware version when the main control chip fails, ensuring system stability.

Benefits of technology

It enables rapid and automatic recovery from partial hardware damage, reduces maintenance costs, improves system stability and self-healing capabilities, avoids human intervention, and supports remote monitoring and data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115729782B_ABST
    Figure CN115729782B_ABST
Patent Text Reader

Abstract

The application discloses a kind of chip operating state monitoring and self-healing method and system, the method includes: in master chip and monitoring chip establish communication connection, master chip sends the care signal to the monitoring chip;Monitoring chip judges master chip state according to the care signal, when the master chip fails, monitoring chip sends modulation command to master chip, for master chip reset;Monitoring chip according to the state after reset of the master chip, corresponding network firmware is re-flashed into or the master chip of failure state;Master chip restarts master chip using network firmware, otherwise returns last version network firmware and restarts.The method and system build chip operating state monitoring and self-healing structure, one chip is master chip, another chip is monitoring chip, and the master chip and monitoring chip are mutually controlled and monitored in function, so two chips can be mutually cared, avoid single chip failure when affecting entire system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of chips, in particular to a chip running state monitoring and self-recovery method and system. BACKGROUND

[0002] Currently, the charging pile business processors are independent of each other, and only rely on serial port or can bus communication link. If a fault occurs in a processor, other business processors can only passively wait for the self-recovery of the fault processor, and the entire charging business cannot be normally performed during the waiting period. The intelligence requirement of the existing charging pile is higher and higher, and more and more modules are required in the charging pile to meet the requirement of intelligence. The increase of the modules has a certain impact on the stability of the entire pile. It is urgent to upgrade the stability of the system without increasing the maintenance cost. That is, the above-mentioned prior art has the following technical problems: any single point fault can cause the entire system to fail; the self-recovery ability is weak, and if the fault processor enters an abnormal cycle, it must be manually intervened to recover. In the case of network disconnection of the chip, the charging pile cannot establish contact with the remote terminal. The on-site debugging method of personnel is used to increase the labor cost. When the charging pile has a problem, the data log cannot be normally collected, which may cause certain hidden troubles for future upgrading. SUMMARY

[0003] One of the purposes of the present application is to provide a chip running state monitoring and self-recovery method and system. The method and system construct the structure of the chip running state monitoring and self-recovery. One of the chips is a master control chip, and the other chip is a monitoring chip. The master control chip and the monitoring chip functionally serve as the master control and the monitoring. Therefore, the two chips can monitor each other, and avoid the influence of the single chip failure on the entire system.

[0004] Another purpose of the present application is to provide a chip running state monitoring and self-recovery method and system. The method and system can perform the self-recovery function between the two chips that monitor each other. In the case of partial non-core hardware damage, the fault chip can be quickly and automatically restarted, thereby reducing the cost of chip recovery.

[0005] Another purpose of the present application is to provide a chip running state monitoring and self-recovery method and system. The method and system perform the re-flashing of the network firmware of the fault chip and the rollback operation of the last network version firmware, thereby avoiding the system problem caused by the software failure or the storage area damage of the fault chip.

[0006] In order to achieve at least one of the above-mentioned purposes, the present application further provides a chip running state monitoring and self-recovery method, which comprises:

[0007] The master control chip and the monitoring chip establish a communication connection, and the master control chip sends a care signal to the monitoring chip;

[0008] The monitoring chip judges the state of the master control chip according to the care signal, and when the master control chip is faulty, the monitoring chip sends a reset command to the master control chip for resetting the master control chip;

[0009] The monitoring chip re-flashes the network firmware to the master control chip in the fault state according to the state of the master control chip after resetting;

[0010] The master control chip restarts the master control chip by using the network firmware, or returns to the previous version of the network firmware and restarts.

[0011] According to one of the preferred embodiments of the present application, the monitoring chip receives the print information of the master control chip through the GPIO port, and the monitoring chip obtains the keyword of the master control chip according to the print information, and judges whether the master control chip is in a fault state according to the keyword.

[0012] According to another preferred embodiment of the present application, the master control chip is provided with a watchdog program, and the care signal includes a dog feeding signal sent by the master control chip to the monitoring chip.

[0013] According to another preferred embodiment of the present application, a signal interruption time threshold is set in the monitoring chip, and when the time for the monitoring chip not receiving the dog feeding signal sent by the master control chip exceeds the interruption time threshold, the master control chip is judged to be in a fault state.

[0014] According to another preferred embodiment of the present application, when the monitoring chip judges that the master control chip is in a fault state, the monitoring chip sends a reset instruction to the master control chip through the debug port, and if the master control chip resets successfully, the communication between the debug port and the master control chip is disconnected.

[0015] According to another preferred embodiment of the present application, when the monitoring chip sends a reset instruction to the master control chip, and the master control chip is still in a fault state, the monitoring chip obtains the network firmware and sends the network firmware to the master control chip through the debug port, and the master control chip re-flashes the network firmware to the corresponding storage area after obtaining the network firmware.

[0016] According to another preferred embodiment of the present application, the care signal includes a date keyword sent by the master control chip to the monitoring chip, the time stamp of the date keyword is compared with the time stamp of the monitoring chip itself, and the keyword signal amount timing sampling method is adopted to calculate whether the number of corresponding fault keywords meets the fault requirement, so as to judge whether the master control chip has a networking module fault.

[0017] According to another preferred embodiment of the present application, if the master chip networking module fails, the monitoring chip acquires the network firmware of the corresponding networking module, and then writes the network firmware of the networking module into the corresponding storage space of the master chip through the debug port, and performs a restart of the master chip.

[0018] According to another preferred embodiment of the present application, if the master chip still cannot restart after the network firmware of the networking module is written, the monitoring chip acquires the last version of the network firmware from the outside, and then writes the last version of the network firmware back into the master chip.

[0019] To achieve at least one of the above-mentioned purposes, the present application further provides a chip running state monitoring and self-recovery system, which executes the above-mentioned chip running state monitoring and self-recovery method.

[0020] The present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to execute the above-mentioned chip running state monitoring and self-recovery method. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A flowchart of a chip running state monitoring and self-recovery method is shown.

[0022] Figure 2 A schematic diagram of a chip running state monitoring and self-recovery system is shown.

[0023] Figure 3 Another schematic diagram of a chip running state monitoring and self-recovery system is shown.

[0024] Figure 4 Still another schematic diagram of a chip running state monitoring and self-recovery system is shown. DETAILED DESCRIPTION

[0025] The following description is provided to enable any person skilled in the art to make and use the present application. The preferred embodiments in the following description are only examples of the present application and modifications can be made by those skilled in the art without departing from the spirit and scope of the present application. The basic principles defined in the following description can be applied to other embodiments, modifications, improvements, equivalents and other technical solutions without departing from the spirit and scope of the present application.

[0026] It can be understood that the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of one element can be one, and in another embodiment, the number of the element can be multiple, and the term "one" cannot be understood as a limitation on the number.

[0027] Please combine Figures 1-4 The application discloses a kind of chip operating state monitoring and self-healing method and system, wherein the method mainly includes the following steps: first need to establish communication connection between main control chip and monitoring chip, wherein the modulation serial port pin, reset pin and debug port pin of main control chip and monitoring chip are connected with each other, so that the main control chip and monitoring chip can be mutually looked after.The main control chip and monitoring chip can send look after signal to each other, and main control chip or monitoring chip will judge the operating state of sender according to the look after signal, specifically, the fault type of sender can be judged according to the look after signal, and further, network firmware is brushed in and restarted according to fault type respectively to the chip of corresponding sender.

[0028] Specifically, the main control chip and monitoring chip can collect the print information of the other party through the debugging serial port, and further obtain the keyword according to the print information, wherein the keyword can analyze the mounting condition of different modules of the other party chip, such as the mounting condition of i2c bus, SPI bus, USB, CAN bus and 485 bus.For example, the keyword in the process of starting can identify the attribute behind by identifying kernel command line keyword.Kernel command line:console=ttyS0,115200earlycon earlyprintk=serial,ttyS0;

[0029] In another keyword:ignore_loglevel root= / dev / mmcblk1p2 memtest=0 rootfstype=erofs ro rootwait cma=256m;

[0030] The keyword mmcblk1:mmc1:0001 88A398 7.28GiB (reading information of the recognized mmcblk, and different partition recognition mounting is different) can be analyzed according to the keyword to obtain the fault information of the bus, and further, corresponding scripts are executed according to the abnormal information to repair the fault bus. The scripts in the application can be configured as a corresponding module network firmware flashing program. For example, when the roots are mounted in the flash during the initialization process of the chip system, the keyword information is captured during the initialization process, and it is judged according to the keyword information that the firmware area of the flash is faulty. After the monitoring chip receives the fault of the above-mentioned firmware area, the network firmware of the corresponding firmware area can be re-flashed, so that the master control chip completes the repair of the corresponding fault type. In the application, the function modules composed of the above-mentioned bus will have their own network firmware. The network firmware referred to in the application is the driver program of the corresponding device of different function modules, which is the underlying software. Different modules need to pass through the network firmware to realize the above-mentioned functions on the corresponding hardware.

[0031] The application further monitors the sending side through the debug serial port, including but not limited to the time stamp, networking module information and software running information, judges the existing function module fault according to the monitored information, respectively executes the re-flashing of the corresponding network firmware of the fault function module to realize the repair of the fault function module, and executes the restart operation of the fault chip.

[0032] Please refer to Figure 2 In one of the preferred embodiments of the application, the master control chip is connected to the monitoring interface through the GPIO port, and a watchdog program is arranged on the master control chip and the monitoring chip. The master control chip sends the dog feeding information to the monitoring chip, and the monitoring chip judges that the master control chip is in a normal state after receiving the dog feeding information sent by the master control chip. At this time, the monitoring chip does not control the master control chip through the debug port. When the master control chip is faulty, the watchdog program of the monitoring chip will be interrupted. In the application, a time threshold is set in the watchdog program of the monitoring program. When the watchdog program is interrupted, the monitoring chip judges that the master control chip is dead, and the monitoring chip sends a reset command to the master control chip through the debug port connected to the master control chip. After receiving the reset command, the master control chip executes the reset operation and restarts.

[0033] When the master chip receives the reset command, the master chip still does not recover normally, at this time, there may be a chip firmware or system damage, the monitoring chip sends the reset command again, and enters the command line mode, wherein the master chip starts to break into the command line mode after BootLoader, after the master chip enters the command line mode, the monitoring chip obtains the network firmware, and flashes the network firmware into the master chip, wherein the master chip executes the restart operation of the master chip after the network firmware is flashed into the corresponding firmware partition of the master chip, wherein the network firmware can be the driver of all functional modules of the chip itself. When the master chip restarts successfully, the monitoring chip can receive the dog feeding signal of the master chip through the watchdog program. The above-mentioned firmware flashing method can effectively solve the problem that the master chip is down due to the program exception of the firmware related program or the damage of the corresponding firmware storage area.

[0034] When the monitoring chip is abnormal due to the firmware related program, the master chip can be converted into the corresponding monitoring chip. That is, the master chip receives the dog feeding signal of the master chip through the watchdog program, when the master chip does not receive the dog feeding signal from the monitoring chip for more than a preset interruption time threshold, the master chip judges that the monitoring chip is faulty, and sends a reset command to the monitoring chip through the debug port. When the monitoring chip receives the reset command and resets successfully, the monitoring chip continuously receives the dog feeding signal from the master chip. When the monitoring chip receives the reset command and still fails to start, at this time, the monitoring chip needs to be flashed with new network firmware according to the above-mentioned network firmware flashing method, and a restart operation is performed. The above-mentioned embodiment discloses a method for solving the problem that the monitoring chip cannot restart.

[0035] Please refer to Figure 3 and Figure 4 In another preferred embodiment of the present application, the master chip or the monitoring chip may be abnormal during system operation. At this time, the master chip or the monitoring chip is still in a running state and may maintain a communication connection state. The master chip or the monitoring chip fails to mount the bus or functional module itself, resulting in abnormality of the corresponding functional module of the master chip or the monitoring chip. The present application performs the following repair operation on the chip abnormality: first, the debug interface (GPIO port) receives the print information of the sender chip, and obtains the keyword according to the print information, and analyzes the functional module abnormality problem of the sender chip through the keyword. The guardian signal in the present application is the information sent by the sender to the receiver, which carries the judgment of the bus or functional module abnormality of the sender, and the guardian signal includes the keyword information in the above-mentioned print information.

[0036] The application takes the abnormality of the networking module of the master control chip as an example: when the master control chip is damaged in the upgrading process or for other reasons, the master control chip cannot update its own time due to the damage of the networking module. Therefore, when the master control chip sends a related care signal to the monitoring chip, the timestamp information in the care signal received on the monitoring chip and the timestamp information of the monitoring chip itself will be different. Therefore, the monitoring chip extracts the timestamp information from the care signal sent by the master control chip, specifically the timestamp information in the date keyword in the master control chip, and compares the timestamp of the care signal with the timestamp of the monitoring chip. If the timestamps are inconsistent, the monitoring chip uses the signal quantity timing sampling method. When the monitoring chip obtains a certain number of keywords corresponding to the master control chip in a unit of time, it can be determined that the networking function module corresponding to the current master control chip cannot self-repair, and the current networking module is determined to be a faulty module. The signal quantity timing sampling method can also be used for fault judgment of other function modules. When it is determined that the master control chip has a fault in the related function module, the corresponding networking firmware can be downloaded from other hosts or networks, and the network firmware of the corresponding networking function module can be burned into the host through the debug port. If the networking module works normally after the corresponding network firmware is burned in, the master chip can be networked and the timestamp can be updated. Thus, the repair operation of the networking module is completed.

[0037] When the master control chip is still in an abnormal state after the corresponding network firmware is re-burned in, the timestamp received by the monitoring chip and the timestamp of the monitoring chip itself are still inconsistent. At this time, the monitoring chip returns the last major version firmware information from other hosts or network ends, and burns the last major version firmware information into the corresponding storage area of the master control chip. In the application, the last version of the entire network firmware information is burned into the entire network firmware information by using the whole system rollback method. After the last version of the entire network firmware information is burned in, the master control chip is restarted normally (all function modules are normal), and the master control chip updates the timestamp through the normal networking module. Otherwise, the monitoring chip reports hardware exception information.

[0038] The above implementation can realize mutual care of multiple chips, and enable the master control chip and the monitoring chip to self-repair each other, thereby saving the cost of manual troubleshooting and repair. In addition, the above implementation strengthens the flow of processor information in the system, and can more conveniently obtain the complete state of the system. When other chips are abnormal, the other chips can report and roll back. The above method and system can prevent the system from being unable to work normally due to chip software exception and storage area damage.

[0039] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments disclosed herein. For example, embodiments of the disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication section, and / or installed from a detachable medium. When the computer program is executed by a central processing unit (CPU), the above-described functions defined in the methods of the present application are performed. It should be noted that the computer readable medium described above in the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, in which a computer readable program code is carried. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that can send, propagate or transfer a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to, wireless, wire, optical cable, RF or the like, or any suitable combination of the above.

[0040] The computer program product of the present application can be a computer program product comprising a computer-readable medium bearing computer program code embodied therein for use with a computer. The computer program code can be code defining and / or implementing the present application. The computer program code can be written in any suitable computer readable programming language. The computer program code can be stored in a computer- readable storage medium, such as, but not limited to, any type of disk including an optical disk, a CD-ROM, a CD-R, a CD-RW, a DVD, a flash memory, a ROM, a RAM, a magnetic disk or hard drive, or any other suitable type of medium including a medium that holds the software for a particular or specialized computing purpose, or any suitable combination of media. The computer program product can be a computer program product distributed to end users, whether as a stand-alone program, as part of a physical system, or as a software download. The computer program product can be distributed on a physical medium, such as, but not limited to, a floppy disk, a CD-ROM, a CD-R, a CD-RW, a DVD, a flash memory, a ROM, a RAM, a magnetic disk or hard drive, or any other suitable type of medium, or any suitable combination of media. The computer program product can be distributed from a program distribution center, either as a tangible medium or via electronic delivery, such as from a Web site via the Internet, or from one computer to another via electronic transfer, such as by e-mail. The computer program product can be distributed in an encrypted manner, such as via encryption or via password protection.

[0041] Those skilled in the art will understand that the application described above and illustrated in the accompanying drawings is presented by way of example only and is not limiting as to the present application. The intent is to cover all modifications and alternatives of the present application falling within the scope of the application.

Claims

1. A method for monitoring and self-healing the operating status of a chip, characterized in that, The method comprises: The master chip and the monitoring chip establish a communication connection, and the master chip and the monitoring chip send a care signal to each other; The monitoring chip judges the state of the master chip according to the care signal, and when the master chip is faulty, the monitoring chip sends a modulation command to the master chip for resetting the master chip; The monitoring chip judges whether the master chip is faulty or not according to the state of the master chip after resetting; The master chip restarts the master chip by using the network firmware, or returns to the last version of the network firmware and restarts; The master chip and the monitoring chip set a watchdog program, and the care signal comprises a dog feeding signal sent by the master chip to the monitoring chip; When the time for which the monitoring chip does not receive the dog feeding signal sent by the master chip exceeds the interruption time threshold, the monitoring chip judges that the master chip is in a faulty state; when the monitoring chip judges that the master chip is in a faulty state, the monitoring chip sends a reset instruction to the master chip through a debug port, and if the master chip is successfully reset, the communication between the debug port and the master chip is disconnected; When the monitoring chip sends a reset instruction to the master chip, and the master chip is still in a faulty state, the monitoring chip acquires the network firmware and sends the network firmware to the master chip through the debug port, and the master chip re-flashes the network firmware to the corresponding storage area after acquiring the network firmware.

2. A chip operation state monitoring and self-healing method, characterized in that, The method comprises: The master chip and the monitoring chip establish a communication connection, and the master chip and the monitoring chip send a care signal to each other; The monitoring chip judges the state of the master chip according to the care signal, and when the master chip is faulty, the monitoring chip sends a modulation command to the master chip for resetting the master chip; The monitoring chip judges whether the master chip is faulty or not according to the state of the master chip after resetting; The master chip restarts the master chip by using the network firmware, or returns to the last version of the network firmware and restarts; The care signal comprises a date keyword sent by the master chip to the monitoring chip, and the number of corresponding fault keywords is calculated by comparing the timestamp of the date keyword with the timestamp of the monitoring chip itself and by using the keyword signal amount timing sampling method, so as to judge whether the master chip has a networking module fault; If the master chip has a networking module fault, the monitoring chip acquires the network firmware of the corresponding networking module, flashes the network firmware of the networking module into the corresponding storage space of the master chip through the debug port, and executes the restart of the master chip; If the master chip still cannot be restarted after the network firmware of the networking module is flashed, the last version of the network firmware is acquired from the outside world by the monitoring chip, and the last version of the network firmware is flashed into the master chip.

3. A chip operating state monitoring and self-healing system, characterized by, The chip detection system executes the chip running state monitoring and self-recovery method in any one of claims 1 or 2.

4. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the chip operation state monitoring and self-recovery method in any one of claims 1 or 2.

Citation Information

Patent Citations

  • Automotive motor controller security monitoring circuit and control method thereof

    CN104199370A

  • Network failure diagnosis method and apparatus, network device, and storage medium

    WO2021017364A1