Multi-core system security enhancement method and system based on watchdog
By using watchdog technology to monitor and detect the processor status and send the fault information to the cloud platform for repair, the problem of poor stability and security of the multi-core system is solved, real-time fault detection and repair is achieved, and the reliability and security of the system are improved.
Patent Information
- Application Number
- CN202510017918.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-06
AI Technical Summary
In the prior art, the stability and security of multi-core systems are poor, making it difficult to realize real-time fault detection and repair of each processor.
The watchdog-based multi-core system security enhancement method is adopted. By pairing multiple processors in pairs as detection groups, the first processor of each detection group monitors the second processor when performing the dog feeding operation of the watchdog timer, obtains monitoring results, and judges whether there is a fault based on the results. If it exists, obtains fault information and sends it to the cloud platform for repair.
Real-time fault detection and repair of multi-core systems is realized, the stability and security of the system are improved, the system architecture is simplified, and the sensitivity and response speed of fault detection are improved.
Smart Images

Figure CN119988072A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic equipment monitoring, and in particular to a multi-core system security enhancement method and system based on a watchdog. Background Art
[0002] Watchdog technology mainly refers to a technology used to monitor and restore the normal operation of computer systems. It is widely used in microcontrollers (MCU) and computer systems. In modern electronic devices, watchdog timers (WDT) are widely used to monitor the operating status of the system. When the system crashes unrecoverably due to faults, errors or other reasons, the watchdog timer triggers a system reset to ensure the reliability of the system.
[0003] As a means of preventing and recovering from system anomalies, watchdog technology has been widely used in single-core systems, but has not been fully utilized in multi-core systems. In modern computer systems, multi-core processors have become a standard configuration to improve computing efficiency and processing power. However, a single processor failure in a multi-core system may cause the performance of the entire system to degrade or become unstable.
[0004] Therefore, in the prior art, the stability and security of multi-core systems are relatively poor. Summary of the invention
[0005] The present invention provides a watchdog-based multi-core system security enhancement method and system, which are used to solve the defects of poor stability and security of multi-core systems in the prior art, realize real-time fault detection of each processor, can locate the faulty processor in time and repair it in time, and improve the stability and security of the multi-core system.
[0006] The present invention provides a multi-core system security enhancement method based on a watchdog, which is applied to a multi-core system. The multi-core system includes a plurality of processors, a watchdog timer corresponding to each processor, and a cloud platform. The plurality of processors are paired in pairs as a detection group. The method includes: For each detection group, the first processor monitors the state of the second processor when executing the watchdog timer feeding operation to obtain the monitoring result; The first processor determines whether the second processor fails according to the monitoring result; If the second processor fails, the first processor obtains fault information corresponding to the second processor and sends the fault information to the cloud platform; The cloud platform repairs the second processor according to the fault information.
[0007] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, when the first processor performs a watchdog timer feeding operation, the first processor monitors the state of the second processor and obtains the monitoring result, including: When the first processor performs a watchdog timer feeding operation, the first processor sends a handshake request signal to the second processor; If the first processor does not receive the response signal sent by the second processor within the predetermined response time, the monitoring result is determined to be that the second processor has a potential fault; otherwise, the monitoring result is determined to be that the second processor is in a normal state.
[0008] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, the first processor determines whether the second processor fails according to a monitoring result, including: When the monitoring result indicates that the second processor has a potential fault, the first processor attempts to resend the handshake request signal to the second processor; If the first processor does not receive a response signal sent by the second processor after sending the handshake request signal to the second processor for multiple consecutive times, it is determined that the second processor has a fault; otherwise, it is determined that the second processor is in a normal state.
[0009] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, the fault information includes: The program counter register value of the second processor, the ID number of the second processor, the fault address of the second processor, and the fault time of the second processor.
[0010] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, sending the fault information to a cloud platform includes: A strong encryption algorithm is used to encrypt the fault information, and the encrypted fault information is sent to the cloud platform through a secure communication protocol.
[0011] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, the cloud platform repairs the second processor according to the fault information, including: The cloud platform performs data analysis on the fault information to identify the fault type of the second processor; The cloud platform provides repair suggestions or automatically triggers a repair process according to the fault type of the second processor.
[0012] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, the multi-core system further includes a user end; after the cloud platform provides a repair suggestion or automatically triggers a repair process according to the fault type of the second processor, the method further includes: The cloud platform generates a fault report and determines a system status of the multi-core system; The cloud platform updates the fault information, the system status of the multi-core system, the fault report, and the repair suggestion to the user end in real time.
[0013] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, after the cloud platform updates the fault information, the system status of the multi-core system, the fault report, and the repair suggestion to the user end in real time, the method further includes: The user terminal visually displays the fault information, the system status of the multi-core system, the fault report, and the repair suggestion through a user interface.
[0014] According to a multi-core system security enhancement method based on a watchdog provided by the present invention, the method further includes: The working mode of the watchdog timer includes a first mode and a second mode; If the working mode of the watchdog timer is the first mode, a direct reset process is performed; wherein the direct reset process includes: if the processor corresponding to the watchdog timer does not perform a dog feeding operation within a dog feeding cycle, the watchdog timer directly resets and restarts the processor corresponding to the watchdog timer; If the working mode of the watchdog timer is the second mode, an interrupt reset process is performed; wherein, the interrupt reset process includes: if the processor corresponding to the watchdog timer does not perform the dog feeding operation within the dog feeding cycle, then the interrupt handling program is entered; if the processor corresponding to the watchdog timer still does not perform the dog feeding operation within the next dog feeding cycle, then the watchdog timer resets and restarts the processor corresponding to the watchdog timer.
[0015] The present invention also provides a multi-core system, the multi-core system comprising: a plurality of processors, a watchdog timer configured corresponding to each processor, and a cloud platform, wherein the plurality of processors are paired in pairs as a detection group; For each detection group, the first processor is used to monitor the state of the second processor when executing the watchdog timer feeding operation and obtain the monitoring result; The first processor is further configured to determine whether the second processor fails according to the monitoring result; The first processor is further configured to obtain fault information corresponding to the second processor if the second processor fails, and send the fault information to the cloud platform; The cloud platform is used to repair the second processor according to the fault information.
[0016] The watchdog-based multi-core system security enhancement method and system provided by the present invention pairs multiple processors into detection groups. For each detection group, two processors monitor each other and detect faults. This grouping strategy not only simplifies the system architecture, but also improves the sensitivity and response speed of fault detection. Furthermore, when the first processor determines that the second processor has a fault based on the monitoring result, it can obtain the fault information corresponding to the second processor in real time, and send the fault information to the cloud platform, so as to repair the second processor in time, thereby ensuring the performance and stability of the entire system. It can be seen that the solution of the present invention improves the reliability and security of the multi-core system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 This is one of the flow charts of the multi-core system security enhancement method based on the watchdog provided by the present invention.
[0019] Figure 2 This is the second flow chart of the watchdog-based multi-core system security enhancement method provided by the present invention.
[0020] Figure 3 It is a flowchart of the watchdog in different modes provided by the present invention.
[0021] Figure 4 It is a schematic diagram of the system architecture of the multi-core system provided by the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0024] The terms "first", "second", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, for example, they can be implemented in an order other than those given in the diagrams or descriptions of the embodiments of this application.
[0025] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such products or devices.
[0026] The term "watchdog" used in each embodiment of the present application is an important system reliability protection mechanism that can monitor the operating status of the system and automatically restart when the system fails. By configuring and controlling the watchdog, we can improve the stability and reliability of the system. In practical applications, we need to select appropriate watchdog configuration parameters according to specific needs and system characteristics, and feed the dog regularly to ensure the normal operation of the system.
[0027] Watchdog technology mainly refers to a technology used to monitor and restore the normal operation of a computer system. It is widely used in microcontrollers (MCUs) and computer systems. The application of watchdog technology includes the following aspects. Watchdog timer (WDT): A timer circuit used to prevent program dead loops or system crashes. It usually has an input terminal (feeding the dog) and an output terminal (usually connected to the reset terminal of the microcontroller). Hardware watchdog: A dedicated hardware timer is used to monitor the operation of the main program. If the timer is not reset (feeding the dog) within the set time, the timer will expire and trigger a reset signal to reset the microcontroller (MCU). Software watchdog: In some systems, the watchdog function can be implemented by software methods, such as using an idle timer / counter in a single-chip microcomputer system to design a software watchdog. Independent Watchdog Timer (IWDG): A watchdog that usually consists of a decrementing counter that generates a reset signal when the counter value is reduced to 0. The value of the counter needs to be reset before the counter is reduced to 0 to avoid reset. Window Watchdog Timer (WWDG): Unlike an independent watchdog, the window watchdog generates a reset signal when the counter value reaches the predefined window upper and lower limits without feeding the dog. The working principle of the watchdog is that when the watchdog is started, it starts counting automatically. If it is not cleared (fed) within the set time, the counter overflow will cause an interrupt or system reset. The watchdog can be reset independently of other components.
[0028] The following specific embodiments are used to describe in detail the technical solution of the present application and how the technical solution of the present application solves the above technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Figure 1-Figure 3 The present invention describes a multi-core system security enhancement method based on a watchdog.
[0029] Figure 1 This is one of the flow charts of the multi-core system security enhancement method based on the watchdog provided by the present invention. In practical applications, the multi-core system security enhancement method based on the watchdog provided by the present invention is applied to a multi-core system, which includes multiple processors, a watchdog timer configured corresponding to each processor, and a cloud platform, and multiple processors are paired in pairs as a detection group. Figure 1 As shown, the method includes steps 101 to 104.
[0030] Step 101: For each detection group, when the first processor performs a watchdog timer feeding operation, the first processor monitors the state of the second processor and obtains a monitoring result.
[0031] Specifically, a multi-core system includes multiple processors. Taking advantage of the multi-core parallel operation, multiple processors are grouped in pairs, that is, each processor is paired with another processor to form a detection group. For example, a multi-core system includes 4 processors, and these 4 processors can be divided into two detection groups. It can be understood that under this grouping method, each processor corresponds to a processor to monitor its own working status and detect faults. Therefore, this grouping strategy not only simplifies the system architecture, but also improves the sensitivity and response speed of fault detection.
[0032] Optionally, other numbers of processors may be provided for each detection group according to actual monitoring and fault detection requirements, which is not specifically limited in this embodiment.
[0033] Optionally, the pairing relationship of each processor can be dynamically adjusted to adapt to different system loads and processor performance.
[0034] The first processor and the second processor belong to the same monitoring group. Specifically, the first processor refers to the processor of the two processors in the detection group that currently executes the dog feeding operation of the watchdog timer. The second processor refers to the processor of the two processors in the detection group that is currently monitored and fault detected. In practical applications, the two processors in each detection group can be used as the first processor or the second processor, that is, the two processors in each detection group monitor and fault detect each other.
[0035] In practical applications, the two processors in each detection group can communicate with each other and send status information, handshake request signals, response signals, etc. Therefore, the two processors in each detection group can monitor each other and detect the working status.
[0036] Among them, the watchdog timer (WDT) is a timer circuit used to prevent program dead loops or system crashes. The watchdog timer has the function of preventing deadlocks. Specifically, the watchdog timer can prevent the system from falling into an infinite loop due to software errors. The watchdog timer also has a system monitoring function. Specifically, the watchdog timer can monitor the system operating status to ensure the normal operation of the system. The watchdog timer also has an automatic restart function. Specifically, if the system fails to reset the timer within the specified time, the watchdog timer will trigger a system restart.
[0037] In this embodiment, a watchdog timer is set for each processor. Each processor has an independent watchdog timer, which can monitor the status of the processor separately, thereby improving the fault tolerance of the system. If a processor fails, it will not affect the watchdog timers of other processors, thereby improving the reliability of the overall system. Each watchdog timer can be configured independently, and different watchdog feeding cycles can be set according to the specific needs and workload of each processor, thereby improving the flexibility of the system. Since a watchdog timer is set for each processor, when a failure occurs, the watchdog timer of each processor can respond quickly when an abnormality is detected, and timely triggering of a restart or error handling program can quickly locate the specific processor, thereby improving the reliability of the multi-core system. An independent watchdog timer helps to maintain the stability of the system, especially in a multi-tasking and multi-threaded environment.
[0038] In this embodiment, the dog feeding operation is a common term for the watchdog timer, which describes the process of preventing the watchdog timer from timing out and triggering a system reset. Within the preset dog feeding cycle of the watchdog timer, the system must perform a dog feeding operation to reset the counter of the watchdog timer. In practice, if the counter of the watchdog timer is not reset within the dog feeding cycle, the watchdog timer will consider that a system failure has occurred (such as a program dead loop or a system hang) and trigger a system reset or interrupt. By regularly feeding the dog, the system can ensure that the watchdog timer will not erroneously trigger a reset, thereby maintaining the normal operation of the system.
[0039] In actual applications, when the system starts, the watchdog timer is activated for initialization and starts counting down. During system operation, the processor will perform the watchdog timer feeding operation within the feeding cycle to reset the watchdog timer, thereby avoiding the watchdog timer timeout and ensuring the normal operation of the system.
[0040] In combination with the above description, the processor generally performs the dog feeding operation of the watchdog timer within a fixed dog feeding cycle. It is understandable that when the first processor performs the dog feeding operation of the watchdog timer, it monitors the state of the second processor and obtains the monitoring results, so that the fault or abnormality of the second processor can be discovered in time, thereby improving the real-time performance of the multi-core system. Furthermore, for each detection group, the two processors monitor each other and detect faults, which can avoid the collapse of the entire system caused by the failure of a single processor, thereby improving the safety and reliability of the multi-core system. In summary, combining the fault detection operation with the dog feeding operation can not only improve the reliability and stability of the system, but also simplify the system design, optimize resource utilization, and reduce the impact of faults on the business. This strategy can meet the requirements of multi-core processor systems for real-time performance and reliability.
[0041] In the present invention, the manner in which the first processor monitors the state of the second processor and obtains the monitoring results is not specifically limited. As an example, the first processor can directly read the memory mapping register of the second processor to obtain the state information, thereby performing state monitoring on the second processor and obtaining the monitoring results. As another example, the first processor and the second processor can share a memory area for exchanging state and monitoring information, and the first processor determines the state information of the second processor by accessing the shared memory area to perform state monitoring on the second processor. As another example, the first processor can monitor the performance counters of the second processor, such as the number of instruction cycles or the number of cache hits, to detect performance degradation or abnormal behavior. As another example, the multi-core system may provide an API that allows the first processor to query the health status and operating status of the second processor. As another example, the first processor can monitor the clock signal and frequency of the second processor to ensure that they are within a normal range.
[0042] Specifically, in one embodiment, Figure 2 FIG. 2 is a flow chart of the multi-core system security enhancement method based on the watchdog provided by the present invention. Figure 2 As shown, the above step 101 includes: step 201 and step 202.
[0043] Step 201: For each detection group, the first processor sends a handshake request signal to the second processor when executing the watchdog timer feeding operation.
[0044] Step 202: If the first processor does not receive a response signal sent by the second processor within a predetermined response time, the monitoring result is determined to be that the second processor has a potential fault; otherwise, the monitoring result is determined to be that the second processor is in a normal state.
[0045] In practice, the first processor sends a handshake request signal to the second processor, and the second processor returns a response signal, which is called a handshake. Exemplarily, the handshake request signal can be a status query signal, which is used to request the second processor to send its current status information, and correspondingly, the response signal is the status information of the second processor. Exemplarily, the handshake request signal can be a heartbeat signal, and correspondingly, the response signal is a heartbeat response signal.
[0046] In practical applications, the first processor and the second processor may be implemented in a variety of ways, including but not limited to: memory-mapped communication, message queues, interrupts, and dedicated communication interfaces.
[0047] Specifically, for each detection group, when the first processor performs the dog feeding operation, it sends a handshake request signal to the second processor. When the second processor is in normal working state, it will send a response signal to the first processor within a predetermined response time after the first processor sends the handshake request signal. Therefore, if the first processor does not receive the response signal sent by the second processor within the predetermined response time, it means that the second processor may have a potential fault, and at this time, the monitoring result is determined to be that the second processor has a potential fault. Relatively speaking, if the first processor receives the response signal generated by the second processor within the predetermined response time, it means that the second processor is in normal state, and at this time, the monitoring result is determined to be that the second processor is in normal state.
[0048] In this embodiment, through this handshake request and response mechanism between processors, on the one hand, the processor status can be monitored in real time, faults or abnormal behaviors can be discovered in time, and the response speed of the system can be improved; on the other hand, the system maintenance and fault diagnosis process is simplified, and the processor that may have a fault can be quickly located. Therefore, this embodiment provides an effective health monitoring method, which helps to maintain the stable operation and high performance of the system.
[0049] Step 102: The first processor determines whether the second processor fails according to the monitoring result.
[0050] In actual applications, after the first processor obtains the monitoring result, it further determines whether the second processor fails according to the monitoring result. Specifically, in one embodiment, the above step 102 includes: When the monitoring result indicates that the second processor has a potential fault, the first processor attempts to resend the handshake request signal to the second processor; If the first processor does not receive a response signal from the second processor after sending a handshake request signal to the second processor for multiple consecutive times, it is determined that the second processor has a fault; otherwise, it is determined that the second processor is in a normal state.
[0051] Specifically, when the first processor performs the watchdog timer feeding operation, it sends a handshake request signal to the second processor. If the first processor does not receive the response signal sent by the second processor within a predetermined response time, the second processor is automatically marked as having a potential fault.
[0052] Furthermore, the first processor attempts to resend a handshake request signal to the second processor. In practice, the handshake request signal can be resent to the second processor multiple times in succession, and the maximum number of transmissions can be set according to actual needs. For example, when the maximum number of transmissions is 4 and the monitoring result shows that the second processor has a potential fault, the first processor attempts to resend the handshake request signal to the second processor three times. If the first processor does not receive a response signal sent by the second processor after sending the handshake request signal to the second processor three times in succession, it is determined that the second processor has a fault; otherwise, it is determined that the second processor is in a normal state.
[0053] In this embodiment, after the first processor sends a handshake request signal to the second processor for the first time, it does not receive a response signal sent by the second processor, and the second processor is automatically marked as having a potential fault. The first processor attempts to resend a handshake request signal to the second processor, and only when the first processor does not receive a response signal sent by the second processor after sending a handshake request signal to the second processor for multiple consecutive times, is it determined that the second processor has a fault. It can be understood that by attempting to handshake multiple times, misjudgments caused by temporary communication problems or momentary processor busyness are reduced, and the fault is determined only after multiple consecutive failures to receive a response, thereby improving the accuracy of fault detection.
[0054] Step 103: If the second processor fails, the first processor obtains the fault information corresponding to the second processor and sends the fault information to the cloud platform.
[0055] In this embodiment, the cloud platform mainly refers to a remote, centralized service that can receive fault information from a multi-core processor system, perform data analysis, identify fault modes, and provide repair suggestions or automatically trigger a repair process. At the same time, the cloud platform is also responsible for updating fault information and system status to the user end in real time to ensure that the user can understand the system status in a timely manner. Such a cloud platform can improve the maintainability, reliability and user experience of the system. Exemplarily, the cloud platform can be a cloud server, a third-party service provider that provides cloud computing resources, a data center with a large number of servers and storage devices, a cloud platform that provides specific software applications, a platform that provides development and operation of applications, a cloud platform that provides big data processing and analysis capabilities, etc.
[0056] In an example, the fault information includes a program counter (PC) register value of the second processor, an ID number of the second processor, a fault address of the second processor, and a fault time of the second processor.
[0057] Among them, PC register value refers to the value of PC register. In practice, PC register is a register in computer processor, which is used to store the address of the next instruction to be executed. It is a key component of processor state, which indicates the location of current program execution. Specifically, PC register holds the address of the instruction currently being executed, or the address of the next instruction to be executed. At each clock cycle of the processor, the value of PC register is incremented to point to the next instruction, and then the instruction is read from this address and executed. At the hardware level, the value of PC register can be monitored for performance analysis and power management. It can be seen that the value of PC register is the key to understanding the current working state of the processor, which is crucial for program execution, debugging, performance analysis and system stability. In the event of a fault, the value of PC register can provide important clues about the cause and location of the fault.
[0058] In practical applications, each processor usually has a unique identifier, i.e., an ID number. This ID number helps to quickly identify the specific processor that has failed in a multi-core system. Therefore, determining the ID number of the second processor can determine the specific processor that has failed.
[0059] The fault address of the second processor refers to a memory address or an instruction address that causes the second processor to fail, which may be a memory access address that causes an exception or an instruction address that causes an error.
[0060] The failure time of the second processor may be a specific timestamp for recording the occurrence of the failure, which is helpful for analyzing the context of the failure, such as system load, executed tasks, etc., and can be used for failure trend analysis.
[0061] In actual applications, after the first processor determines that the second processor fails, the first processor will read the PC register value and other related status information of the failed processor (ie, the second processor) to provide data support for fault analysis.
[0062] In this embodiment, there is no specific limitation on the manner in which the first processor obtains the fault information corresponding to the second processor. For example, the system can be configured to automatically record the fault information corresponding to the faulty processor into a log file when a fault is detected, and the second processor can read the fault information corresponding to the second processor from the log file. For another example, special monitoring tools and software are used to track the status of the processor and capture fault information when a fault occurs. For another example, the internal state of the processor, including the PC register value and other status information, can be accessed through debugging interfaces such as JTAG and Core Sight.
[0063] Furthermore, after the first processor obtains the fault information corresponding to the second processor, the first processor sends the fault information to the cloud platform so that the cloud platform can analyze the fault information and repair the second processor in time.
[0064] Optionally, in an example, the above step 103 sends the fault information to the cloud platform, including: using a strong encryption algorithm to encrypt the fault information, and sending the encrypted fault information to the cloud platform through a secure communication protocol.
[0065] In this example, a strong encryption algorithm is used to encrypt the fault information, and the encrypted fault information is sent to the cloud platform through a secure communication protocol. This can ensure that the fault information is transmitted to the cloud platform safely and losslessly, and that the cloud platform can subsequently analyze the fault information and repair the second processor in a timely manner, thereby improving the reliability and security of the multi-core system.
[0066] Step 104: The cloud platform repairs the second processor according to the fault information.
[0067] Specifically, in one embodiment, the above step 104 includes: the cloud platform performs data analysis on the fault information to identify the fault type of the second processor; the cloud platform provides repair suggestions or automatically triggers the repair process according to the fault type of the second processor.
[0068] In actual applications, after receiving the fault information, the cloud platform can perform further data analysis to identify the fault type. For example, the fault type can be a hardware fault, a software error, an overheating problem, a power supply problem, etc. Furthermore, after the cloud platform identifies the fault type of the second processor, it can provide repair suggestions in a timely manner or automatically trigger a repair process according to the fault type of the second processor.
[0069] Specifically, according to the fault type of the second processor, the cloud platform may provide repair suggestions. Exemplarily, the repair suggestions may include operation steps, configuration changes, software updates, etc. In actual applications, after the cloud platform provides the repair suggestions, it supports the user to choose whether to repair the second processor according to the repair suggestions.
[0070] Specifically, for different fault types, corresponding repair processes can be pre-set. For some known fault types that can be automatically handled, the cloud platform can automatically trigger the repair process, such as restarting services, resetting configurations, applying patches, etc.
[0071] In this implementation, corresponding repair suggestions are provided or repair processes are automatically triggered for different fault types, which can simplify the processor repair process, save processor repair time, and improve the reliability and security of the multi-core system.
[0072] In the solution of this embodiment, multiple processors are paired in pairs as detection groups. For each detection group, two processors monitor each other and detect faults. This grouping strategy not only simplifies the system architecture, but also improves the sensitivity and response speed of fault detection. Furthermore, when the first processor determines that the second processor has a fault based on the monitoring results, it can obtain the fault information corresponding to the second processor in real time, and send the fault information to the cloud platform to repair the second processor in time, thereby ensuring the performance and stability of the entire system. It can be seen that the solution of this embodiment improves the reliability and security of the multi-core system.
[0073] In addition, in a possible implementation manner, the multi-core system further includes a user end, and after the cloud platform provides a repair suggestion or automatically triggers a repair process according to the fault type of the second processor, the method further includes: The cloud platform generates a fault report and determines the system status of the multi-core system; The cloud platform updates the fault information, system status of multi-core systems, fault reports, and repair suggestions to the user end in real time.
[0074] In actual applications, the first processor sends fault information to the cloud platform, including PC register value, processor ID number, fault address, and fault time. The cloud platform collects fault information corresponding to all faulty second processors in the multi-core system. Furthermore, the cloud platform uses its data analysis capabilities to conduct in-depth analysis of the collected fault information to identify the fault type and root cause. Based on the analysis results, the cloud platform generates a detailed fault report, including the fault type, possible cause, impact range, and recommended operation steps.
[0075] In practical applications, the cloud platform determines the system status of a multi-core system, which usually involves collecting, analyzing, and processing data from various parts of the system. For example, the cloud platform collects status data through interfaces with the multi-core system (such as APIs, network interfaces, etc.). This data may include processor load, temperature, memory usage, storage usage, network status, etc. The cloud platform integrates all the collected status data and analyzes and determines the system status of the multi-core system.
[0076] Furthermore, the cloud platform updates the fault information, system status of the multi-core system, fault reports, and repair suggestions to the user end in real time, ensuring that the user can understand the system status in a timely manner.
[0077] In this embodiment, after receiving the fault information, the cloud platform can perform further data analysis, identify the fault type, and provide repair suggestions or automatically trigger the repair process. The cloud platform updates the fault information and system status to the user end in real time, ensuring that the user can understand the system status in a timely manner, reducing the inconvenience caused by system failures.
[0078] In addition, in a possible implementation manner, after the cloud platform updates the fault information, the system status of the multi-core system, the fault report, and the repair suggestion to the user end in real time, the method further includes: The user end visualizes the fault information, system status of the multi-core system, fault reports, and repair suggestions through the user interface.
[0079] In this implementation, the user terminal provides a friendly interface to display system status, fault reports and repair suggestions, so that users can easily manage the multi-core processor system and realize human-computer interaction.
[0080] In practical applications, a watchdog in human-machine interaction (HMI) usually refers to a system-level monitoring mechanism that is used to ensure stable system operation and prevent program dead loops or system crashes. In the human-machine interaction interface, the watchdog generally does not interact directly with the user, but its role is crucial to improving the user experience. For example, the following are scenarios where watchdogs may involve human-computer interaction: system status indication. On some devices, the system status may be indicated by an LED light or display screen, which may include the status of the watchdog, such as normal operation, reset, etc.; error message prompt. If the watchdog triggers a system reset, it may be necessary to display an error message or warning on the human-computer interaction interface to inform the user that the system has automatically restarted; configuration interface. In some advanced devices, users may be allowed to configure watchdog parameters such as timeout period, dog feeding cycle, etc. through the human-computer interaction interface; log recording. The system may record logs of watchdog events and provide them to users or maintenance personnel through the human-computer interaction interface for viewing to facilitate fault diagnosis; maintenance mode. In maintenance mode, technicians can test the watchdog through the human-computer interaction interface to ensure its normal operation; safety mechanism. The watchdog can be used as part of the safety mechanism to trigger the safety program when a system abnormality is detected, and notify the user through the interface.
[0081] In addition, in a possible implementation manner, the working mode of the watchdog timer includes a first mode and a second mode. Figure 3 Schematic diagram of the watchdog flow in different modes provided by the present invention, such as Figure 3 As shown, the method further includes: step 301 to step 302.
[0082] Step 301, determining whether the working mode of the watchdog timer is the first mode; Step 302: If the working mode of the watchdog timer is the first mode, a direct reset process is performed.
[0083] The direct reset process includes: if the processor corresponding to the watchdog timer does not perform the dog feeding operation within the dog feeding cycle, the watchdog timer directly resets and restarts the processor corresponding to the watchdog timer; Step 303: If the working mode of the watchdog timer is the second mode, an interrupt reset process is executed.
[0084] Among them, the interrupt reset processing includes: if the processor corresponding to the watchdog timer does not perform the dog feeding operation within the dog feeding cycle, then enter the interrupt processing program; if the processor corresponding to the watchdog timer still does not perform the dog feeding operation within the next dog feeding cycle, then the watchdog timer resets and restarts the processor corresponding to the watchdog timer.
[0085] Specifically, in the first mode, once the watchdog timer times out, the system will be reset immediately, which is suitable for fast recovery systems, such as in safety-critical applications. The first mode does not require complex interrupt processing logic, and the system design and implementation are relatively simple.
[0086] Specifically, in the second mode, before the watchdog timer times out, the system has the opportunity to handle potential errors or exceptions through the interrupt service routine. Allow the system to perform error recovery operations such as logging, state saving, etc. before the actual reset. If the problem can be solved in time, the system can avoid unnecessary restarts, thereby reducing system downtime. For systems that require complex error handling and recovery strategies, the second mode provides more flexibility. By allowing the system to perform self-checks and state recovery before reset, the stability and reliability of the system are improved. Optionally, in the second mode, after entering the interrupt handler, a warning or notification can be provided to the user so that the user can save work or take other necessary measures.
[0087] In this embodiment, the working modes of the watchdog timer include a first mode and a second mode. The first mode provides a fast and simple system reset mechanism, which is suitable for occasions with high requirements on response time; while the second mode provides more flexibility and error handling capabilities, which is suitable for occasions that require complex error recovery strategies. Users can choose the most appropriate mode according to specific application requirements and system characteristics, which improves the flexibility of multi-core systems.
[0088] The present embodiment provides a multi-core system security enhancement method based on a watchdog, which is applied to a multi-core system. The multi-core system includes multiple processors, a watchdog timer configured corresponding to each processor, and a cloud platform. Multiple processors are paired in pairs as a detection group; the method includes: for each detection group, when the first processor performs the dog feeding operation of the watchdog timer, the state of the second processor is monitored and the monitoring result is obtained; the first processor determines whether the second processor fails according to the monitoring result; if the second processor fails, the first processor obtains the fault information corresponding to the second processor and sends the fault information to the cloud platform; the cloud platform repairs the second processor according to the fault information. The solution of this embodiment, by pairing multiple processors in pairs as detection groups, for each detection group, the two processors monitor and detect faults with each other. This grouping strategy not only simplifies the system architecture, but also improves the sensitivity and response speed of fault detection. Further, when the first processor determines that the second processor fails according to the monitoring result, the fault information corresponding to the second processor can be obtained in real time, and the fault information can be sent to the cloud platform, so as to repair the second processor in time, thereby ensuring the performance and stability of the entire system. It can be seen from this that the solution of this embodiment improves the reliability and security of the multi-core system.
[0089] The multi-core system provided by the present invention is described below. The multi-core system described below and the multi-core system security enhancement method based on the watchdog described above can be referenced to each other.
[0090] Figure 4 Schematic diagram of the system architecture of the multi-core system provided by the present invention. Figure 4 As shown, the multi-core system includes: a plurality of processors 40 , a watchdog timer 41 configured corresponding to each processor 40 , and a cloud platform 42 , wherein the plurality of processors 40 are paired in pairs to form a detection group 43 .
[0091] For each detection group 43, the first processor is used to monitor the state of the second processor when executing the dog feeding operation of the watchdog timer 41, and obtain the monitoring result; The first processor is further used to determine whether the second processor fails according to the monitoring result; The first processor is further configured to obtain fault information corresponding to the second processor if a fault occurs to the second processor, and send the fault information to the cloud platform 42; The cloud platform 42 is used to repair the second processor according to the fault information.
[0092] Specifically, the multi-core system includes a plurality of processors 40. Taking advantage of the multi-core parallel operation, the plurality of processors 40 are grouped in pairs, that is, each processor 40 is paired with another processor 40 to form a detection group 43. For example, the multi-core system includes four processors 40, and the four processors 40 can be divided into two detection groups 43. It can be understood that under this grouping method, each processor 40 corresponds to a processor 40 to monitor its own working state and detect faults. Therefore, this grouping strategy not only simplifies the system architecture, but also improves the sensitivity and response speed of fault detection.
[0093] Optionally, other numbers of processors 40 may be provided for each detection group 43 according to actual monitoring and fault detection requirements, which is not specifically limited in this embodiment.
[0094] Optionally, the pairing relationship of each processor 40 can be dynamically adjusted to adapt to different system loads and processor 40 performance.
[0095] The first processor and the second processor belong to the same monitoring group. Specifically, the first processor refers to the processor 40 of the two processors 40 in the detection group 43 that currently executes the dog feeding operation of the watchdog timer 41. The second processor refers to the processor 40 of the two processors 40 in the detection group 43 that is currently monitored and fault detected. In practical applications, the two processors 40 in each detection group 43 can be used as the first processor or the second processor, that is, the two processors 40 in each detection group 43 monitor and fault detect each other.
[0096] In practical applications, the two processors 40 in each detection group 43 can communicate with each other and send status information, handshake request signals, response signals, etc. Therefore, the two processors 40 in each detection group 43 can monitor each other and detect the working status.
[0097] Among them, the watchdog timer 41 (WDT) is a timer circuit used to prevent program dead loops or system crashes. The watchdog timer 41 has a function of preventing deadlocks. Specifically, the watchdog timer 41 can prevent the system from falling into an infinite loop due to software errors. The watchdog timer 41 also has a system monitoring function. Specifically, the watchdog timer 41 can monitor the system operating status to ensure the normal operation of the system. The watchdog timer 41 also has an automatic restart function. Specifically, if the system fails to reset the timer within the specified time, the watchdog timer 41 will trigger a system restart.
[0098] In this embodiment, a watchdog timer 41 is set corresponding to each processor 40. Each processor 40 has an independent watchdog timer 41, which can monitor the status of the processor 40 independently, thereby improving the fault tolerance of the system. If a processor 40 fails, it will not affect the watchdog timers 41 of other processors 40, thereby improving the reliability of the overall system. Each watchdog timer 41 can be configured independently, and different watchdog feeding cycles can be set according to the specific needs and workload of each processor 40, thereby improving the flexibility of the system. Since a watchdog timer 41 is set corresponding to each processor 40, when a failure occurs, the watchdog timer 41 of each processor 40 can respond quickly when an abnormality is detected, and timely triggering of a restart or error handling program can quickly locate the specific processor 40, thereby improving the reliability of the multi-core system. An independent watchdog timer 41 helps to maintain the stability of the system, especially in a multi-tasking and multi-threaded environment.
[0099] In this embodiment, the dog feeding operation is a common term for the watchdog timer 41, which describes the process of preventing the watchdog timer 41 from timing out and triggering a system reset. Within the preset dog feeding cycle of the watchdog timer 41, the system must perform a dog feeding operation to reset the counter of the watchdog timer 41. In practice, if the counter of the watchdog timer 41 is not reset within the dog feeding cycle, the watchdog timer 41 will consider that a system failure has occurred (such as a program dead loop or system hang) and trigger a system reset or interrupt. By regularly feeding the dog, the system can ensure that the watchdog timer 41 will not trigger a reset by mistake, thereby maintaining the normal operation of the system.
[0100] In actual application, when the system starts, the watchdog timer 41 is activated for initialization and starts counting down. During the operation of the system, the processor 40 will perform the feeding operation of the watchdog timer 41 within the feeding cycle to reset the watchdog timer 41, thereby avoiding the watchdog timer 41 from timing out and ensuring the normal operation of the system.
[0101] In combination with the above description, the processor 40 generally performs the dog feeding operation of the watchdog timer 41 within a fixed dog feeding cycle. It is understandable that when the first processor performs the dog feeding operation of the watchdog timer 41, it monitors the state of the second processor and obtains the monitoring result, so that the fault or abnormality of the second processor can be discovered in time, thereby improving the real-time performance of the multi-core system. Further, for each detection group 43, the two processors 40 monitor and detect faults with each other, which can avoid the collapse of the entire system caused by the failure of a single processor 40, thereby improving the safety and reliability of the multi-core system. In summary, combining the fault detection operation with the dog feeding operation can not only improve the reliability and stability of the system, but also simplify the system design, optimize resource utilization, and reduce the impact of faults on the business. This strategy can meet the requirements of the multi-core processor system for real-time performance and reliability.
[0102] In the present invention, the manner in which the first processor monitors the state of the second processor and obtains the monitoring results is not specifically limited. As an example, the first processor can directly read the memory mapping register of the second processor to obtain the state information, thereby performing state monitoring on the second processor and obtaining the monitoring results. As another example, the first processor and the second processor can share a memory area for exchanging state and monitoring information, and the first processor determines the state information of the second processor by accessing the shared memory area to perform state monitoring on the second processor. As another example, the first processor can monitor the performance counters of the second processor, such as the number of instruction cycles or the number of cache hits, to detect performance degradation or abnormal behavior. As another example, the multi-core system may provide an API that allows the first processor to query the health status and operating status of the second processor. As another example, the first processor can monitor the clock signal and frequency of the second processor to ensure that they are within a normal range.
[0103] Specifically, in one embodiment, the above-mentioned first processor is used to monitor the status of the second processor when executing the dog feeding operation of the watchdog timer 41, and when obtaining the monitoring result, it is specifically used to: when executing the dog feeding operation of the watchdog timer 41, send a handshake request signal to the second processor; if the response signal sent by the second processor is not received within the predetermined response time, then it is determined that the monitoring result is that there is a potential fault in the second processor; otherwise, it is determined that the monitoring result is that the status of the second processor is normal.
[0104] In this embodiment, through the handshake request and response mechanism between the processors 40, on the one hand, the processor 40 status can be monitored in real time, faults or abnormal behaviors can be discovered in time, and the response speed of the system can be improved; on the other hand, the system maintenance and fault diagnosis process is simplified, and the processor 40 that may have a fault can be quickly located. Therefore, this embodiment provides an effective health monitoring method, which helps to maintain the stable operation and high performance of the system.
[0105] In actual applications, after the first processor obtains the monitoring result, it further determines whether the second processor has a fault based on the monitoring result. Specifically, in one embodiment, when the first processor is used to determine whether the second processor has a fault based on the monitoring result, it is specifically used to: when the monitoring result shows that the second processor has a potential fault, the first processor attempts to resend a handshake request signal to the second processor; if the first processor does not receive a response signal sent by the second processor after sending a handshake request signal to the second processor for multiple consecutive times, it is determined that the second processor has a fault; otherwise, it is determined that the second processor is in a normal state.
[0106] In this embodiment, after the first processor sends a handshake request signal to the second processor for the first time, it does not receive a response signal sent by the second processor, and the second processor is automatically marked as having a potential fault. The first processor attempts to resend a handshake request signal to the second processor, and only when the first processor does not receive a response signal sent by the second processor after sending a handshake request signal to the second processor for multiple consecutive times, is it determined that the second processor has a fault. It can be understood that by attempting to handshake multiple times, misjudgments caused by temporary communication problems or momentary busyness of the processor 40 are reduced, and the fault is determined only after multiple consecutive failures to receive a response, thereby improving the accuracy of fault detection.
[0107] In an example, the fault information includes a program counter (PC) register value of the second processor, an ID number of the second processor, a fault address of the second processor, and a fault time of the second processor.
[0108] Furthermore, after the first processor obtains the fault information corresponding to the second processor, the first processor sends the fault information to the cloud platform 42 so that the cloud platform 42 can analyze the fault information and repair the second processor in time.
[0109] Optionally, in one example, when the first processor is used to send fault information to the cloud platform 42, it is specifically used to: use a strong encryption algorithm to encrypt the fault information, and send the encrypted fault information to the cloud platform 42 through a secure communication protocol.
[0110] In this example, a strong encryption algorithm is used to encrypt the fault information, and the encrypted fault information is sent to the cloud platform 42 through a secure communication protocol. This can ensure that the fault information is transmitted to the cloud platform 42 safely and losslessly, and that the cloud platform 42 can subsequently analyze the fault information and repair the second processor in a timely manner, thereby improving the reliability and security of the multi-core system.
[0111] Specifically, in one embodiment, when the cloud platform 42 is used to repair the second processor based on fault information, it is specifically used for: the cloud platform 42 performs data analysis on the fault information to identify the fault type of the second processor; the cloud platform 42 provides repair suggestions or automatically triggers the repair process based on the fault type of the second processor.
[0112] In this embodiment, corresponding repair suggestions are provided or the repair process is automatically triggered for different fault types, which can simplify the processor 40 repair process, save the processor 40 repair time, and improve the reliability and security of the multi-core system.
[0113] In addition, in a possible implementation, the multi-core system also includes a user end, and the cloud platform 42 is also used for: the cloud platform 42 generates a fault report and determines the system status of the multi-core system; the cloud platform 42 updates the fault information, the system status of the multi-core system, the fault report, and the repair suggestions to the user end in real time.
[0114] In this embodiment, after receiving the fault information, the cloud platform 42 can perform further data analysis, identify the fault type, and provide repair suggestions or automatically trigger the repair process. The cloud platform 42 updates the fault information and system status to the user end in real time, ensuring that the user can understand the system status in a timely manner, reducing the inconvenience caused by system failures.
[0115] In addition, in a possible implementation, the above-mentioned user terminal is used to visualize fault information, system status of the multi-core system, fault reports, and repair suggestions through a user interface.
[0116] In this implementation, the user terminal provides a friendly interface to display system status, fault reports and repair suggestions, so that users can easily manage the multi-core processor system and realize human-computer interaction.
[0117] In addition, in a possible implementation, the working mode of the watchdog timer 41 includes a first mode and a second mode. The watchdog timer 41 is used to determine whether the working mode of the watchdog timer 41 is the first mode; the watchdog timer 41 is also used to perform a direct reset process if the working mode of the watchdog timer 41 is the first mode; wherein the direct reset process includes: if the processor 40 corresponding to the watchdog timer 41 does not perform the dog feeding operation during the dog feeding cycle, the watchdog timer 41 directly resets and restarts the processor 40 corresponding to the watchdog timer 41; the watchdog timer 41 is also used to perform an interrupt reset process if the working mode of the watchdog timer 41 is the first mode. wherein the interrupt reset process includes: if the processor 40 corresponding to the watchdog timer 41 does not perform the dog feeding operation during the dog feeding cycle, then enter the interrupt handling program; if the processor 40 corresponding to the watchdog timer 41 still does not perform the dog feeding operation during the next dog feeding cycle, then the watchdog timer 41 resets and restarts the processor 40 corresponding to the watchdog timer 41.
[0118] In this embodiment, the working modes of the watchdog timer 41 include a first mode and a second mode. The first mode provides a fast and simple system reset mechanism, which is suitable for occasions with high requirements on response time; while the second mode provides more flexibility and error handling capabilities, which is suitable for occasions requiring complex error recovery strategies. Users can choose the most appropriate mode according to specific application requirements and system characteristics, thereby improving the flexibility of the multi-core system.
[0119] The present embodiment provides a multi-core system. The multi-core system includes multiple processors, a watchdog timer configured corresponding to each processor, and a cloud platform. Multiple processors are paired in pairs as a detection group; the method includes: for each detection group, when the first processor performs the dog feeding operation of the watchdog timer, the state of the second processor is monitored and the monitoring result is obtained; the first processor determines whether the second processor fails according to the monitoring result; if the second processor fails, the first processor obtains the fault information corresponding to the second processor and sends the fault information to the cloud platform; the cloud platform repairs the second processor according to the fault information. The solution of this embodiment, by pairing multiple processors in pairs as detection groups, for each detection group, the two processors monitor and detect faults with each other. This grouping strategy not only simplifies the system architecture, but also improves the sensitivity and response speed of fault detection. Furthermore, when the first processor determines that the second processor fails according to the monitoring result, the fault information corresponding to the second processor can be obtained in real time, and the fault information can be sent to the cloud platform, so as to repair the second processor in time, thereby ensuring the performance and stability of the entire system. It can be seen from this that the solution of this embodiment improves the reliability and security of the multi-core system.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-core system security enhancement method based on a watchdog, characterized in that: Applied to a multi-core system, the multi-core system includes multiple processors, a watchdog timer configured corresponding to each processor, and a cloud platform, the multiple processors are paired in pairs as a detection group; the method includes: For each detection group, the first processor monitors the state of the second processor when executing the watchdog timer feeding operation to obtain the monitoring result; The first processor determines whether the second processor fails according to the monitoring result; If the second processor fails, the first processor obtains fault information corresponding to the second processor and sends the fault information to the cloud platform; The cloud platform repairs the second processor according to the fault information.
2. The multi-core system security enhancement method based on watchdog according to claim 1 is characterized in that: When the first processor executes the watchdog timer feeding operation, the second processor is monitored for status and a monitoring result is obtained, including: When the first processor performs a watchdog timer feeding operation, the first processor sends a handshake request signal to the second processor; If the first processor does not receive the response signal sent by the second processor within the predetermined response time, the monitoring result is determined to be that the second processor has a potential fault; otherwise, the monitoring result is determined to be that the second processor is in a normal state.
3. The multi-core system security enhancement method based on watchdog according to claim 2 is characterized in that: The first processor determines, according to the monitoring result, whether the second processor fails, including: When the monitoring result indicates that the second processor has a potential fault, the first processor attempts to resend the handshake request signal to the second processor; If the first processor does not receive a response signal sent by the second processor after sending the handshake request signal to the second processor for multiple consecutive times, it is determined that the second processor has a fault; otherwise, it is determined that the second processor is in a normal state.
4. The multi-core system security enhancement method based on watchdog according to claim 1 is characterized in that: The fault information includes: The program counter register value of the second processor, the ID number of the second processor, the fault address of the second processor, and the fault time of the second processor.
5. The multi-core system security enhancement method based on watchdog according to claim 1 is characterized in that: The sending the fault information to the cloud platform includes: A strong encryption algorithm is used to encrypt the fault information, and the encrypted fault information is sent to the cloud platform through a secure communication protocol.
6. The multi-core system security enhancement method based on watchdog according to claim 1 is characterized in that: The cloud platform repairs the second processor according to the fault information, including: The cloud platform performs data analysis on the fault information to identify the fault type of the second processor; The cloud platform provides repair suggestions or automatically triggers a repair process according to the fault type of the second processor.
7. The multi-core system security enhancement method based on watchdog according to claim 6 is characterized in that: The multi-core system further includes a user end; after the cloud platform provides a repair suggestion or automatically triggers a repair process according to the fault type of the second processor, the method further includes: The cloud platform generates a fault report and determines a system status of the multi-core system; The cloud platform updates the fault information, the system status of the multi-core system, the fault report, and the repair suggestion to the user end in real time.
8. The multi-core system security enhancement method based on watchdog according to claim 7 is characterized in that: After the cloud platform updates the fault information, the system status of the multi-core system, the fault report, and the repair suggestion to the user end in real time, the method further includes: The user terminal visually displays the fault information, the system status of the multi-core system, the fault report, and the repair suggestion through a user interface.
9. The multi-core system security enhancement method based on a watchdog according to any one of claims 1 to 8, characterized in that: The method further comprises: The working mode of the watchdog timer includes a first mode and a second mode; If the working mode of the watchdog timer is the first mode, a direct reset process is performed; wherein the direct reset process includes: if the processor corresponding to the watchdog timer does not perform a dog feeding operation within a dog feeding cycle, the watchdog timer directly resets and restarts the processor corresponding to the watchdog timer; If the working mode of the watchdog timer is the second mode, an interrupt reset process is performed; wherein, the interrupt reset process includes: if the processor corresponding to the watchdog timer does not perform the dog feeding operation within the dog feeding cycle, then the interrupt handling program is entered; if the processor corresponding to the watchdog timer still does not perform the dog feeding operation within the next dog feeding cycle, then the watchdog timer resets and restarts the processor corresponding to the watchdog timer.
10. A multi-core system, characterized in that: The multi-core system includes: a plurality of processors, a watchdog timer configured corresponding to each processor, and a cloud platform, wherein the plurality of processors are paired in pairs as a detection group; For each detection group, the first processor is used to monitor the state of the second processor when executing the watchdog timer feeding operation and obtain the monitoring result; The first processor is further configured to determine whether the second processor fails according to the monitoring result; The first processor is further configured to obtain fault information corresponding to the second processor if the second processor fails, and send the fault information to the cloud platform; The cloud platform is used to repair the second processor according to the fault information.
Citation Information
Patent Citations
System and method for multiprocessor reset control and watchdog monitoring
CN114518975A
System abnormal operation processing method and device, storage medium and electronic equipment
CN117931578A
Watchdog circuit fault monitoring method
CN118626296A
Runtime Software-Based Self-Test with Mutual Inter-Core Checking
US20190278677A1
Cited By
Fatal fault isolation methods, apparatus, storage media, and products
CN122526899A