Application program operation monitoring method and device, storage medium and electronic equipment
By monitoring and resetting the process's timer in the storage system and triggering an alarm when the process does not respond, the problem of missing storage system response is solved, improving the stability and reliability of the system.
Patent Information
- Application Number
- CN202412000537.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the storage system is lost during startup or operation due to hardware failure, external interference or operation locking, resulting in low system operation efficiency and lack of effective solutions.
It provides an application operation monitoring method, by determining the current process in the target application, resetting the process monitoring timer to the adapted process duration, and triggering a monitoring alarm when the process does not respond, thereby starting the exception processing process.
It can promptly trigger monitoring alarms and start the exception processing process when an abnormality occurs in the target process, improve the stability and reliability of the storage system and avoid business interruptions caused by the system's unresponsiveness.
Smart Images

Figure CN120045423A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers, and more particularly, to a method and apparatus for monitoring the running of an application program, a storage medium, and an electronic device. Background Art
[0002] During the operation of a storage system, it may sometimes fail to respond normally due to certain reasons. For example, during the startup of the storage system, due to a short-term hardware failure or a temporary external interference, the BIOS cannot be initialized normally, especially from the CPU initialization stage to the operating system kernel loading stage. Another example is that during the operation of the storage system, deadlocks may occur due to improper operation of locks or the peripheral hardware may not respond, or during the operation of the storage system, the kernel may freeze due to reasons such as the CPU, causing the operating system to hang and not respond to any operations.
[0003] In the related art, when the above-mentioned problem of response loss occurs in the storage system, the system will always maintain an unresponsive state until the operator discovers it and performs operations such as restarting the system to make the storage system resume normal operation. In other words, there is a technical problem of low operating efficiency of the storage system in the prior art.
[0004] For the above technical problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present application provide a method and apparatus for monitoring the running of an application program, a storage medium, and an electronic device to at least solve the technical problem of low operating efficiency of the storage system in the related art.
[0006] According to one aspect of the embodiments of the present application, a method for monitoring the running of an application program is provided, including: determining a target process currently running in a target application program; resetting a process monitoring timer configured for the target application program to a process monitoring duration adapted to the target process; and triggering a monitoring alarm when the elapsed running time of the target process reaches the reset process monitoring duration but no heartbeat monitoring signal of the target process is received.
[0007] Optionally, resetting the process monitoring timer configured for the target application program to a process monitoring duration adapted to the target process includes: when the target process is a daemon process, determining a predicted duration of tasks to be executed by the daemon process; resetting the process monitoring duration of the process monitoring timer to the predicted duration of tasks; and when the target process is an input / output control process, resetting the process monitoring duration of the process monitoring timer to a target threshold duration.
[0008] Optionally, resetting the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process includes: when receiving a process exit instruction indicating to exit the input / output control process, extending the process monitoring duration to a transition threshold duration, where the transition threshold duration is greater than the target threshold duration, and the transition threshold duration is determined based on the switching duration from the input / output control process back to the daemon process.
[0009] Optionally, after triggering the monitoring alarm, it further includes: when determining that the monitoring alarm is triggered by the daemon process, waking up the user-mode thread in the sleep state based on the monitoring alarm and restarting the hardware device where the target application is located; when determining that the monitoring alarm is triggered by the input / output control process, waking up the user-mode thread in the sleep state based on the monitoring alarm, and after closing the input / output control process through the daemon process, restarting a new input / output control process.
[0010] Optionally, before determining the target process currently running in the target application, it further includes: when the hardware device where the target application is located is started, initializing the device components and loading the bootloader into the memory; executing the bootloader and loading the operating system to be applied; when a preset monitoring character is detected in the operating system, configuring the process monitoring timer to the enabled state, where the process monitoring timer in the enabled state is in the on-timing state; when the preset monitoring character is not detected in the operating system, configuring the process monitoring timer to the disabled state, where the process monitoring timer in the disabled state is in the off-timing state.
[0011] Optionally, after initializing the device components, it further includes: through the communication link between the central processing component and the logic processing component in the hardware device, initializing and configuring the process monitoring timer among the following configuration options: when the configuration option indicates that the operating system is in the debug state, configuring the process monitoring timer to the disabled state; when the configuration option indicates that the operating system is in the crash state, configuring the process monitoring timer to the disabled state; when the configuration option indicates that the operating system is in the update state, configuring the process monitoring timer to the disabled state; when the configuration option indicates that the process monitoring timer is in the uninstalled state, configuring the process monitoring timer to the disabled state.
[0012] According to another aspect of the embodiments of the present application, there is provided a method and apparatus for monitoring the running of an application program, including: a determination unit for determining a target process currently running in a target application program; a reset unit for resetting a process monitoring timer configured for the target application program to a process monitoring duration adapted to the target process; a trigger unit for triggering a monitoring alarm when the running duration of the target process reaches the reset process monitoring duration but no heartbeat monitoring signal of the target process is received.
[0013] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the above-mentioned method for monitoring the running of an application program when running.
[0014] According to still another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for monitoring the running of an application program as described above.
[0015] According to still another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned method for monitoring the running of an application program through the computer program.
[0016] Through the above implementation manners of the present application, first, a target process currently running in a target application program is determined; then, a process monitoring timer configured for the target application program is reset to a process monitoring duration adapted to the target process; further, when the running duration of the target process reaches the reset process monitoring duration but no heartbeat monitoring signal of the target process is received, a monitoring alarm is triggered. It can trigger a monitoring alarm in time when the target process has an abnormality and start an exception handling process without waiting for manual intervention, thereby effectively improving the stability and reliability of the storage system and solving the technical problem of low running efficiency of the storage system in the related art. Description of the Drawings
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is a hardware structural block diagram of a server device for an optional method for monitoring the running of an application program according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of an optional method for monitoring the running of an application according to an embodiment of the present application;
[0020] Figure 3 is another flowchart of an optional method for monitoring the running of an application according to an embodiment of the present application;
[0021] Figure 4 is yet another flowchart of an optional method for monitoring the running of an application according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of an optional method for monitoring the running of an application according to an embodiment of the present application;
[0023] Figure 6 is yet another flowchart of an optional method for monitoring the running of an application according to an embodiment of the present application;
[0024] Figure 7 is a schematic diagram of the structure of an optional device for monitoring the running of an application according to an embodiment of the present application;
[0025] Figure 8 is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0026] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 is a hardware structure block diagram of a server device for a data processing method according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0030] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to a method for adjusting the read voltage in a memory according to an embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0031] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0032] As an alternative embodiment, the method for monitoring the running of the above application program is as Figure 2 shown, including:
[0033] S202. Determine the target process currently running in the target application;
[0034] S204. Reset the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process;
[0035] S206. Trigger a monitoring alarm when the running duration of the target process reaches the reset process monitoring duration but no heartbeat monitoring signal of the target process is received.
[0036] It should be noted that the above steps S202 to S206 are a monitoring method that can timely detect and respond to the abnormal state of an application or process for alarm without human intervention. Specifically, first execute the above step S202 to determine the target process currently running in the target application.
[0037] For the above step S202, it can be understood that in the operating environment of a storage system, there are usually multiple applications and processes. Different applications and processes may perform different tasks and have different running cycles and performance characteristics. To effectively monitor and prevent the abnormal freezing of an application, the target application directly related to the storage service and its running target process in the current system can be determined first, and no specific limitation is made here.
[0038] Optionally, the determination of the above target process can be based on the process ID. For example, the process list of all currently running processes can be obtained first. After obtaining the process list, the target process in the target application can be further filtered out. Specifically, according to the context information of the process, such as the name of the application to which it belongs, the process ID (PID), or a specific tag, the target process is identified, so as to implement the determination of the target process currently running in the target application in the above step S202.
[0039] After determining the target process, then execute the above step S204 to reset the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process. That is to say, once the target process is determined, the next step is to adjust the above process monitoring timer so that its duration can match the running characteristics of the target process. It should be noted that the above process monitoring timer can be a timer in the CPU of the storage device or other chips in the storage device, such as a timer in a CPLD (Complex Programmable Logic Device). No specific limitation is made here.
[0040] It should be noted that the above process monitoring timer is used to monitor the running status of the target process within a set period of time. If no heartbeat monitoring signal is received from the target process during this period, it is considered that the above target process may have an abnormality and corresponding handling measures need to be taken.
[0041] It is worth noting that different differential handling measures can be taken for different types of the above target processes. Specifically, the above step S204 may include:
[0042] S204-1, when the target process is a daemon process, determine the predicted duration of the task to be executed by the daemon process; reset the process monitoring duration of the process monitoring timer to the task predicted duration;
[0043] S204-2, when the target process is an input / output control process, reset the process monitoring duration of the process monitoring timer to the target threshold duration.
[0044] The following will explain the two sub-steps of S204-1 and S204-2 in detail. For the above step S204-1, it should be noted that the above daemon process is responsible for executing background tasks and has the characteristic of continuous operation to ensure the normal operation of specific services or functions. For example, in a storage system, the daemon process can be responsible for operations such as regular data backup and continuous monitoring of the system status. Since the execution time of the daemon process can be predicted, a reasonable monitoring duration can be determined by analyzing the task characteristics.
[0045] Specifically, first, identify the tasks to be executed by the daemon process, including but not limited to data read and write operations, system status checks, resource allocation, etc. Then, analyze its execution time based on the characteristics of the task. For example, the execution time can be analyzed from the following aspects:
[0046] Average execution time, calculate the average duration of task execution based on historical operation data;
[0047] Maximum execution time, considering the possible delays or abnormalities during task execution, determine the maximum execution duration of a task;
[0048] Priority or importance of the task, allocate a shorter monitoring duration for high-priority or important tasks to ensure a quick response to their potential abnormalities.
[0049] Furthermore, based on the analysis of the above three aspects, set a reasonable time threshold as the above task predicted duration to ensure that the task can be completed within the monitoring duration under normal circumstances and that timeout abnormalities can be detected in a timely manner.
[0050] Furthermore, after determining the task prediction duration of the daemon process, the next step is to reset the process monitoring duration of the above-mentioned process monitoring timer to this task prediction duration. It can be understood that setting the timeout of the timer according to the predicted execution time of the next task of the daemon process can accurately monitor the execution of the task. In addition, this setting is completed before the daemon process starts to execute a new task to ensure that the monitoring duration matches the actual requirements of the task.
[0051] Regarding step S204-2 above, it should be noted that the above input / output control process is responsible for operations such as data reading, writing, and transmission in the storage system. The execution time of the above input / output process is usually affected by various factors such as external hardware status, network latency, and resource contention. Therefore, its execution time may be difficult to accurately predict. Thus, the strategy of using the target threshold duration in step S204-2 as the monitoring duration can be adopted.
[0052] Specifically, in step S204-2 above, the above target threshold duration can be set based on an estimate of the potentially longest execution time of the above input / output control process, such as 1.5 seconds. It can be understood that the above input / output control process can respond and complete tasks within a short time (such as the above target threshold). Therefore, a delay exceeding the target threshold duration may indicate potential hardware failures, network problems, or resource bottlenecks.
[0053] Furthermore, when the target process is the above input / output control process, the monitoring duration of the process monitoring timer is fixedly reset to the set target threshold duration. During the execution of the above input / output control process, regardless of the characteristics of the task currently being executed by the above input / output control process, its running state will be monitored at fixed time intervals (the target threshold duration).
[0054] Through the above steps S204-1 to S204-2, it is possible to adopt a differentiated monitoring strategy for different types of processes, that is, for the daemon process, the monitoring duration is dynamically adjusted based on the task prediction duration; while for the input / output control process, a fixed target threshold duration is set. This not only takes into account the characteristics and complexity of task execution but also takes into account the efficiency and accuracy of system monitoring, thereby improving the stability and reliability of the entire storage system in the unattended situation.
[0055] Furthermore, after resetting the process monitoring timer configured for the target application to the process monitoring duration adapted to the target process through the above step S204, execute the above step S206. When the running duration of the target process reaches the reset process monitoring duration but the heartbeat monitoring signal of the target process has not been received, trigger a monitoring alarm.
[0056] Specifically, for the above-mentioned step S206, when it is monitored that the running duration of the target process reaches or exceeds the reset process monitoring duration, and no heartbeat monitoring signal is received from the target process during this period, this usually means that the target process may have stopped responding or entered an abnormal state. At this time, a monitoring alarm is triggered and an exception handling process is started to avoid service interruption caused by process exceptions.
[0057] Optionally, the above-mentioned monitoring alarm includes but is not limited to sending an alarm message to the system log, sending an alarm signal to a specific application or service, collecting exception information for fault diagnosis, notifying the system administrator, or executing a custom exception recovery policy, etc., which are not specifically limited here.
[0058] Through the above implementation method, first determine the target process currently running in the target application; then reset the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process; further, when the running duration of the target process reaches the reset process monitoring duration but no heartbeat monitoring signal of the target process is received, trigger a monitoring alarm. It can timely trigger a monitoring alarm and start an exception handling process when the target process is abnormal, without waiting for manual intervention, thereby effectively improving the stability and reliability of the storage system and solving the technical problem of low running efficiency of the storage system in the related technology.
[0059] As an optional implementation method, resetting the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process includes:
[0060] S1, when a process exit instruction for indicating to exit the input / output control process is received, extend the process monitoring duration to a transition threshold duration, where the transition threshold duration is greater than the target threshold duration, and the transition threshold duration is determined based on the switching duration from the input / output control process back to the daemon process.
[0061] It can be understood that, as described above, during the operation of the storage system, the above-mentioned input / output control process is responsible for executing tasks closely related to hardware interaction, such as data reading, writing, cache management, etc. When the above-mentioned input / output control process is running, it is monitored by a preset target threshold duration in order to respond quickly in the case of abnormal timeout of the input / output control process.
[0062] However, it should be noted that when the above-mentioned input and output control process exits normally or is forced to exit due to a fault, control needs to be smoothly returned to the daemon process to maintain the stable operation of the storage system. The return process involves operations such as resource release and state synchronization of the input and output control process, and restart and state recovery of the daemon process, and these operations themselves also take a certain amount of time. Therefore, simply monitoring according to the target threshold duration may lead to false alarms, that is, when the input and output control process is performing a normal exit operation, it will be mistakenly considered that the system is abnormal due to the short monitoring duration. Therefore, it is necessary to extend the above-mentioned process monitoring duration to the above-mentioned transition threshold duration when the exit instruction of the input and output control process is received.
[0063] Specifically, when an instruction is detected that the above-mentioned input and output control process is about to exit (it can be an active exit from within the process, or it can be an exit command issued by an external manager or daemon process, which is not specifically limited here), the switching time from the input and output control process back to the daemon process is predicted based on historical operating data, current system load conditions, and the specific type and status of the input and output control process.
[0064] Furthermore, based on the predicted result of the switching duration, the transition threshold duration is calculated. It should be noted that the transition threshold duration should be slightly longer than the predicted switching duration to ensure that no abnormal response is triggered due to the monitoring duration before the resources of the input / output control process are completely released and smoothly switched to the daemon process.
[0065] Through the above implementation, when a process exit instruction for instructing to exit the input / output control process is received, the process monitoring duration is extended to the transition threshold duration, thereby avoiding false alarms triggered by normal system switching. At the same time, it also ensures that during the input / output control process exit, the storage system can be given sufficient time for self-adjustment and recovery, thereby improving the stability and reliability of the storage system when facing process switching.
[0066] As an optional implementation, after the monitoring alarm is triggered, the method further includes:
[0067] S1, when it is determined that the daemon process triggers the monitoring alarm, wake up the user-mode thread in the sleeping state based on the monitoring alarm, and restart the hardware device where the target application is located;
[0068] S2, when it is determined that the input / output control process triggers the monitoring alarm, wake up the user-mode thread in the sleeping state based on the monitoring alarm, and after closing the input / output control process through the daemon process, restart the new input / output control process.
[0069] It is understandable that the above steps S1 to S2 are response steps after triggering the monitoring alarm. The following is a specific description of the above steps S1 to S2.
[0070] Regarding the above step S1, it should be noted that the daemon process is responsible for providing core services in the storage system and the continuous operation and monitoring of the overall storage system. Once it is detected that the daemon process triggers the monitoring alarm, it means that the daemon process has entered an abnormal state and may not be able to continue to execute its tasks normally, such as data management, system status monitoring, etc. In this case, emergency measures need to be taken to prevent a wider range of service interruptions or data loss.
[0071] Specifically, based on the monitoring alarm, the user-mode threads in the sleep state are awakened. It should be noted that in the case of the daemon process being abnormal, there is usually one or more user-mode threads in the sleep state in the storage system, waiting for signals or instructions from the daemon process. The above user-mode threads can be threads for monitoring system status, handling abnormal situations, or executing specific tasks. When the monitoring alarm is triggered, the user-mode threads in the sleep state will be awakened and respond quickly to execute the predefined exception handling logic.
[0072] In addition, more importantly, when the daemon process triggers the monitoring alarm, measures will be taken to restart the hardware device to completely eliminate the system instability state that may be caused by the abnormal daemon process, and restore the normal operation of the entire system through a hardware-level restart.
[0073] It should be noted that restarting the hardware device will reset the software states of all software running on the hardware, including the operating system, application programs, and their runtime environments. The operation of restarting the hardware can ensure that the storage system starts running from a normal initial state, thereby improving the stability of the system.
[0074] Regarding the above step S2, in the case where it is determined that the input / output control process triggers the monitoring alarm, based on the monitoring alarm, the user-mode threads in the sleep state are awakened, and after the input / output control process is closed by the daemon process, a new input / output control process is restarted.
[0075] It should be noted that in the case where it is determined that the input / output control process triggers the monitoring alarm, similar to the above step S1, the user-mode threads in the sleep state will also be awakened to execute further exception handling logic. Different from the above step S1, in step S2, the input / output control process is closed by the daemon process.
[0076] Specifically, since the input / output control process and the daemon process work in cooperation, when shutting down an abnormal input / output control process, operating through the daemon process can ensure the correct release of resources and the synchronization of states. After triggering the monitoring alarm, the daemon process will be responsible for shutting down the abnormal input / output control process, including but not limited to operations such as releasing the resources occupied by the input / output control process, updating the system state, and recording abnormal information, which are not specifically defined herein.
[0077] Furthermore, after shutting down the abnormal input / output control process, a new input / output control process will be restarted by the daemon process to resume the normal operation of functions such as data reading and writing, cache management, etc. Optionally, the process of the daemon process restarting a new input / output control process can include operations such as initializing the new input / output control process by the daemon process, resource configuration, and state synchronization, to ensure that the new process can immediately take over the work of the abnormal process without causing service interruption.
[0078] Through the above implementation manners, for the abnormality of the daemon process, measures of restarting the hardware device are taken to completely eliminate the unstable state of the system; while for the abnormality of the input / output control process, the daemon process is used to shut down the abnormal process and restart a new process, minimizing the time and scope of influence of service interruption as much as possible. This enhances the stability and reliability of the storage system in unattended situations, and at the same time provides more accurate means of abnormal positioning and handling for system administrators.
[0079] As an optional implementation manner, before determining the target process currently running in the target application program, it further includes:
[0080] S1, when the hardware device where the target application program is located is started, initialize the device components and load the bootloader into the memory;
[0081] S2, execute the bootloader and load the operating system to be applied;
[0082] S3, when a preset monitoring character is detected in the operating system, configure the process monitoring timer to the enabled state, where the process monitoring timer in the enabled state is in the on-timing state;
[0083] S4, when the preset monitoring character is not detected in the operating system, configure the process monitoring timer to the disabled state, where the process monitoring timer in the disabled state is in the off-timing state.
[0084] It should be noted that before determining the target process currently running in the target application program, the storage system and the hardware device where it is located will be initialized and configured to ensure that the monitoring mechanism can be started accurately and timely and adapted to a specific operating environment.
[0085] Specifically, for the above step S1, when the hardware device where the target application is located is started, the device components are initialized, and the bootloader is loaded into the memory. It should be noted that when the hardware device of the storage system is powered on and starts to boot, the device initialization process is executed first. The device initialization process includes, but is not limited to, the initialization of key hardware components such as the CPU, memory, and I / O controller. Subsequently, the bootloader, such as GRUB (Grand Unified Bootloader), will be loaded into the memory.
[0086] Furthermore, execute the above step S2, execute the bootloader, and load the operating system to be applied. Specifically, after the bootloader is loaded into the memory, it will start to execute its functions. For example, first, it performs a power-on self-test (POST), and then, according to the preset boot order, selects a device from the available boot devices and loads the operating system kernel into the memory. It can be understood that the above step S2 is the transition of the storage system from startup to operation. Once the operating system is loaded, it will take over the control of the hardware resources and start to execute system initialization and the startup of user applications.
[0087] Further, for the above steps S3 to S4, it should be noted that in order to adapt to different operating environments, there is a detection mechanism for identifying whether the operating system needs to enable the process monitoring timer. Specifically, it is judged by detecting a preset monitoring character. In the specific environment of the storage system, if the operating system detects this preset monitoring character at startup, it means that the system needs to enable the process monitoring timer to enhance the stability and response ability of the system. At this time, the process monitoring timer is configured to the enabled state, that is, the timer starts timing, and the process monitoring timer enters the working state and starts to monitor the running states of the daemon processes and input / output control processes related to the storage service.
[0088] If the operating system fails to detect the preset monitoring character at startup, this may mean that the currently running operating system is not the dedicated operating system of the storage system, or the system administrator has chosen not to enable the process monitoring timer function. In this case, the process monitoring timer will be configured to the disabled state, that is, the process monitoring timer will not start timing.
[0089] Through the above implementation manner, the intelligent initialization and configuration of the process monitoring timer are realized in the operating environment of the storage system, ensuring the timely enabling of the process monitoring timer and its adaptive adjustment for specific scenarios, thereby improving the stability and reliability of the entire system in unattended situations, and at the same time providing convenience for system management and maintenance.
[0090] As an alternative implementation, after initializing the device components, it further includes:
[0091] S1. Through the communication link between the central processing component and the logic processing component in the hardware device, initialize the configuration of the process monitoring timer among the following configuration options:
[0092] S1-1. When the configuration option indicates that the operating system is in the debug state, configure the process monitoring timer to the disabled state;
[0093] S1-2. When the configuration option indicates that the operating system is in the crash state, configure the process monitoring timer to the disabled state;
[0094] S1-3. When the configuration option indicates that the operating system is in the update state, configure the process monitoring timer to the disabled state;
[0095] S1-4. When the configuration option indicates that the process monitoring timer is in the uninstalled state, configure the process monitoring timer to the disabled state.
[0096] It should be noted that after the initialization of the storage system hardware device, it is necessary to intelligently configure the process monitoring timer according to the current state and configuration options of the storage system to adapt to different operating environments and requirements. Specifically, for the above step S1, after the initialization of the hardware device is completed, that is, after components such as the CPU, memory, and I / O controller are configured to a known state, the next step is to initialize the configuration of the process monitoring timer through the communication link between the central processing component (CPU) and the logic processing component (such as CPLD) in the hardware device. Specifically, it is based on the current state of the operating system and the set configuration options. The following specifically describes the specific situations indicated by the configuration options in steps S1-1 to S1-4.
[0097] For step S1-1, when the configuration option indicates that the operating system is in the debug state, configure the process monitoring timer to the disabled state. It should be noted that when the operating system is in the debug mode, the system may perform operations such as exception monitoring, performance testing, or development debugging. These operations may take a long time, and the system behavior may not follow the regular response time. To avoid accidentally triggering the monitoring alarm due to the irregular system behavior during debugging, the process monitoring timer is set to the disabled state, that is, the process monitoring timer will not start timing, and the process monitoring timer will not perform timeout monitoring on the processes in the system in this state, thus providing a timeout-free environment for the debug personnel.
[0098] For step S1-2, when the configuration option indicates that the operating system is in a crashed state, configure the process monitoring timer to the disabled state. It should be noted that if the operating system detects a crashed state, there may be unpredictable system behaviors at this time due to reasons such as hardware failures, software errors, or system resource exhaustion. Therefore, when the system crashes, the process monitoring timer may have lost its monitoring significance because the system may not be able to respond to any normal timeout recovery mechanism. Configuring the process monitoring timer to the disabled state can avoid meaningless monitoring attempts by the system in the crashed state.
[0099] For step S1-3, when the configuration option indicates that the operating system is in an update state, configure the process monitoring timer to the disabled state. It should be noted that operating system updates are a normal maintenance activity, and during the operating system update, it may involve downtime upgrades of system software or hardware drivers. The operating system update takes a certain amount of time to complete, and during this period, it may not be able to respond to normal business requests. Disabling the process monitoring timer can ensure that the update process is not affected by timeout monitoring, thus avoiding the situation of restarting or terminating critical processes due to misjudgment during the update process and ensuring the smooth progress of the update operation.
[0100] For step S1-4, when the configuration option indicates that the process monitoring timer is in an uninstalled state, configure the process monitoring timer to the disabled state. It should be noted that when the software watchdog module is uninstalled by the user or system administrator, it means that the system no longer requires the operation of the process monitoring timer. In this case, configuring the process monitoring timer to the disabled state, that is, turning off the monitoring of the daemon process and the input / output control process, avoids misoperations or resource waste that may be caused by the remaining monitoring mechanism after uninstallation.
[0101] Through the above embodiments, the process monitoring timer can intelligently adapt to various operating environments in the storage system, including scenarios such as debugging, crashing, updating, and module uninstallation. It not only improves the stability and security of the system but also provides users with flexible configuration options. Users can adjust the operating state of the process monitoring timer according to actual needs, thereby reducing misoperations and resource consumption that may be caused by system state changes while ensuring the normal operation of the system, and enhancing the self-management ability of the storage system in complex environments.
[0102] Figure 3 is another optional flowchart of the method for monitoring the operation of an application according to an embodiment of the present application. To address the situation where the storage system is abnormally stuck during the startup and operation phases, the watchdog module 302 in Figure 3 can wake up the user-mode waiting-for-dog-bark thread 304 to simulate dog barking and achieve abnormal alarm.
[0103] Specifically, the above watchdog module 302 consists of a timer and a wait queue. The watchdog module 302 provides the functions of setting the watchdog timeout, feeding the dog, and waking up the thread after the dog barks to the outside. The main implementation method is that after loading the dog-feeding module, the wait queue is initialized and the timer is started; when the timer reaches the set time, the wake-up flag is set to wake up the wait queue, and after the wait queue is awakened, the thread is awakened to simulate the dog barking.
[0104] Specifically, in S302, the timer starts timing.
[0105] In S304, wait for timeout.
[0106] In S306, reset the timer timeout time. If the timer timeout time is reset, go back to step S302; if the timer timeout time is not reset, execute step S308.
[0107] In S308, let the timer start timing again. If the timer can start timing again, go back to step S302; if it cannot start timing again, execute step S310.
[0108] In S310, check if it times out. If it times out, execute step S312; if it does not time out, go back to step S304.
[0109] In S312, check if the wait queue is awakened. If not, execute step S314.
[0110] In S314, the thread sleeps.
[0111] In S316, check if the thread is awakened. If so, execute step S318; if not, go back to step S314.
[0112] In S318, the dog barks, and handle the dog barking logic.
[0113] It can be understood that the entire watchdog module 302 provides two ways of feeding the dog to the outside:
[0114] Method 1: Reset the timer timeout time and let the timer start timing again. This is the way for the daemon process to feed the dog.
[0115] Method 2: Set to start timing again and let the timer start timing again. This is the way for the I / O process to feed the dog.
[0116] It should be noted that the user-mode process is in a sleep state through a separate thread and waits to be awakened after the dog barks to implement the watchdog timeout operation. At this time, there are also two ways:
[0117] Method 1: If the I / O process times out, the daemon process will kill the process and restart a new process.
[0118] Method 2: If the daemon process times out, the device will be restarted.
[0119] It should be noted that the storage system's service first starts the daemon process, and then the daemon process starts the I / O process. Through the watchdog module 302 in Figure 3 , it monitors whether the daemon process and the I / O process are stuck due to deadlocks or system call exceptions. Specifically, first, during the stage from the daemon process to starting the I / O process, when the daemon process predicts that it will take time T1 to complete job1 (task 1), it will reset the watchdog timeout to T1. Similarly, when it plans to complete job2 (task 2), it will set the time T2 as the new timeout. When a task exceeds the planned time, the watchdog timeout will be triggered.
[0120] Furthermore, after the I / O process is started, the I / O process takes over the control of the software watchdog, and the daemon process no longer feeds the dog. At the same time, the timeout is set to a fixed value, such as 1.5s, and the timer is reset every 100ms to simulate feeding the dog once every 100ms. If the dog is not fed for 1.5s, the watchdog timeout will be triggered.
[0121] During the abnormal / normal exit stage of the I / O process, it should be noted that before the I / O process exits, the watchdog timeout needs to be extended, for example, from 1.5s to 20s, so as to leave sufficient time for the entire I / O process to exit and the daemon process to take over the watchdog, preventing the 1.5s from being too short and accidentally triggering the watchdog timeout. After the daemon process takes over the watchdog, it continues to feed the dog (reset the timer) in the same way as in the stage from the daemon process to starting the I / O process. For the working process of the overall watchdog module, see Figure 4 .
[0122] Specifically, S402, the daemon process executes.
[0123] S404, initialize the software watchdog. After step S404, step S406 and step S416 are executed in parallel.
[0124] S406, the daemon process feeds the dog. Ensure that the watchdog module 402 does not wake up the user-mode thread.
[0125] S408, start the I / O process.
[0126] S409, reset the watchdog timeout to 1.5s.
[0127] S410, the I / O process feeds the dog.
[0128] S412, the I / O process takes over the watchdog and starts feeding the dog.
[0129] S414, Watchdog timeout.
[0130] S416, Create a new thread and wait for the watchdog to time out.
[0131] S418, The thread sleeps.
[0132] S420, Check if the thread is awakened. If yes, execute step S422; if no, go back to step S418.
[0133] S422, Check if there is an I / O process. If yes, execute step S424; if no, execute step S428.
[0134] S424, Send a SIGCHLD signal to the I / O process to terminate the I / O process.
[0135] S426, Restart the I / O process.
[0136] S428, Restart the storage system.
[0137] It should be noted that there is a CPLD (Complex Programmable Logic Device) inside the storage device. The CPLD has sufficient logic resources and timer functions, and can implement the watchdog timing and device reset functions without adding additional hardware resources. Specifically, it is mainly achieved by using the GPIO pins reserved between the CPU and the CPLD. Among them, C0 is the GPIO from the CPU to the CPLD direction and serves as a switch. When this pin is set to "low", it means open, and the CPLD enables the timing function; when set to "high", it means closed, and the CPLD disables the timing function. C1 is the pulse output from the CPU to the CPLD, and it simulates feeding the dog through a pulse signal with 1 second high level and 1 second low level. For the specific process, please refer to Figure 5 。
[0138] When the timing function of the CPLD is turned on and no pulse signal is detected for 120 seconds, the CPLD considers the system abnormal, simulates the watchdog timeout, and powers on and off the device again to restart.
[0139] It should be noted that the CPLD is connected to the outside through GPIO pins to control its enable signal and dog-feeding signal. In the method of feeding the dog through the CPLD, it also includes:
[0140] S1, Support flexible configuration of the watchdog timeout time, which can be specifically configured through the I2C link between the CPU and the CPLD through the protocol, so as to achieve different timeout times in different scenarios.
[0141] S2. Provide configuration options to select whether to enable the watchdog function when the storage system starts. For example, in special scenarios such as debugging problems, if the watchdog reset is not required, the watchdog function can be turned off by changing the configuration options.
[0142] S3. For the same hardware, when installing different operating systems, adaptively enable / disable the watchdog function. Specifically, when installing the storage system, write "Enable watchdog" in the boot partition. When the BIOS boot loader GRUB detects this special character, it will select to enable the watchdog and start the detection in the boot stage; when installing other storage systems, if this special character is not detected, the watchdog is prohibited.
[0143] S4. Provide a flexible shutdown method. For special scenarios, when collecting kernel crashdumps, system updates, or uninstalling the watchdog feeding module, the watchdog function needs to be turned off, and the CPLD no longer detects abnormal states.
[0144] The process of implementing watchdog monitoring through the CPLD is shown in Figure 6 。
[0145] Specifically, S602. The system boots up.
[0146] S604. The CPLD receives the rest_L signal.
[0147] S606. Enable the Watchdog function during the system startup phase. If the Watchdog function is enabled, execute step S608; if not, skip step S608 and directly execute step S610.
[0148] S608. The Watchdog starts timing.
[0149] S610. The BIOS initializes.
[0150] S612. Whether the flag indicating the completion of BIOS initialization is received within the T1 time. If yes, execute step S614; if not, execute step S622.
[0151] S614. The CPLD clears the timer.
[0152] S616. Whether the boot loader enables Watchdog enable. If yes, execute step S618; if not, execute steps S626 and S629.
[0153] S618. The CPLD receives the Watchdog enable signal and enables the Watchdog function to detect the kernel status during system operation.
[0154] S620. If the CPLD has not detected a heartbeat signal for more than T2 time, if yes, execute step S624; if no, execute step S630.
[0155] S622. The flag indicating the completion of bios initialization has not been received for more than T1 time.
[0156] S624. Power down the system.
[0157] S626. enable has not been received.
[0158] S628. Disable the Watchdog function.
[0159] S629. Start to enter the OS.
[0160] S630. The system runs normally.
[0161] It should be noted that after the system is powered on and the CPLD is initialized, if the watchdog function is enabled, the watchdog starts timing. If the pet watchdog heartbeat signal is not received within the specified time, the system is powered down and reset.
[0162] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0163] According to another aspect of the embodiments of the present application, there is also provided an application running monitoring device for implementing the above application running monitoring method. As Figure 7 shown, the device includes:
[0164] A determination unit 702, configured to determine a target process currently running in a target application;
[0165] A reset unit 704, configured to reset a process monitoring timer configured for the target application to a process monitoring duration adapted to the target process;
[0166] A trigger unit 706, configured to trigger a monitoring alarm when the running duration of the target process reaches the reset process monitoring duration but the heartbeat monitoring signal of the target process has not been received.
[0167] Optionally, the above reset unit 704 includes a predicted task duration module, configured to determine the predicted task duration to be executed by the daemon process when the target process is a daemon process; a reset task duration module, configured to reset the process monitoring duration of the process monitoring timer to the task predicted duration; a reset threshold duration module, configured to reset the process monitoring duration of the process monitoring timer to the target threshold duration when the target process is an input / output control process.
[0168] Optionally, the above reset unit 704 further includes a duration extension module, configured to extend the process monitoring duration to a transition threshold duration when receiving a process exit instruction for instructing to exit the input / output control process, where the transition threshold duration is greater than the target threshold duration, and the transition threshold duration is determined based on the switching duration from the input / output control process back to the daemon process.
[0169] Optionally, the above application running monitoring device further includes a thread wake-up unit, configured to wake up the user-mode thread in the sleep state based on the monitoring alarm and restart the hardware device where the target application is located when it is determined that the monitoring alarm is triggered by the daemon process; wake up the user-mode thread in the sleep state based on the monitoring alarm and restart a new input / output control process after the input / output control process is closed by the daemon process when it is determined that the monitoring alarm is triggered by the input / output control process.
[0170] Optionally, the above application running monitoring device further includes an initialization unit, configured to initialize device components and load the boot loader into the memory when the hardware device where the target application is located is started; execute the boot loader and load the operating system to be applied; configure the process monitoring timer to the enabled state when a preset monitoring character is detected in the operating system, where the process monitoring timer in the enabled state is in the on-timing state; configure the process monitoring timer to the disabled state when the preset monitoring character is not detected in the operating system, where the process monitoring timer in the disabled state is in the off-timing state.
[0171] Optionally, the above application running monitoring device further includes a post-processing unit, configured to perform initialization configuration for the process monitoring timer through the communication link between the central processing component and the logic processing component in the hardware device: configure the process monitoring timer to the disabled state when the configuration option indicates that the operating system is in the debug state; configure the process monitoring timer to the disabled state when the configuration option indicates that the operating system is in the crash state; configure the process monitoring timer to the disabled state when the configuration option indicates that the operating system is in the update state; configure the process monitoring timer to the disabled state when the configuration option indicates that the process monitoring timer is in the uninstalled state.
[0172] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above-mentioned method for monitoring the operation of an application program. The electronic device may be Figure 1 the terminal device or server shown in the figure. In this embodiment, the electronic device is taken as an example of a mobile phone or a computer. As Figure 8 shown in the figure, the electronic device includes a memory 802 and a processor 804. A computer program is stored in the memory 802, and the processor 804 is configured to execute the steps in any one of the above method embodiments through the computer program.
[0173] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.
[0174] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0175] S1, determine the target process currently running in the target application program;
[0176] S2, reset the process monitoring timer configured for the target application program to a process monitoring duration adapted to the target process;
[0177] S3, trigger a monitoring alarm when the running duration of the target process reaches the reset process monitoring duration but no heartbeat monitoring signal of the target process is received.
[0178] Optionally, those of ordinary skill in the art can understand that Figure 8 the structure shown in the figure is only schematic. The electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (MID), a PAD and other terminal devices. Figure 8 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 8 , or have a different configuration from that shown in Figure 8 .
[0179] Among them, the memory 802 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for monitoring the running of the application program in the embodiments of the present application. The processor 804 executes various functional applications and data processing by running the software programs and modules stored in the memory 802, that is, implements the method for monitoring the running of the above application program. The memory 802 may include the determination unit 702, the reset unit 704, and the trigger unit 706 in the above device for monitoring the running of the application program. The memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 802 may further include a memory remotely disposed relative to the processor 804, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 802 may specifically but not limitedly be used to store information such as page elements and page styles. This will not be elaborated in this example.
[0180] Optionally, the above transmission device 806 is used to receive or send data via a network. Specific examples of the above network may include a wired network and a wireless network. In one instance, the transmission device 806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one instance, the transmission device 806 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0181] In addition, the above electronic device further includes: a display 808, which is used to display the above target page; and a connection bus 810, which is used to connect each module component in the above electronic device.
[0182] In other embodiments, the above terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes may form a point-to-point network, and any form of computing device, such as a server, a terminal, and other electronic devices, can become a node in the blockchain system by joining the point-to-point network.
[0183] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present application.
Claims
1. A method for monitoring the operation of an application, characterized in that: include: Determine the target process currently running in the target application; Resetting the process monitoring timer configured for the target application to a process monitoring duration that is compatible with the target process; When the running time of the target process reaches the process monitoring time after reset, but the heartbeat monitoring signal of the target process is not received, a monitoring alarm is triggered.
2. The method according to claim 1, characterized in that The step of resetting the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process includes: In the case where the target process is a daemon process, determining a predicted duration of a task to be executed by the daemon process; resetting the process monitoring duration of the process monitoring timer to the predicted duration of the task; When the target process is an input / output control process, the process monitoring duration of the process monitoring timer is reset to a target threshold duration.
3. The method according to claim 2, characterized in that The step of resetting the process monitoring timer configured for the target application to a process monitoring duration adapted to the target process includes: When a process exit instruction is received to instruct exiting the input / output control process, the process monitoring duration is extended to a transition threshold duration, wherein the transition threshold duration is greater than the target threshold duration, and the transition threshold duration is determined based on a switching duration from the input / output control process back to the daemon process.
4. The method according to claim 2, characterized in that: After the monitoring alarm is triggered, the method further includes: In the case where it is determined that the daemon process triggers the monitoring alarm, waking up the user state thread in the sleeping state based on the monitoring alarm, and restarting the hardware device where the target application is located; When it is determined that the input-output control process triggers the monitoring alarm, the user-mode thread in sleep state is awakened based on the monitoring alarm, and after the input-output control process is closed by the daemon process, a new input-output control process is restarted.
5. The method according to claim 1, characterized in that Before determining the target process currently running in the target application, the method further includes: When the hardware device where the target application is located is started, initializing the device components and loading the boot loader into the memory; Execute the boot loader and load the operating system to be applied; In the case where a preset monitoring character is detected in the operating system, the process monitoring timer is configured to be in an enabled state, wherein the process monitoring timer in the enabled state is in an on timing state; In the case that the preset monitoring character is not detected in the operating system, the process monitoring timer is configured to be in a disabled state, wherein the process monitoring timer in the disabled state is in a closed timing state.
6. The method according to claim 5, characterized in that After the initialization of the device components, the method further includes: The process monitoring timer is initialized and configured in the following configuration options through the communication link between the central processing component and the logic processing component in the hardware device: When the configuration option indicates that the operating system is in a debugging state, configuring the process monitoring timer to be in a disabled state; When the configuration option indicates that the operating system is in a crash state, configuring the process monitoring timer to be in a disabled state; When the configuration option indicates that the operating system is in an updating state, configuring the process monitoring timer to be in a disabled state; When the configuration option indicates that the process monitoring timer is in an uninstalled state, the process monitoring timer is configured to be in a disabled state.
7. A device for monitoring the operation of an application, characterized in that: include: A determination unit, used to determine a target process currently running in a target application; A reset unit, configured to reset a process monitoring timer configured for the target application to a process monitoring duration adapted to the target process; The trigger unit is used to trigger a monitoring alarm when the running time of the target process reaches the process monitoring time after reset but the heartbeat monitoring signal of the target process is not received.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 6 when executed by a processor.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Storage system process management method, electronic equipment, storage medium and program product
CN120723586A