Downtime processing method and electronic equipment
By monitoring the status of the test-end equipment, distinguishing between true and false downtime, and adopting corresponding processing strategies, the problem of low efficiency of manual downtime processing in automated testing is solved, automated downtime processing is achieved, and testing efficiency is improved.
Patent Information
- Application Number
- CN202511233449.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
In automated testing scenarios, downtime of test-end equipment requires manual processing, resulting in low testing efficiency.
By monitoring the user interaction behavior, power status and kernel status of the test end device, we can distinguish between real and false crashes, and use the solid-state drive's power-off protection mode and user space checkpoint/recovery tools to automatically handle crashes.
It realizes automatic downtime processing, improves downtime processing efficiency, reduces test interruption time, and ensures test continuity and efficiency.
Smart Images

Figure CN120743685A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method for processing a computer outage and an electronic device. Background Art
[0002] In automated testing scenarios, the test-end equipment needs to ensure test continuity. However, during the test process, the test-end equipment may experience downtime. In this case, operation and maintenance personnel need to manually handle the downtime. However, manual downtime handling is inefficient, affecting overall test efficiency. Summary of the Invention
[0003] The present application provides a downtime processing method and electronic device to at least solve the problem in the related art that the efficiency of manual downtime processing is low, which affects the overall testing efficiency.
[0004] This application provides a method for handling downtime, including:
[0005] Monitoring the state of the test end device; the state includes user interaction behavior, the user interaction behavior includes input events of a physical input device; the state also includes at least one of a power state and a kernel state;
[0006] Determining whether a test end device has experienced a downtime based on monitored status data and determining the type of downtime; the downtime type is either a true downtime or a false downtime; the status data includes input event-related data; a false downtime is determined based on the input event-related data; the status data also includes at least one of power supply status-related data and kernel status-related data; a true downtime is determined based on at least one of the power supply status-related data and kernel status-related data;
[0007] In the event of a crash on the test-end device, the crash is handled according to the preset processing strategy corresponding to the crash type; the preset processing strategy corresponding to a real crash includes the power-off protection mode of the solid-state drive; the preset processing strategy corresponding to a fake crash includes the user space checkpoint / recovery tool.
[0008] This application also provides a downtime processing device, including:
[0009] A monitoring module for monitoring the state of the test end device; the state includes user interaction behavior, which includes input events of physical input devices; the state also includes at least one of power state and kernel state;
[0010] a determination module, configured to determine whether a test end device has experienced a downtime based on monitored status data, and to determine the type of downtime; the downtime type being either a true downtime or a false downtime; the status data including input event-related data; a false downtime being determined based on the input event-related data; the status data also including at least one of power supply status-related data and kernel status-related data; a true downtime being determined based on at least one of the power supply status-related data and kernel status-related data;
[0011] The processing module is used to handle the downtime according to the preset processing strategy corresponding to the downtime type when the test end device crashes; the preset processing strategy corresponding to the real downtime includes the power-off protection mode of the solid-state drive; the preset processing strategy corresponding to the fake downtime includes the user space checkpoint / recovery tool.
[0012] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for handling downtime when executing the computer program.
[0013] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for handling system downtime are implemented.
[0014] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for handling system downtime when executed by a processor.
[0015] The downtime processing method and electronic device provided by the present application can timely know the current situation of the test end device by monitoring the status of the test end device, such as at least one of the user interaction behavior, power status and kernel status, so as to determine whether the test end device has crashed based on the monitored status data, and can further determine whether it is a real crash or a false crash. In the case of a real crash, the downtime processing is performed according to the power-off protection mode of the solid-state hard drive. In the case of a false crash, the downtime processing is performed according to the user space checkpoint / recovery tool, thereby realizing automated downtime processing. Compared with the manual downtime processing method in the related art, the present application improves the efficiency of downtime processing, thereby improving the overall testing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1A schematic diagram of a system architecture provided for this application;
[0018] Figure 2 A flowchart of a method for handling downtime provided in this application;
[0019] Figure 3 A flowchart of another downtime handling method provided for this application;
[0020] Figure 4 A schematic diagram of the structure of a downtime processing device provided in this application;
[0021] Figure 5 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0022] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, other embodiments obtained by ordinary technicians in this field without making any creative work are all within the scope of protection of this application.
[0023] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0024] Linux: An operating system.
[0025] An outage is a condition in which a computer, server, or system stops responding, crashes, or freezes.
[0026] In automated testing scenarios, the test-end equipment needs to ensure test continuity. However, during the test process, the test-end equipment may experience downtime. In this case, operation and maintenance personnel need to manually handle the downtime. However, manual downtime handling is inefficient, affecting overall test efficiency.
[0027] Regarding the above technical issues, consider that crashes can be either true or false. A true crash occurs when the test device completely stops functioning, unable to be operated or recovered. This is typically caused by a serious hardware failure or system crash. A false crash occurs when the test device's operating system is unresponsive, but the kernel is alive. Different crash types require different handling methods. Therefore, when a test device experiences a crash, first determine the crash type and then address the issue.
[0028] Specifically, the present application monitors the status of the test end device, such as at least one of user interaction behavior, power status and kernel status, to timely know the current situation of the test end device, and thus determine whether the test end device has crashed based on the monitored status data, and can further determine whether it is a real crash or a false crash. In the case of a real crash, the crash is handled according to the power-off protection mode of the solid-state hard drive. In the case of a false crash, the crash is handled according to the user space checkpoint / recovery tool, thereby realizing automated crash handling. Compared with the manual crash handling method in the related technology, the present application improves the efficiency of crash handling, thereby improving the overall testing efficiency.
[0029] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0030] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the downtime handling method depends, the specific application environment architecture or specific hardware architecture is described here.
[0031] Figure 1 A schematic diagram of a system architecture provided for this application, such as Figure 1 As shown, the test end device 10 and the monitoring end device 20 are communicatively connected.
[0032] The test end device 10 is used for automated testing. For example, the test end device 10 is equipped with automated testing software to execute pre-set test cases and collect performance data and logs during the test. After the test is completed, the test end device 10 generates and stores a test report.
[0033] Exemplarily, the test end device 10 is a server or a terminal device, such as a notebook computer, a desktop computer, or other devices capable of automated testing.
[0034] The monitoring device 20 receives the remote desktop stream and test data transmitted by the test device 10 in real time, and displays the test execution screen and key indicators (such as progress and error prompts) synchronously through a graphical interface. Users can actively send control commands (such as pausing or terminating the test) to the test device 10 through the monitoring device 20.
[0035] Exemplarily, the monitoring terminal device 20 is a terminal device, such as a laptop computer, a desktop computer, or other remotely controllable device.
[0036] Optionally, the test end device 10 and the monitoring end device 20 use a remote desktop protocol for data exchange, for example, sharing a graphical desktop via Virtual Network Computing (VNC).
[0037] The monitoring terminal device 20 is in communication with a physical input device, which is a device connected to a user terminal via a Universal Serial Bus (USB), referred to as a USB device, such as a mouse, keyboard, or the like.
[0038] In one scenario, a user inputs a control instruction by operating a physical input device, wherein the user's operation of the physical input device can be regarded as a user interaction behavior, which triggers an input event of the physical input device.
[0039] In another scenario, an automation program is pre-installed on the monitoring terminal device 20, which is used to periodically trigger input events. The automation program simulates user interaction behavior, so that the user does not need to be on duty in front of the monitoring terminal device 20 at all times, saving labor costs.
[0040] The monitoring end device 20 sends the detected input event to the test end device 10 in real time, so that the test end device 10 executes the downtime processing method provided in the embodiment of the present application.
[0041] This application is suitable for long-term stress testing where test continuity must be guaranteed.
[0042] The execution subject of the downtime processing method provided in this application is a downtime processing device, which is integrated into an electronic device. For example, the electronic device is a test terminal device 10.
[0043] Figure 2 A flowchart of a downtime processing method provided in this application is shown as follows: Figure 2 As shown, the method includes the following steps.
[0044] S201: Monitor the status of the test end device.
[0045] The state of the test end device includes user interaction behavior.
[0046] User interaction refers to a user's operation of a physical input device on the monitoring device, which triggers an input event on the physical input device. For the testing device, monitoring user interaction is equivalent to monitoring input events on the physical input device. Therefore, user interaction can be considered to include input events on the physical input device.
[0047] The state of the test end device further includes at least one of a power state and a kernel state.
[0048] The power status is used to indicate whether the power supply of the test end device is abnormal.
[0049] The kernel status is used to indicate whether the network status of the kernel of the test end device is abnormal.
[0050] Optionally, the user interaction behavior, power state, and kernel state of the test end device are monitored separately. For example, the user interaction behavior of the test end device is monitored to obtain input event-related data; the power state of the test end device is monitored to obtain power state-related data; and the kernel state of the test end device is monitored to obtain kernel state-related data.
[0051] S202: Determine whether the test end device has experienced a downtime based on the monitored status data, and determine the downtime type.
[0052] The downtime type is either true downtime or false downtime.
[0053] A true downtime occurs when the test equipment completely stops working and cannot be operated or recovered in any way. This is usually caused by a serious hardware failure or system crash.
[0054] A false crash occurs when the operating system of the test device becomes unresponsive but the kernel survives.
[0055] When the state of the test end device includes user interaction behavior, the monitored state data is input event-related data, and the input event-related data includes the loss duration of the mouse event and the switching delay of the keyboard number lock key (NumLock).
[0056] The loss duration is the time difference between the current time and the last time a mouse event was triggered. The switching delay is the response time of the Num Lock key. Accordingly, a false downtime is determined based on the loss duration and switching delay. In other words, the loss duration of the mouse event and the switching delay of the keyboard Num Lock key can be used to determine whether the test end device is experiencing a false downtime.
[0057] If the test end device's status includes power status, the monitored status data includes power status data. This power status data indicates whether the power supply is normal or abnormal. Normal power means the test end device's power supply is operating normally. Abnormal power means there's a problem with the test end device's power supply, such as a power failure, unstable voltage, or a disconnected power supply. Accordingly, the power status data can be used to determine whether the test end device is truly down.
[0058] If the test-end device's status includes kernel status, the monitored status data includes kernel status data. Kernel status data indicates whether the kernel network is normal or abnormal. A normal kernel network indicates that the kernel network of the test-end device is connected properly. A kernel network abnormality indicates that a problem has occurred in the kernel network of the test-end device, such as a network connection interruption. Accordingly, the kernel status data can be used to determine whether the test-end device is truly down.
[0059] The true shutdown is determined based on at least one of the power state related data and the kernel state related data.
[0060] Exemplarily, a decision engine is pre-deployed in the test end device, and step S202 is implemented by the decision engine. For example, the decision engine can be implemented by a rule tree.
[0061] Optionally, when it is determined that the test end device has experienced a false downtime, a false downtime signal is output; or when it is determined that the test end device has experienced a true downtime, a true downtime signal is output, so that the test end device executes step S203 in response to the true downtime signal or the false downtime signal.
[0062] S203: When a downtime occurs on the test end device, downtime processing is performed according to a preset processing strategy corresponding to the downtime type.
[0063] This application sets corresponding preset processing strategies for real crashes and fake crashes: the preset processing strategy corresponding to real crashes includes the power loss protection (PLP) mode of the solid state drive (SSD); the preset processing strategy corresponding to fake crashes includes the user space checkpoint / restore (CRIU) tool.
[0064] The above preset processing strategies can automatically handle downtime.
[0065] Exemplarily, in response to a true downtime signal, downtime processing is performed according to a preset processing strategy corresponding to a true downtime; or, in response to a false downtime signal, downtime processing is performed according to a preset processing strategy corresponding to a home downtime.
[0066] By monitoring the status of the test end device, such as at least one of the user interaction behavior, power status and kernel status, the present application can timely know the current status of the test end device, thereby determining whether the test end device has crashed based on the monitored status data, and can further determine whether it is a real crash or a false crash. In the case of a real crash, the crash is handled according to the power-off protection mode of the SSD. In the case of a false crash, the crash is handled according to the user space checkpoint / recovery tool, thereby realizing automated crash handling. Compared with the manual crash handling method in the related art, the present application improves the efficiency of crash handling, thereby improving the overall testing efficiency.
[0067] The following describes the process of determining and handling false downtime.
[0068] In some embodiments, the input event includes a mouse event and a switching event of the keyboard number lock key. For example, the mouse event can be a mouse movement event or a mouse click event. Accordingly, when the state of the test end device includes user interaction behavior, monitoring the state of the test end device in S201 is specifically implemented as the following steps (1)-(2).
[0069] (1): In response to a mouse event being triggered, the number of mouse events and the time at which the mouse events are triggered are recorded. In addition, in response to the number lock key being pressed, the switching delay of the number lock key is recorded. The switching delay is the time difference between the key being pressed and the status indicator light changing.
[0070] Among them, the switching delay can be regarded as the response time of the number lock key.
[0071] (2): Periodically detect the trigger time and switching delay of the most recent record according to the preset sampling rate, and calculate the loss duration of the mouse event.
[0072] Optionally, the trigger time of the most recently recorded mouse event is periodically detected at a preset sampling rate, and the time difference between the current time and the trigger time is calculated each time. This time difference can be regarded as the duration of the mouse event loss, and a determination is made as to whether the duration of the loss is greater than a preset time threshold. If the duration of the loss is greater than the preset time threshold, it indicates that the number of occurrences of the mouse event has not been updated for a long period of time, i.e., the mouse has been unresponsive for a long period of time, and the mouse event can be considered lost. If the duration of the loss is less than or equal to the preset time threshold, it indicates that the mouse event has not been lost.
[0073] Optionally, the most recently recorded Num Lock key switching delay is periodically checked at a sampling rate, and during each test, a determination is made as to whether the switching delay is less than a preset delay threshold. If the switching delay is less than the preset delay threshold, the keyboard is responding normally; conversely, if the switching delay is greater than or equal to the preset delay threshold, the keyboard is stuck.
[0074] The sampling rate, preset time threshold and preset delay threshold can all be set according to actual needs. For example, the sampling rate is 100 Hz, the preset time threshold is 5 seconds (s), and the preset delay threshold is 200 milliseconds (ms). This application does not limit this.
[0075] Exemplarily, when a user moves the mouse or presses the keyboard on the monitoring end device, or when an automated program in the monitoring end device triggers an input event, the monitoring end device captures these input events. Accordingly, the monitoring end device encapsulates the event type of the input event (such as mouse movement, mouse button, keyboard button, etc.) and the specific parameters of the input event (such as mouse coordinates, key code of the button, etc.) into a data packet in a specific format and sends the data packet to the test end device.
[0076] The specific process for capturing input events includes: collecting current signals from the keyboard and mouse using a USB hub and current sensors; and converting the analog current signals into digital signals using an analog-to-digital converter (ADC). For example, the capture includes the mouse displacement and the rising edge delay of the keyboard number lock key. Displacement can be expressed as coordinates.
[0077] Accordingly, upon receiving the data packet, the test-end device parses it and performs the corresponding operations. These operations include: determining the device that triggered the input event based on the event type; if the device that triggered the input event is a mouse, incrementing the global counter by 1 to determine the number of mouse events and recording the current time, which can be considered the trigger time of the mouse event. If the device that triggered the input event is a keyboard and the pressed key is the number lock key, obtaining the timestamp of the key press and recording the timestamp of the status indicator change; and calculating the time difference between the two timestamps in nanoseconds, which can be considered the switching delay.
[0078] Exemplarily, the above corresponding operations are performed by a Field-Programmable Gate Array (FPGA).
[0079] On the basis of steps (1)-(2), S202 determines whether the test end device has crashed according to the monitored status data, and determines the crash type as follows: when the duration of the loss of the mouse event is greater than the preset time threshold and the switching delay of the number lock key is less than the preset delay threshold, it is determined that the test end device has crashed, and the crash type is determined to be a false crash.
[0080] If the duration of mouse event loss is greater than the preset time threshold and the switching delay of the number lock key is less than the preset delay threshold, it means that the mouse is unresponsive for a long time but the keyboard responds quickly. In this case, it can be determined that the test end device has a false crash.
[0081] The duration of mouse event loss directly reflects the responsiveness of the graphical user interface (GUI). If the GUI is stuck, it will be unable to process mouse events, causing the loss duration to increase. Pressing the keyboard number lock key triggers two actions: a key event is sent to the operating system (at the software level), and a physical change in the status indicator light (at the hardware level, but controlled by the operating system). In a healthy system, these two actions are almost instantaneous, with minimal latency. However, if the system is busy or stuck, software processing may slow, but the indicator light control signal may still be sent through the underlying driver, resulting in an abnormally long switching delay. By simultaneously monitoring these two indicators, an accurate judgment can be made: a long loss duration indicates an unresponsive GUI, while a short switching delay indicates that the underlying input driver and kernel interrupt processing are likely functioning properly (otherwise, the status indicator light would not illuminate), thus eliminating the possibility of a true crash. Combining these two indicators can accurately determine whether the test end device has experienced a false crash.
[0082] In addition, monitoring mouse events and the switching delay of a single key takes up less system resources, and based on continuous event monitoring and periodic time difference calculation, it can quickly detect unresponsive states, which is conducive to quickly triggering preset processing strategies and reducing downtime.
[0083] On the basis of the above embodiment, the kernel survival status can be combined to further determine whether a false crash has occurred. Optionally, determining whether the test end device has crashed based on the monitored status data and determining the crash type in S202 is specifically implemented as follows: when the loss duration is greater than a preset time threshold and the switching delay is less than a preset delay threshold, sending a process status check command to the kernel; when the process status check command can be executed and returns a result, determining that the kernel survival status is kernel survival; if the kernel survival status is determined to be kernel survival, determining that the test end device has crashed and determining the crash type is false crash.
[0084] Exemplarily, the process status viewing command is the ps command, which is a command used to display the current process status in the Linux system.
[0085] If the process status check command executes normally and returns results, the kernel is managing processes normally and can be considered to be in normal working order. In other words, the kernel is considered alive. If the kernel is alive but the mouse remains unresponsive for a long time while the keyboard responds quickly, it can be confirmed that the test end device has experienced a false crash.
[0086] By combining the kernel survival status to form a dual judgment mechanism, the possibility of misjudgment is reduced.
[0087] If the duration of the mouse event loss is less than or equal to the preset time threshold, and the switching delay of the number lock key is less than the preset delay threshold, it means that the mouse and keyboard responses are normal and no processing is performed.
[0088] If the duration of the loss of the mouse event is less than or equal to the preset time threshold, and the switching delay of the numeric lock key is greater than or equal to the preset delay threshold, it means that the mouse responds normally, but the keyboard is stuck. In this case, it can be determined that a false crash has occurred in the test end device. Alternatively, the kernel survival status can be combined to further determine whether a false crash has occurred. For example, a process status check command is sent to the kernel; if the process status check command can be executed and a result is returned, the kernel survival status is determined to be kernel survival; if the kernel survival status is determined to be kernel survival, it is determined that the test end device has crashed, and the crash type is determined to be a false crash.
[0089] If the duration of mouse event loss exceeds a preset time threshold, and the numeric lock key toggle delay is greater than or equal to the preset delay threshold, indicating that the mouse has been unresponsive for an extended period and the keyboard has been stuck, this situation can be combined with the kernel survival status to further determine whether a false crash has occurred. For example, a process status check command is sent to the kernel; if the process status check command is executed and returns a result, the kernel survival status is determined to be kernel alive; if the kernel survival status is determined to be kernel alive, then it is determined that the test end device has crashed, and the crash type is determined to be a false crash.
[0090] Regarding the aforementioned situation where kernel liveness is used to determine whether a crash is a false one, if the process status check command is unresponsive, this indicates the kernel may be stuck in a deadlock, infinite loop, or other abnormal state, preventing it from processing userspace requests. This situation is neither a false nor a true crash and requires further evaluation. For example, a notification message can be sent to a monitoring device to alert the user of the kernel anomaly, allowing them to promptly troubleshoot the cause.
[0091] One application scenario of the present application is: in the process of running a test process on a test end device, a downtime processing is performed. In response to the above-mentioned false downtime situation, in some embodiments, the downtime processing method provided by the present application also includes: in the process of running the test process, periodically freezing the process status of the test process through the CRIU tool. Accordingly, in S203, when a downtime occurs on the test end device, downtime processing is performed according to the preset processing strategy corresponding to the downtime type, which is specifically implemented as follows: in the case of a false downtime on the test end device, the test process is restored according to the most recently frozen process status through the CRIU tool; the currently running graphics service process is terminated, and the current desktop environment type is determined; and a new graphics service process is recreated according to the desktop environment type.
[0092] Among them, the freezing period can be set according to actual needs, such as 10 minutes, 20 minutes, etc., and this application does not limit this.
[0093] This application establishes a test process breakpoint resumption mechanism through the process recovery function of the CRIU tool. By periodically freezing the process status of the test process, when a false crash occurs, the test process can be restored according to the previously frozen process status, thus eliminating the need for retesting, saving test time, improving test efficiency, and reducing test progress deviations caused by crashes.
[0094] During a fake crash, the GUI freezes, indicating that the currently running graphics service process may be dead. By terminating the graphics service process and recreating a new one, a hot restart of the graphics service is achieved, quickly restoring the GUI without restarting the device. For example, the graphics service is the Xorg service (a display server software).
[0095] The desktop environment type is the type of the current desktop environment. For example, an environment variable is used to determine the desktop environment currently in use. This variable is set by the desktop environment when it is started and is used to indicate the current desktop environment type. When creating a new graphics service process, it is necessary to create an image service process corresponding to the desktop environment type.
[0096] In some embodiments, freezing the process state of a test process is specifically implemented by collecting state information of the test process; the state information includes memory data, file descriptors, and register status of the test process; and saving the state information to a checkpoint directory. Exemplarily, the state information is compressed into a successful package using the LZ4 compression algorithm, which is a high-performance lossless data compression algorithm, and then saved to the checkpoint directory.
[0097] Accordingly, the test process is restored according to the most recently frozen process state through the user space checkpoint / restore tool, including: loading the most recently frozen process state from the checkpoint directory; recreating a new test process, and controlling the newly created test process to continue execution from the most recently frozen process state.
[0098] Optionally, the CRIU tool freezes the test process's state information and uploads it to a Network File System (NFS) for shared storage. The CRIU tool also automatically generates checkpoints every 10 minutes for incremental saves. Furthermore, the CRIU tool responds to false crash signals and performs process recovery.
[0099] By using the checkpoint and recovery mechanism of the CRIU tool, the progress deviation during process recovery can be reduced and the accuracy and efficiency of recovery can be improved.
[0100] Using the CRIU tool for process recovery, by freezing and restoring the process state in user space while leveraging kernel support to accurately save file descriptor offsets and perform CRC32 checks on memory states, this enables a process breakpoint-resume mechanism, reducing test progress deviations. This approach is suitable for services and applications requiring high availability, as it allows for rapid recovery from process failures, minimizing service interruption. CRC32 (Cyclic Redundancy Check 32) is a 32-bit checksum based on the cyclic redundancy check (CRC) algorithm.
[0101] For example, during Serial Advanced Technology Attachment Hard Disk Drive (SATA HDD) compatibility testing, if a defective driver causes an Xorg service freeze—for example, if mouse events are lost for 8 seconds or the number lock key toggle delay is 120ms—a false crash is detected, and the decision engine generates a false crash signal, triggering the corresponding preset handling strategy. Specifically, the CRIU tool is used to resume the test process and hot-restart the graphics service. This allows for rapid test resumption and reduces downtime handling time.
[0102] The following describes the process of determining and handling a true outage.
[0103] In some embodiments, when the state of the test end device includes a power state, the monitored state data includes power state-related data. Monitoring the state of the test end device in S201 includes: monitoring the power state of the test end device via a baseboard management controller (BMC); correspondingly, determining whether the test end device has experienced a downtime and determining the downtime type based on the monitored state data in S202 includes: when the power state-related data indicates that the power supply of the test end device is abnormal, determining that the test end device has experienced a downtime and determining the downtime type as a true downtime.
[0104] On the contrary, when the power status-related data indicates that the power supply of the test end device is normal, it means that the test end device is not really down, and no processing is performed.
[0105] It is more accurate to determine whether the test end device has truly crashed by the power status.
[0106] In some embodiments, when the state of the test end device includes the kernel state, the monitored state data includes kernel state-related data; monitoring the state of the test end device in S201 includes: periodically sending a network connectivity check command to the kernel; when the response to the network connectivity check command times out, determining that the kernel state-related data is a kernel network abnormality; accordingly, determining whether the test end device has crashed and determining the type of crash based on the monitored state data in S202 includes: when the kernel state-related data is a kernel network abnormality, determining that the test end device has crashed and determining the type of crash as a true crash.
[0107] Exemplarily, the network connectivity check command is a ping command. The ping command is a command provided by the Linux system for checking whether a network interface is responsive. The ping command can reflect the kernel's network processing capabilities. If the ping command can execute normally and return a result, it indicates that the kernel's network subsystem is functioning normally, and the kernel network can be determined to be normal. If the ping command times out, it means that the expected response was not received within a specific time period. This may be because the network connection is interrupted, the test end device is unreachable, or the test end device kernel does not respond to the network request. In this case, the kernel network is abnormal.
[0108] It is relatively quick to determine whether the test end device has truly crashed by using the kernel status.
[0109] In actual applications, whether a system crash occurs can be determined based on the power status, the memory status, or a combination of the power status and the memory status.
[0110] In some embodiments, in S203, when the test end device crashes, crash processing is performed according to a preset processing strategy corresponding to the crash type, including: when the test end device actually crashes, activating the PLP mode of the SSD; using direct memory access (DMA) through the capacitor power supply window of the SSD to store the data in the memory into the NAND gate flash memory of the SSD; the data in the memory includes the value of the register specified by the test port; after the storage is completed, a shutdown operation is performed and a log is uploaded to the remote server.
[0111] Among them, the NAND flash memory refers to NAND flash memory, which is a non-volatile storage medium.
[0112] The length of the capacitor power supply window is determined by the solid state drive (SSD), for example, 500ms.
[0113] For example, using the capacitor power supply window, direct DMA transfer of key data in the memory to the NAND flash memory, with the source address being 0xFFFF0-0x1FFFF. DMA is a hardware-accelerated data transfer method with high data transfer efficiency and speed.
[0114] Exemplarily, the data is stored in a specific address range of the NAND flash memory, which is reserved for storing critical data to ensure that the critical data can be quickly found and restored when the system is restored.
[0115] Through PLP lightning backup technology, in the event of a real downtime, critical data in the memory can be quickly saved to the SSD, thus avoiding data loss and improving reliability and data security.
[0116] For example, the PLP circuit timing is as follows: at 0ms, the BMC detects a power anomaly; at 20ms, the PLP mode of the SSD is activated; at 50ms, a DMA transfer is initiated (source address: 0xFFFF0-0x1FFFF); at 450ms, data is written to the NAND flash memory; at 480ms, a completion signal is sent to the BMC; at 500ms, the system is safely shut down.
[0117] The above timing design complies with the Non-Volatile Memory Express (NVMe) 1.4 standard. Compared with the related art mode of collecting logs after the device is restarted, this application can ensure the integrity of data capture.
[0118] For example, the process for handling a true power outage in a power failure scenario involves: If a power anomaly is detected, a PLP backup is initiated, saving the register state of the critical central processing unit (CPU), specifically the debug port register state: EAX = 0x78AB, EIP = 0x7FFE82C1. After the data backup is complete, a safe shutdown is performed, and an encrypted log is uploaded.
[0119] The above registers contain the CPU's execution state at the time of power failure and are crucial for subsequent fault analysis and recovery. EAX = 0x78AB stores the value of the general-purpose register EAX, which is typically used to store function return values or as operands. EIP = 0x7FFE82C1 stores the value of the instruction pointer register EIP, which points to the address of the next instruction to be executed.
[0120] For example, when the BMC detects a power supply anomaly, it outputs an alarm code, for example, the alarm code 0xA1 represents a power supply anomaly.
[0121] Safe shutdown specifically involves shutting down running services and processes and saving necessary system status information. Logging: This records the system status and operation logs at the time of the power outage, including register values and memory page data. Encrypted upload: This log data is encrypted and uploaded to a remote server or cloud storage to ensure data security and integrity, which facilitates subsequent fault analysis and system recovery. For example, the cause of the fault is a power management firmware stack overflow.
[0122] In the related art, when a test-end device experiences a complete crash (such as a kernel crash or power failure), traditional solutions rely on the BMC to collect fault logs. However, the BMC only captures logs after a reboot, losing memory or register states at the moment of the crash. This results in a low preservation rate for critical events and poor data integrity. Furthermore, the time from crash to log generation is long, making it inadequate for high-density testing. However, this application utilizes the SSD's PLP mode after detecting a true crash to quickly save critical data from memory to the SSD, thereby avoiding data loss and improving reliability and data security.
[0123] Figure 3 A flowchart of another downtime processing method provided for this application is shown in FIG. Figure 3 As shown, this crash handling method includes: monitoring the status of the test end device, using the decision engine to determine whether a crash has occurred and the type of crash. If it is a true crash, a PLP flash backup is performed, followed by a safe shutdown and log generation. If it is a false crash, the CRIU process is restored, the graphics service is hot-restarted, and the test continues.
[0124] In some embodiments, the test end device is a master node in a multi-node test cluster, for example, a Remote Dictionary Server (Redis) cluster.
[0125] Accordingly, in S203, when a test-end device crashes, crash handling is performed according to a preset handling strategy corresponding to the crash type, including: activating the SSD's PLP mode in the event of a true crash; storing data in the memory to the SSD's NAND flash memory using DMA via the SSD's capacitor power supply window; performing a shutdown operation after the storage is complete, and uploading a log to a remote server. Furthermore, the crash handling method provided in this application further includes: restoring the cluster state based on the data in the NAND flash memory after the test-end device restarts.
[0126] The data in the memory includes cluster state related data, such as the cluster state data structure and the cluster node mapping table.
[0127] It should be noted that when Sentinel (a tool of Redis) detects that the master node is down, it selects a node from multiple slave nodes and promotes it to the new master node to ensure the high availability of the cluster.
[0128] After resolving the issue that caused the downtime, the master node can be restarted. After the master node restarts, the cluster state is restored using the data backed up by the PLP, specifically restoring the cluster state data structure and the cluster node mapping table.
[0129] This application can be applied to a multi-node test cluster. When the master node goes down, the cluster status can be backed up, so that after the master node is restarted, the cluster status can be quickly restored based on the backed-up data.
[0130] This application aims at the automated testing scenario and builds an unattended automatic processing solution for real / fake downtime, which achieves the purpose of capturing on-site data instantly during real downtime, minimizing the deviation of test progress after self-healing of fake downtime, and reducing manual intervention.
[0131] 1. Millisecond-level on-site capture of real downtime: Traditional BMC log solutions only capture data after restart, resulting in a high loss rate of key registers. This application reduces the data loss rate.
[0132] 2. Self-healing from fake crashes in seconds: Graphics service freezes, input driver failures, memory leaks, and other issues occur frequently, and traditional manual handling solutions take a long time. This application automates the handling of fake crashes, shortening the processing time and, in turn, the test interruption time caused by crashes.
[0133] 3. This application breaks through the bottleneck of test progress being reset to zero due to restarting the entire machine, and achieves rapid continuation of the test.
[0134] In addition to automated testing scenarios, the true and false crash discrimination mechanism provided in this application can also be used in industrial controllers, such as programmable logic controllers (PLCs).
[0135] Alternatively, it can be applied to driving scenarios. When the vehicle system experiences a frozen interface (e.g., the central control screen freezes) but the underlying program is operating normally, the steering wheel torque sensor response delay and gear response delay can be used to determine whether it is a false crash. If so, the display service is automatically restarted to ensure driving continuity. For example, if the steering wheel torque sensor response delay and gear response delay are greater than the preset response delay, a false crash is determined. Otherwise, no action is taken.
[0136] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0137] Figure 4 This is a schematic diagram of the structure of a downtime processing device provided by this application. Figure 4 As shown, the downtime processing device 40 includes the following modules.
[0138] Monitoring module 401, for monitoring the state of the test end device; the state includes user interaction behavior, which includes input events of physical input devices; the state also includes at least one of power state and kernel state;
[0139] Determination module 402 is configured to determine whether a test end device has experienced a downtime based on the monitored status data, and to determine the downtime type; the downtime type is either a true downtime or a false downtime; the status data includes input event-related data; a false downtime is determined based on the input event-related data; the status data also includes at least one of power state-related data and kernel state-related data; a true downtime is determined based on at least one of the power state-related data and kernel state-related data;
[0140] Processing module 403 is used to handle the downtime according to the preset processing strategy corresponding to the downtime type when the test end device downtime occurs; the preset processing strategy corresponding to the real downtime includes the power-off protection mode of the solid-state hard disk; the preset processing strategy corresponding to the fake downtime includes the user space checkpoint / recovery tool.
[0141] Optionally, the input events include mouse events and switching events of the keyboard number lock key; the input event-related data include the loss duration of the mouse event and the switching delay of the keyboard number lock key; when the state includes user interaction behavior, the monitoring module 401 is specifically used to: in response to the triggering of the mouse event, record the number of occurrences of the mouse event and the triggering time of the mouse event, and, in response to the keyboard number lock key being pressed, record the switching delay of the keyboard number lock key, where the switching delay is the time difference between the key being pressed and the status indicator light changing; periodically detect the trigger time and switching delay of the most recent record according to the preset sampling rate, and calculate the loss duration of the mouse event; the loss duration is the time difference between the current time and the trigger time; the determination module 402 is specifically used to: when the loss duration is greater than the preset time threshold and the switching delay is less than the preset delay threshold, determine that the test end device has crashed and determine that the crash type is a false crash.
[0142] Optionally, the determination module 402 is specifically used to: send a process status check command to the kernel when the loss duration is greater than a preset time threshold and the switching delay is less than a preset delay threshold; determine that the kernel survival state is kernel survival when the process status check command can be executed and return a result; if the kernel survival state is determined to be kernel survival, determine that the test end device has crashed, and determine that the crash type is a false crash.
[0143] Optionally, the downtime processing device 40 also includes a freezing module, which is used to: periodically freeze the process status of the test process through the user space checkpoint / recovery tool during the running of the test process; accordingly, the processing module 403 is specifically used to: in the event of a false downtime of the test end device, restore the test process according to the most recently frozen process status through the user space checkpoint / recovery tool; terminate the currently running graphics service process and determine the current desktop environment type; and recreate a new graphics service process according to the desktop environment type.
[0144] Optionally, the freezing module, when freezing the process state of the test process, is specifically used to: collect state information of the test process; the state information includes memory data, file descriptors and register state of the test process; and save the state information to a checkpoint directory; the processing module 403, when restoring the test process according to the most recently frozen process state through the user space checkpoint / restore tool, is specifically used to: load the most recently frozen process state from the checkpoint directory; recreate a new test process, and control the newly created test process to continue execution from the process state.
[0145] Optionally, when the status includes the power status, the status data includes power status-related data; the monitoring module 401 is specifically used to: monitor the power status of the test end device through the baseboard management controller; the determination module 402 is specifically used to: when the power status-related data indicates that the power supply of the test end device is abnormal, determine that the test end device has crashed, and determine that the crash type is a true crash.
[0146] Optionally, when the status includes the kernel status, the status data includes kernel status-related data; the monitoring module 401 is specifically used to: periodically send a network connectivity check command to the kernel; the determination module 402 is specifically used to: when the response to the network connectivity check command times out, determine that the test end device has crashed, and determine that the crash type is a true crash.
[0147] Optionally, the processing module 403 is specifically used to: activate the power-off protection mode of the solid-state hard disk in the event of a real crash of the test end device; store the data in the memory into the NAND gate flash memory of the solid-state hard disk using direct memory access through the capacitor power supply window of the solid-state hard disk; the data includes the value of the register specified by the test port; execute the shutdown operation after the storage is completed, and upload the log to the remote server.
[0148] Optionally, the test end device is a master node in a multi-node test cluster; the processing module 403 is specifically used to: in the event of a real downtime of the test end device, activate the power-off protection mode of the solid-state hard disk; through the capacitor power supply window of the solid-state hard disk, use direct memory access to store the data in the memory into the NAND gate flash memory of the solid-state hard disk; the data includes a cluster status data structure and a cluster node mapping table; after the storage is completed, perform a shutdown operation and upload a log to a remote server; after the test end device is restarted, restore the cluster status according to the data in the NAND gate flash memory.
[0149] For the description of the features in the embodiment corresponding to the downtime processing device 40, reference can be made to the relevant description of the embodiment corresponding to the downtime processing method, which will not be repeated here.
[0150] The present application also provides an electronic device, which may be the above-mentioned test end device.
[0151] Figure 5 This is a schematic diagram of the structure of an electronic device provided by this application. Figure 5 As shown, the electronic device 50 provided in this embodiment includes: a processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.
[0152] During the specific implementation process, the processor 501 executes the computer program stored in the memory 502, so that the processor 501 executes the above-mentioned embodiment of the method for handling system downtime.
[0153] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0154] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0155] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0156] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, and control buses.
[0157] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned embodiments of the method for handling system downtime when running.
[0158] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0159] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned embodiments of the downtime processing method are implemented.
[0160] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned embodiments of the downtime processing method are implemented.
[0161] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0162] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0163] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0164] It should be understood that the above-described device embodiments are merely illustrative, and the devices of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0165] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.
[0166] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for specific applications, but such implementation should not be considered to be beyond the scope of this application.
[0167] The above is a detailed introduction to a downtime processing method and electronic device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for handling downtime, characterized in that: include: Monitoring the state of the test end device; the state includes user interaction behavior, the user interaction behavior includes input events of a physical input device; the state also includes at least one of a power state and a kernel state; Determine whether the test end device has experienced a downtime and the downtime type based on the monitored status data; the downtime type is a true downtime or a false downtime; the status data includes input event related data; The false downtime is determined based on the input event related data; the state data further includes at least one of power state related data and kernel state related data; The true shutdown is determined based on at least one of the power state related data and the kernel state related data; In the event that the test end device crashes, processing the crash is performed according to a preset processing strategy corresponding to the crash type; The preset processing strategy corresponding to the real crash includes a power-off protection mode of the solid-state hard disk; the preset processing strategy corresponding to the fake crash includes a user space checkpoint / recovery tool.
2. The method for handling system downtime according to claim 1, wherein: The input events include mouse events and keyboard number lock key switching events; the input event related data include the loss duration of the mouse events and the keyboard number lock key switching delay; In the case where the state includes the user interaction behavior, monitoring the state of the test end device includes: In response to the mouse event being triggered, recording the number of occurrences of the mouse event and the triggering time of the mouse event; and, in response to the keyboard number lock key being pressed, recording the switching delay of the keyboard number lock key, where the switching delay is the time difference between the key being pressed and the status indicator light changing; Periodically detecting the trigger time and the switching delay recorded most recently according to a preset sampling rate, and calculating the loss duration of the mouse event; the loss duration is the time difference between the current time and the trigger time; Determining whether the test end device has experienced a downtime and determining the downtime type based on the monitored status data includes: When the loss duration is greater than a preset time threshold and the switching delay is less than a preset delay threshold, it is determined that the test end device has crashed, and the crash type is determined to be a false crash.
3. The method for handling system downtime according to claim 2, wherein: Determining whether the test end device has experienced a downtime and determining the downtime type based on the monitored status data includes: When the loss duration is greater than a preset time threshold and the switching delay is less than a preset delay threshold, sending a process status check command to the kernel; When the process status checking command can be executed and returns a result, determining that the kernel survival state is kernel survival; If it is determined that the kernel survival state is kernel survival, it is determined that the test end device has crashed, and the crash type is determined to be a false crash.
4. The method for handling system downtime according to claim 1 or 2, wherein: The method further comprises: During the running of the test process, periodically freezing the process state of the test process by the user space checkpoint / recovery tool; When the test end device crashes, performing crash processing according to a preset processing strategy corresponding to the crash type includes: In the event of a false downtime of the test end device, recovering the test process according to the most recently frozen process state through the user space checkpoint / recovery tool; Terminate the currently running graphics service process and determine the current desktop environment type; A new graphics service process is recreated according to the desktop environment type.
5. The method for handling system downtime according to claim 4, wherein: Freezing the process state of the test process, including: Collecting status information of the test process; the status information includes memory data, file descriptors and register status of the test process; Saving the state information to a checkpoint directory; The recovering the test process according to the most recently frozen process state by the user space checkpoint / recovery tool comprises: Load the most recently frozen state of the process from the checkpoint directory; A new test process is recreated, and the newly created test process is controlled to continue execution from the process state.
6. The method for handling system downtime according to claim 1, wherein: In the case where the state includes the power state, the state data includes data related to the power state; and monitoring the state of the test end device includes: Monitoring the power status of the test end device by a baseboard management controller; Determining whether the test end device has experienced a downtime and determining the downtime type based on the monitored status data includes: When the power state-related data indicates that the power supply of the test end device is abnormal, it is determined that the test end device has crashed, and the crash type is determined to be a true crash.
7. The method for handling system downtime according to claim 1, wherein: In the case where the state includes the kernel state, the state data includes kernel state related data; The monitoring of the status of the test end device includes: Periodically send network connectivity check commands to the kernel; Determining whether the test end device has experienced a downtime and determining the downtime type based on the monitored status data includes: In the case where the response to the network connectivity check command times out, it is determined that the test end device has crashed, and the crash type is determined to be a true crash.
8. The method for handling system downtime according to claim 6 or 7, wherein: When the test end device crashes, performing crash processing according to a preset processing strategy corresponding to the crash type includes: In the event of a true downtime of the test end device, activating a power-off protection mode of the solid-state hard disk; The data in the memory is stored in the NAND gate flash memory of the solid state drive by using a direct memory access method through the capacitor power supply window of the solid state drive; the data includes the value of the test port designated register; After the storage is completed, the shutdown operation is performed and the logs are uploaded to the remote server.
9. The method for handling system downtime according to claim 6 or 7, wherein: The test end device is a master node in a multi-node test cluster; When the test end device crashes, performing crash processing according to a preset processing strategy corresponding to the crash type includes: In the event of a true downtime of the test end device, activating a power-off protection mode of the solid-state hard disk; The data in the memory is stored in the NAND flash memory of the solid state drive by direct memory access through the capacitor power supply window of the solid state drive; the data includes a cluster state data structure and a cluster node mapping table; After the storage is completed, the shutdown operation is executed and the logs are uploaded to the remote server; The method further comprises: After the test end device is restarted, the cluster state is restored according to the data in the NAND gate flash memory.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the downtime processing method according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Crash recovery method and device, equipment and storage medium
CN115756924A
Server downtime fault position obtaining method and device and program product
CN118193269A
System downtime detection method and device, electronic equipment and storage medium
CN118301045A
Programmable logic controller, downtime diagnosis method and device thereof and readable storage medium
CN119806882A
Recovery Method for Terminal Device Startup Failure and Terminal Device
US20190294490A1