Load-aware software suspension processing method and device, equipment and medium
Identifying software suspend issues through hierarchical tracking and inspection rules, solving infinite loops and heavy load tasks, improving system reliability and performance.
Patent Information
- Application Number
- CN202510506940.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art is difficult to effectively detect and deal with software suspend issues, especially infinite cycles and heavy load tasks, resulting in impaired system reliability and performance.
The hierarchical tracking method is used to detect at the application layer and the system layer, and the branch instruction record and eBPF tracking system calls are tracked through hardware, combined with inspection rules and early warning methods, LHB and heavy load tasks are identified, and configuration items are generated for optimization.
It improves the reliability of the software during long-term operation, reduces detection overhead, and achieves efficient detection and optimization of reliability problems.
Smart Images

Figure CN120336154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a load-aware software suspension processing method, apparatus, device, and medium. Background Art
[0002] Infrastructure software (such as HDFS) is the foundation for building large-scale systems because they need to provide underlying and stable long-term services for high-level applications. In this case, the reliability of such software plays an extremely important role in the normal operation of the entire system. However, when the software is running, it often faces complex and changing task loads. At this time, the software may get stuck in an abnormally long-running task (far exceeding the normal running time) due to special loads, or even become unresponsive to other requests, seriously affecting the reliability of the software system. Configuration, as an important means to control software behavior and adjust resource allocation, can help the software adapt to different loads and avoid / fix possible software reliability problems. However, not all such unresponsive problems are caused by loads, and the underlying reason may also be an infinite suspension of the software (Software Hang Bug) caused by code defects (which can also be triggered by normal task loads). Such suspension problems can also lead to task failures of upper-layer applications and even further damage the entire system.
[0003] Existing work on detecting software suspension problems mainly uses:
[0004] (1) Runtime log information to analyze the software state and discover potential execution anomalies. Such a method highly depends on the quality of the software's own logs;
[0005] (2) Static program analysis and SAT or SMT-based techniques to prove the termination of the code. These methods are limited by the complexity of the program.
[0006] When the software system is running an "abnormally long" task or even stuck in an "unresponsive" state, it is usually very difficult for users to decide whether to kill this task because they don't know whether this task is currently in an infinite suspension state or still in progress. Killing a running work task (although running slowly) may affect service stability and thus result in huge overhead caused by software service restart, while waiting for an infinite suspension may ultimately cause serious damage to the reliability and performance of the system. Summary of the Invention
[0007] The main objective of the embodiments of the present invention is to propose a load-aware software suspension processing method, apparatus, device, and medium, which improves the reliability of the application program during long-term operation.
[0008] One aspect of the present invention provides a load-aware software suspension processing method, including:
[0009] Use a hierarchical tracking method to detect the target application at the application layer and the system layer, and obtain the application layer tracking result and the system layer tracking result;
[0010] Use inspection rules and warning triggering methods for the application layer tracking result and the system layer tracking result to determine the software reliability inspection result;
[0011] Determine the reliability problem type according to the application reliability inspection result, where the reliability problem type includes LHB and heavy load tasks;
[0012] Generate configuration items according to the reliability problem type, and optimize the reliability of the target software through the configuration items.
[0013] According to the load-aware software suspension processing method described above, where a hierarchical tracking method is used to detect the target application at the application layer and the system layer, and obtain the application layer tracking result and the system layer tracking result, including:
[0014] Use hardware tracing to obtain branch instruction records, determine the address distribution of branch instructions according to the branch instruction records, and determine the application layer tracking result according to the address distribution of branch instructions;
[0015] Use the eBPF tracing method for events and system calls during the operation of the target application to obtain the system layer tracking result, where the task types of events and system calls include file management, communication, memory management, I / O control, process management, and synchronization control.
[0016] According to the load-aware software suspension processing method described above, where hardware tracing is used to obtain branch instruction records, and the address distribution of branch instructions is determined according to the branch instruction records, including:
[0017] Use the Linux performance analysis tool to obtain branch instruction records, and after filtering the branch instruction records for the execution information of non-target applications, obtain the mapping of the branch target address and the calling method during the operation of the target application.
[0018] According to the load-aware software suspension processing method described above, where inspection rules and warning triggering methods are used for the application layer tracking result and the system layer tracking result to determine the application reliability inspection result, including:
[0019] The inspection rules include intensive repeated code execution, accessing the same resources, and becoming invalid after specific operations. Among them, the intensive repeated code execution determines the application behavior of the target application through the distribution of branch target addresses; among them, the access to the same resources uses a quadruple to record the accessed resources of the target application, where the quadruple includes the accessed resource, the access operation, the access unit, and the access content; the becoming invalid after specific operations uses a quadruple group, where the quadruple includes the accessed resource, the access operation, the access unit, and the error flag;
[0020] The method for triggering a warning uses a time window to detect the repeated behaviors of each inspection rule and obtains the application reliability inspection result.
[0021] According to the load-aware software suspension processing method described above, where the method for triggering a warning uses a time window to detect the repeated behaviors of each inspection rule and obtains the application reliability inspection result, including:
[0022] According to the time period and the number of windows of the time window, when the inspection rule is intensive repeated code execution, collect the distribution of branch target addresses within the preset number of windows, determine the number of branches and the branch percentage of different targets in each time window, and then determine the LHB during the execution of the target application;
[0023] When the inspection rule is accessing the same resources or becoming invalid after specific operations, detect the first time window. If there are repeated access units, repeat the tracking record sequence of the preset number of time windows to obtain the LHB during the execution of the target application. If there are no repeated access units, reset the start time of the time window.
[0024] According to the load-aware software suspension processing method described above, where according to the type of reliability problem, generate configuration items, and optimize the reliability of the target software through the configuration items, including:
[0025] When the type of reliability problem is LHB, determine the location of the reliability problem by checking the code, and optimize the LHB using the first configuration item;
[0026] When the type of reliability problem is a heavy-load task, optimize the heavy-load task using the second configuration item.
[0027] According to the load-aware software suspension processing method described above, where optimizing the LHB using the first configuration item includes:
[0028] When the LHB enters a loop, add check code, obtain the loop variable and the loop termination condition through the check code, and execute the loop termination process of the LHB according to the loop variable and the loop termination condition;
[0029] Alternatively, add the maximum number of retries and timeout for LHB through the first configuration item.
[0030] Another aspect of the embodiments of the present invention provides a load-aware software suspension processing device, including:
[0031] A first module for detecting a target application at the application layer and the system layer by using a hierarchical tracing method to obtain an application layer tracing result and a system layer tracing result;
[0032] A second module for determining a software reliability inspection result by using an inspection rule and a warning triggering method for the application layer tracing result and the system layer tracing result;
[0033] A third module for determining a reliability problem type according to the application reliability inspection result, where the reliability problem type includes LHB and heavy load tasks;
[0034] A fourth module for generating a configuration item according to the reliability problem type, and optimizing the reliability of the target software through the configuration item.
[0035] Another aspect of the embodiments of the present invention provides an electronic device, including a processor and a memory;
[0036] The memory is used to store a program;
[0037] The processor executes the program to implement the method described above.
[0038] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described above.
[0039] The beneficial effects of the present invention are as follows: Use the method of tracking the target address of branch instructions and lightweight system-level tracing to detect infinite loops that may cause system suspension. Monitor the execution of the application program by tracking the target address of the CPU branch instruction. Compared with traditional application layer tracing, branch target tracing is supported by hardware, with relatively low overhead and no need for any instrumentation in the source code; Analyze the branch target traces and system call sequences within a continuous time period, and can accurately obtain whether the application program is in LHB or heavy load tasks for a long time, and perform corresponding optimizations according to the application program status, improving the reliability of the application program and achieving high reliability problem detection with low consumption. Description of the Drawings
[0040] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:
[0041] Figure 1 is a schematic flowchart of a load-aware software suspension processing method according to an embodiment of the present invention.
[0042] Figure 2 is an example diagram of system call trace records according to an embodiment of the present invention.
[0043] Figure 3 is an example diagram of repeated access to the same resource according to an embodiment of the present invention.
[0044] Figure 4 is an example diagram of "failure" in the same stage according to an embodiment of the present invention.
[0045] Figure 5 is a framework and flowchart of RwChecker according to an embodiment of the present invention.
[0046] Figure 6 is a schematic diagram of a load-aware software suspension processing device according to an embodiment of the present invention. Detailed implementation manners
[0047] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. In the following description, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of description of the present invention and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably. "First", "second", etc. are only used to distinguish technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In the following description, the consecutive numbering of method steps is for the convenience of review and understanding. Combining the overall technical solution of the present invention and the logical relationship between each step, adjusting the implementation order between steps will not affect the technical effect achieved by the technical solution of the present invention. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0048] RwChecker, a load-aware software suspension processing tool according to an embodiment of the present invention.
[0049] Refer to Figure 1 , Figure 1 is a schematic flowchart of a load-aware software suspension processing method according to an embodiment of the present invention, which includes but is not limited to steps S100 to S400:
[0050] S100. Detect the target application at the application layer and the system layer using a hierarchical tracing method to obtain an application layer tracing result and a system layer tracing result.
[0051] In some embodiments, obtain branch instruction records by hardware tracing, determine the address distribution of branch instructions according to the branch instruction records, and determine the application layer tracing result according to the address distribution of the branch instructions.
[0052] In some embodiments, use the eBPF tracing method for events and system calls during the runtime of the target application to obtain a system layer tracing result, where the task types of the events and system calls include file management, communication, memory management, I / O control, process management, and synchronization control.
[0053] For the application layer tracing records, RwChecker collects the recent branch instruction records LBR through hardware tracing and counts the address distribution of the branch instructions to characterize the application behavior; for the infinite loop execution of code blocks, the branch target addresses should be concentrated on the same or multiple addresses; branch target tracing has low time overhead and does not require code instrumentation; filter out the branch records in the kernel through RwChecker and only collect the execution information of the application program.
[0054] For example, RwChecker uses Linux_perf_events for tracing and reporting, and uses perfmap-agent to generate a mapping between branch target addresses and example methods of system call tracing records at runtime.
[0055] In some embodiments, refer to Figure 2 With reference to the example diagram of system call tracing records, RwChecker adopts the tracing technology of eBPF to trace specific system calls. RwChecker focuses on six main task types: file management, communication, memory management, I / O control, process management, and synchronization control. These are common tasks for application programs to interact with the system and access corresponding system resources. Each task instance consists of its sub-operations (and their combinations). For example, a file management task instance may first create a non-existent file and then perform several read and write operations. Different tasks have different combinations and sequences of sub-operations. For each task instance record, it contains a unique resource ID to record the accessed resource, and RwChecker also records specific parameters in the system call to compare different task instances. The resource ID and parameters can be obtained in each system call.
[0056] S200. Determine the software reliability check result by using check rules and warning triggering methods for the application layer tracing result and the system layer tracing result.
[0057] In some embodiments, the occurrence of LHB in an application is determined in the following manner, including:
[0058] Detect through inspection rules to obtain possible duplicate behavior rules. The inspection rules include intensive repeated code execution, accessing the same resource, and failure after a specific operation. Among them, intensive repeated code execution determines the application behavior of the target application through the distribution of branch target addresses; accessing the same resource records the accessed resources of the target application using a quadruple, where the quadruple includes the accessed resource, access operation, access unit, and access content; failure after a specific operation uses a quadruple group, where the quadruple includes the accessed resource, access operation, access unit, and error identifier;
[0059] The method of triggering a warning uses a time window to detect the repeated behavior of each inspection rule to obtain the application reliability inspection result.
[0060] In some embodiments, for intensive repeated code execution, based on the branch target address tracking record, when the software falls into an infinite loop, the same code block (function method) will be repeatedly executed without executing other functions. For example, in HBase - 14621, in the loadWALsFromQueues method, a retry infinite loop may cause the task to hang because the error logic of updating v0 causes it to never be equal to v1. The code snippet is as follows:
[0061] 1private Set <string>loadWALsFromQueues(){
[0062] 2-int v0 = replicationQueues.getQueuesZNodeCversion();
[0063] 3for(int retry = 0;; retry++){
[0064] 4+int v0 = replicationQueues.getQueuesZNodeCversion();
[0065] 5... / / load wal from queue
[0066] 6int v1 = replicationQueues.getQueuesZNodeCversion();
[0067] 7if(v0 == v1) return wals;
[0068] 8}
[0069] 9}
[0070] In this case, the loadWALsFromQueues method will be repeatedly executed without calling other methods or returning to the calling method above. To monitor and detect such code execution behavior, RwChecker records the branch target addresses for a specific period of time to show the execution of different code blocks. Since the number of branch instructions is usually very large, it is not feasible to directly analyze them on the branch stack. To solve this problem, RwChecker uses the branch target address distribution to indirectly characterize the program execution behavior. In the example shown in the above case, the method loadWALsFromQueues occupies more than 98% of the branch targets in a 5-second record. RwChecker collects the user-mode CPU branch instruction address targets in consecutive time windows.
[0071] For access to the same resource and invalidation after a specific operation, it is based on system call trace records.
[0072] In some embodiments, a process that accesses the same resource application and frequently accesses a resource with defects during operation usually repeatedly accesses exactly the same resource with the same content without accessing other different resources during execution. This means that the task may fall into an abnormal resource interaction state. Taking HBase-20865 as an example. At the beginning, the CreateTableProcedure fails to write a file to HDFS. At this time, part of the file has been successfully written. During subsequent retries, the CREATE_TABLE_WRITE_FS_LAYOUT state checks the existence status of the file each time. Since the old file is not cleaned up, this check fails in each retry, triggering an infinite retry loop. The code snippet is as follows:
[0073] 1FLow executeFromState()
[0074] 2switch(state){
[0075] 3case CREATE_TABLE_WRITE_FS_LAYOUT:
[0076] 4+DeleteTableProcedure.deleteFromFs(...);
[0077] 5newRegions=createFsLayout(); 6...
[0079] 7} 8
[0081] 9HRegionFileSystem createRegionOnFileSystem(){ 10...
[0083] 11Path regionDir=regionFs.getRegionDir();
[0084] 12if fs.exists(regionDir))
[0085] 13throw new IOException(”The specified region already
[0086] exists on disk”); 14...
[0088] 15}
[0089] Reference Figure 3 Refer to the example diagram of repeated access to the same resource shown below. This is the system call behavior of such abnormal resource interaction. In this case, the CreateTableProcedure will obtain the block file information from the meta file and check the status of the file at each retry. In this infinite retry loop, the CreateTableProcedure will repeatedly access the same file content without performing other operations and will not make any real modifications to these files.
[0090] Reference Figure 4 Refer to the example diagram of "failure" in the same phase shown below. Embodiments of the present invention describe potential instances (iLoop) of LHB through a tuple of four elements.
[0091] Among them, Resource (resource) records the resources being accessed by the process. In this rule, the resources should be the same, which means that the software wastes time accessing the same resources. In the above example, the resources are:
[0092] iLoop.Resource = {meta_file_path(r1), block_file_path(r2)}
[0093] Operation (operation) includes all system calls for operating on the resources. The operation set is fixed in the same case, but is replaceable in different cases and for different resources. The operation set for this example is:
[0094] iLoop.operation = {open(o1), read(o2), close(o3), stat(o4)}
[0095] AccessUnit (access unit) describes the overall repeated behavior of accessing resources. In this example, it is shown as:
[0096] iLoop.accessUnit = {o1(r1), o2(r1, Content), o3(r1), o4(r2)}
[0097] This access unit will appear in a cyclic manner in the system call sequence without being interspersed with other operations. The last element is Content (accessed content), which is related to a specific operation and can be obtained through the specific parameters and return values of the system call. In this file access case, the accessed content is the content read from the file / socket. In an LHB, usually the content in each access unit is the same, indicating that the process has fallen into repeated work.
[0098] For failures at the same stage after the same operation, local errors are common in system execution. In LHB, the same error may occur at the same stage after repeated operations. This stage can be a specific system call or signal handling. This means that the system does not take effective countermeasures against errors. Figure 4 Shows an example of an infinite while loop caused by integer overflow (Yarn-8833). The erroneous code snippet is shown below:
[0099] 1void computeSharesInternal(...) 2...
[0101] 3while(resourceUsedWithWeightToResourceRatio(rMax,...)
[0102] 4<totalResource){
[0103] 5rMax *= 2.0;
[0104] 6} 7
[0106] 8-static int resourceUsedWithWeightToResourceRatio(...){
[0107] 9+static long resourceUsedWithWeightToResourceRatio(...){
[0108] 10-int share = computeShare(sched,w2rRatio,type);
[0109] 11+long share = computeShare(sched,w2rRatio,type);
[0110] 12+if(Long.MAX_VALUE-resourcesTaken<share);
[0111] 13+return Long.MAX_VALUE;
[0112] 14...}
[0113] In this error scenario, the computing shared thread (running computeSharesInternal) will hold the write lock and prevent the scheduling process from performing related operations. The system call trace records of the computing thread and the scheduling thread containing the error are as Figure 4 shown. The computing thread continuously executes mprotect() on the same memory address addr. Since the memory state switches between no_access and read_only, other threads will repeatedly "fail" (receive the SIGSEGV signal, where SEGV_ACCER indicates accessing memory in a way that violates protection) after attempting to access the memory address addr.
[0114] Embodiments of the present invention further describe such a situation through a quadruple, including:
[0115] (iLoop: Resource, Operations, AccessUnit, ErrorSign)
[0116] For Error Sign (error flag), it indicates that an error occurs after the access unit. These flags often represent a more serious situation encountered by the system rather than common errors that often occur during normal task execution.
[0117] In some embodiments, according to the time period and the number of windows of the time window, when the inspection rule is dense repeated code execution, the branch target address distribution within a preset number of windows is collected, the number of branches and the branch percentage of different targets in each time window are determined, and then the LHB during the execution of the target application is determined.
[0118] In some embodiments, when the inspection rule is accessing the same resource or becoming invalid after a specific operation, the first time window is detected. If there are duplicate access units, the trace record sequence of a preset number of time windows is repeatedly executed to obtain the LHB during the execution of the target application. If there are no duplicate access units, the start time of the time window is reset.
[0119] In some embodiments, RwChecker checks whether there is a repeated occurrence of the loop instance iLoop to determine whether there is a potential LHB. RwChecker uses continuous time windows for the check instead of a fixed time threshold to ensure that the repeated behavior is consistent within a continuous time period. There are two parameters to adjust the matching process, which are the time period Δt and the number of windows n respectively. The user can configure specific parameters according to the complexity of the task.
[0120] For the execution of code with dense and repetitive rules, at the window start time T0 (when the user starts running RwChecker), RwChecker collects the branch target address distribution within n time windows. For each time window W i (0 < i < n), whose end time is T0 + i * Δt, RwChecker records the number of branches BN i for different targets and the branch percentage BP i in each window W i . For the top 3 most dense target addresses in the distribution, if BP i and BN i / (i * Δt) are equal in each time window (the difference does not exceed 0.1). RwChecker considers that there is a potential LHB during program execution.
[0121] For regular access to the same resource and invalidation after a specific operation, RwChecker tries to find possible abnormal access units (these access units continuously loop within consecutive time). Specifically, at the start time T0, RwChecker first analyzes the system call sequence of the first time window and finds potential loop access units. If there are no potentially recurring LHB access units in the first window, RwChecker will directly reset the start time. Otherwise, RwChecker will continue to collect the trace record sequences of n time windows (the end time of W i is T0 + i * Δt). If there are recurring access units for regular access to the same resource or invalidation after a specific operation in each window W i ((0 < i < n)), RwChecker will report a potential LHB, otherwise it will refresh the window start time and return to the first window state.
[0122] S300, determine the type of reliability problem according to the application reliability check result, where the type of reliability problem includes LHB and heavy load tasks.
[0123] S400, generate configuration items according to the type of reliability problem, and optimize the reliability of the target software through the configuration items.
[0124] In some embodiments, when LHB enters a loop, check code is added. The loop variable and loop termination condition are obtained through the check code, and the loop termination process of LHB is executed according to the loop variable and loop termination condition; or, the maximum number of retries and timeout time of LHB are added through the first configuration item.
[0125] In some embodiments, the first configuration item is used to limit the method of loop execution. The specific example is as follows in the code segment:
[0126] / *add Hive configuration parameter* /
[0127] 2public static enum ConfVars{ 3...
[0129] 4+HIVE_ICEBERG_METADATA_REFRESH_MAX_RETRIES(
[0130] 5+”hive.iceberg.metadata.refresh.max.retries”,2,
[0131] 6+”Max retry count for trying to access metadata location”+
[0132] 7+”in order to refresh metadata during Iceberg table load.”),
[0133] 8} 9
[0135] 10protected void doRefresh(){ 11...
[0137] 12-refreshFromMetadataLocation(...);
[0138] 13+refreshFromMetadataLocation(...,
[0139] 14+HiveConf.getIntVar(conf,
[0140] 15+HIVE_ICEBERG_METADATA_REFRESH_MAX_RETRIES));
[0141] 16}
[0142] In some embodiments, for the optimization of heavy-load tasks, the method of controlling the running / retry duration of tasks by using time-related configuration items can be adopted, and the control of resource usage by using resource-related configuration items can also be achieved through configuration items.
[0143] In some embodiments, refer to Figure 5 The framework and flowchart of RwChecker are shown. A framework that automatically monitors software behavior and identifies potential loop-induced hang problems in long-running tasks. RwChecker can run when the user becomes aware of a performance alert (e.g., an unusually long task), or periodically check for potential LHBs in long-running processes. The workflow of RwChecker is as Figure 5 shown. The framework first obtains software behavior from the application layer and system layer using hierarchical tracing. After that, RwChecker will identify potential LHBs from the tracing records using checking rules over consecutive time periods.
[0144] Refer to the results of RwChecker identifying LHBs in a real production environment in Table 1. In the experiment, RwChecker successfully identified 21 out of 24 LHBs. The results are shown in Table 1. This section classifies each case according to the rule instance that was first triggered during the monitoring period. Rules #2 and #3 are very effective in identifying LHBs, accounting for 81.0% (17 / 21) of all detected cases, where #1 - #3 represent intensive repeated code execution, accessing the same resource, and becoming ineffective after a specific operation, respectively.
[0145] Table 1 Results of RwChecker identifying LHBs in a real production environment
[0146] Software Identified LHB Rule #1 Rule #2 Rule #3 HBase 8 / 9 1 4 3 HDFS 4 / 5 1 2 1 Yarn 3 / 3 2 / 1 ZooKeeper 2 / 3 / 1 1 Hive 4 / 4 / 3 1 Total 21 / 24 4 10 7
[0147] The runtime overhead of RwChecker was evaluated through experiments on 25 test sets. The results include that the overhead of branch target tracing is on average less than 1%, and it performs relatively stably under different environments and task loads. The overhead of system call tracing is relatively high, averaging 5.9%, and it is affected by the task load. Its overhead depends on the way and frequency of interaction between the software and the environment. Since RwChecker does not instrument the software source code, all overheads are only introduced after RwChecker is started and do not affect the normal operation of the software system. Additionally, both the branch target and system call tracing processes generate log data.
[0148] Regarding the test results: RwChecker incurs a runtime overhead of 0.9% and 5.9% from branch target and system call tracing, respectively, in normal long-running tasks, which is acceptable for production runs. In particular, when the software is running normally, no additional overhead will be introduced if RwChecker is not enabled.
[0149] Figure 6 It is a diagram of a load-aware software suspension processing analysis device according to an embodiment of the present invention. The device includes a first module 610, a second module 620, a third module 630, and a fourth module 640, where:
[0150] The first module is used to detect a target application at the application layer and the system layer by using a hierarchical tracing method, and obtain an application layer tracing result and a system layer tracing result; the second module is used to determine a software reliability inspection result for the application layer tracing result and the system layer tracing result by using inspection rules and a warning triggering method; the third module is used to determine a reliability problem type according to the application reliability inspection result, where the reliability problem type includes LHB and heavy load tasks; the fourth module is used to generate a configuration item according to the reliability problem type, and optimize the reliability of the target software through the configuration item.
[0151] Exemplarily, with the cooperation of the first module, the second module, the third module, and the fourth module in the device, the embodiment device can implement any of the foregoing load-aware software suspension processing methods, that is, detect a target application at the application layer and the system layer by using a hierarchical tracing method, and obtain an application layer tracing result and a system layer tracing result; determine a software reliability inspection result for the application layer tracing result and the system layer tracing result by using inspection rules and a warning triggering method; determine a reliability problem type according to the application reliability inspection result, where the reliability problem type includes LHB and heavy load tasks; generate a configuration item according to the reliability problem type, and optimize the reliability of the target software through the configuration item. The beneficial effects of the present invention are as follows: Use the method of tracking the target address of branch instructions and lightweight system-level tracing to detect infinite loops that may cause system suspension, monitor the execution of the application program by tracking the target address of the CPU branch instruction. Compared with traditional application layer tracing, branch target tracing is supported by hardware, with relatively low overhead and no need for any instrumentation in the source code; analyze the branch target traces and system call sequences within a continuous time period, and can accurately obtain whether the application program is in LHB or a heavy load task for a long time, and perform corresponding optimizations according to the application program state, improving the reliability of the application program and achieving high reliability problem detection with low consumption.
[0152] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory;
[0153] The memory stores a program;
[0154] The processor executes a program to perform the foregoing load-aware software suspension processing method; the electronic device has the function of carrying and running the software system of the load-aware software suspension processing provided by the embodiments of the present invention. For example, a personal computer, a minicomputer, a mainframe, a workstation, a network or a distributed computing environment, a separate or integrated computer platform, or communicating with charged particle tools or other imaging devices, etc.
[0155] Embodiments of the present invention also provide a computer-readable storage medium, and the storage medium stores a program, and the program is executed by a processor to implement the load-aware software suspension processing method as described above.
[0156] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0157] Embodiments of the present invention also disclose a computer program product or a computer program, and the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the foregoing load-aware software suspension processing method.
[0158] In addition, although the present invention is described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, and the scope of the present invention is determined by the full scope of the appended claims and their equivalents.
[0159] When the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., all kinds of media that can store program codes.
[0160] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0161] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways when necessary, and then storing it in a computer memory.
[0162] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0163] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0164] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0165] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.< / string>
Claims
1. A load-aware software suspension processing method, characterized in that Including: Using a hierarchical tracing method to detect the target application at the application layer and the system layer, obtaining the application layer tracing result and the system layer tracing result; Using inspection rules and a warning triggering method to determine the software reliability inspection result for the application layer tracing result and the system layer tracing result; Determining the reliability problem type according to the application reliability inspection result, where the reliability problem type includes LHB and heavy load tasks; Generating configuration items according to the reliability problem type, and optimizing the reliability of the target software through the configuration items.
2. The load-aware software suspension processing method according to claim 1, wherein The using a hierarchical tracing method to detect the target application at the application layer and the system layer, obtaining the application layer tracing result and the system layer tracing result, includes: Using hardware tracing to obtain branch instruction records, determining the address distribution of branch instructions according to the branch instruction records, and determining the application layer tracing result according to the address distribution of branch instructions; Using the eBPF tracing method for events and system calls during the runtime of the target application to obtain the system layer tracing result, where the task types of the events and system calls include file management, communication, memory management, I / O control, process management, and synchronization control.
3. The load-aware software suspension processing method according to claim 2, characterized in that, The using hardware tracing to obtain branch instruction records and determining the address distribution of branch instructions according to the branch instruction records, includes: Using a Linux performance analysis tool to obtain branch instruction records, and after filtering the execution information of non-target applications from the branch instruction records, obtaining the mapping of the branch target addresses and call methods during the runtime of the target application.
4. The load-aware software suspension processing method according to claim 1, wherein The using inspection rules and a warning triggering method to determine the application reliability inspection result for the application layer tracing result and the system layer tracing result, includes: The inspection rules include dense and repeated code execution, accessing the same resource, and failure after a specific operation. Among them, the dense and repeated code execution determines the application behavior of the target application through the distribution of branch target addresses; accessing the same resource uses a quadruple to record the accessed resources of the target application, where the quadruple includes the accessed resource, access operation, access unit, and access content; the failure after a specific operation uses a quadruple group, where the quadruple includes the accessed resource, access operation, access unit, and error flag; Using a warning triggering method to detect the repeated behavior of each inspection rule using a time window, obtaining the application reliability inspection result.
5. The load-aware software suspension processing method according to claim 4, wherein The using a warning triggering method to detect the repeated behavior of each inspection rule using a time window, obtaining the application reliability inspection result, includes: According to the time period and the number of windows of the time window, when the inspection rule is dense and repeated code execution, collecting the branch target address distribution within a preset number of windows, determining the number of branches and branch percentages of different targets in each time window, and further determining the LHB during the execution of the target application; When the inspection rule is accessing the same resource or failure after a specific operation, detecting the first time window. If there are repeated access units, repeating the tracing record sequence of a preset number of time windows to obtain the LHB during the execution of the target application. If there are no repeated access units, resetting the start time of the time window.
6. The load-aware software suspension processing method according to claim 1, wherein Generating configuration items according to the reliability problem type and optimizing the reliability of the target software through the configuration items includes: When the reliability problem type is LHB, determining the location of the reliability problem by checking the code and optimizing LHB with the first configuration item; When the reliability problem type is a heavy-load task, optimizing the heavy-load task with the second configuration item.
7. The load-aware software suspension processing method according to claim 1, wherein The optimizing LHB with the first configuration item includes: When LHB enters a loop, adding check code, obtaining loop variables and loop termination conditions through the check code, and performing loop termination processing on LHB according to the loop variables and loop termination conditions; Alternatively, adding the maximum retry times and timeout of LHB through the first configuration item.
8. A load-aware software suspension processing device, characterized in that Including: A first module for detecting the target application at the application layer and the system layer by using a hierarchical tracing method to obtain an application layer tracing result and a system layer tracing result; A second module for determining the software reliability inspection result by using inspection rules and a warning triggering method for the application layer tracing result and the system layer tracing result; A third module for determining the reliability problem type according to the application reliability inspection result, where the reliability problem type includes LHB and heavy-load tasks; A fourth module for generating configuration items according to the reliability problem type and optimizing the reliability of the target software through the configuration items.
9. An electronic device, characterized in that, Including a processor and a memory; The memory is used for storing programs; The processor executes the program to implement the load-aware software suspension processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by the processor to implement the load-aware software suspension processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Method and system for tracking target variable in executable program
CN113672499A
Availability test method and device, electronic equipment and storage medium
CN115080438A
Function debugging method and device, storage medium and electronic equipment
CN116627850A
Load adjusting method and terminal equipment
CN117632460A
Full-link tracking analysis method and device, computer equipment and readable storage medium
CN117834267A