Linux error state capturing method and device, Linux system and storage medium

By listening to shortcut keys on Linux systems and automatically capturing memory state, running environment and log information during crashes, the problems of complex use of existing debugging tools and information loss are solved, and debugging efficiency and accuracy are improved.

CN120029882APending Publication Date: 2025-05-23天固信息安全系统(深圳)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983924.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

On Linux systems, when an application crashes or exceptions, existing debugging tools are complex to use, making it difficult to quickly capture memory state, running environment and log information, resulting in low debugging efficiency and loss of information.

Method used

By listening to preset shortcut keys, the error state capture instruction is triggered, and the memory status information, operating environment information and log information of the current application are automatically captured, and debugged and analyzed through the preset debugging tool.

Benefits of technology

Simplifies the debugging process, improves the accuracy and efficiency of capturing critical information in the event of crashes, reduces the risk of information loss, and reduces the risk of errors caused by human operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029882A_ABST
    Figure CN120029882A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a Linux error state capturing method and device, a Linux system and a storage medium, an error state capturing instruction is triggered by monitoring a preset shortcut key, and once the instruction is detected, the memory state, the running environment and the log information of a current application program are captured immediately, so that the error state of the current application program is captured. According to the technical scheme, the preset debugging tool is conveniently used for debugging analysis subsequently, manual starting and breakpoint setting are not needed, the problem that key information is lost after crash is effectively avoided, a developer can obtain complete detailed information of the crash moment, and therefore the debugging efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a Linux error state capturing method, device, Linux system and storage medium. Background Art

[0002] On Linux systems, application debugging usually relies on tools such as GDB (GNU debugger), strace, and valgrind. Although these tools are powerful, their use is complicated, especially when the program crashes or exceptions occur. Since the crash may be difficult to trigger repeatedly, developers need to manually start the debugging tool, set breakpoints, and rerun the program to reproduce the crash environment and collect relevant memory, thread, and log information. The whole process is cumbersome and time-consuming, and it is especially difficult to deal with occasional crashes. In addition, existing debugging tools rely on manual triggering and operation, and usually can only start capturing relevant status information after the program exception or crash occurs. That is, at the moment of the crash, the key memory status, thread information, and log data may have been overwritten or lost, resulting in developers being unable to obtain complete detailed information at the moment of the crash. Summary of the invention

[0003] In order to overcome the shortcomings of the prior art, the present invention provides a Linux error status capture method, device, Linux system and storage medium, which can quickly capture the memory status, operating environment and log information at the moment of crash through a simple shortcut key trigger mechanism, greatly improving the debugging efficiency and the accuracy of problem location, not only simplifying the debugging process, but also reducing the error risks caused by human operation, so that developers can focus more on solving practical problems.

[0004] A first aspect of the present application provides a Linux error state capturing method, the method comprising: Monitor whether the developer triggers the preset shortcut key to generate an error status capture instruction; When the error state capture instruction is detected, the memory state information, operating environment information and log information of the current application are captured; The memory status information, the operating environment information and the log information are debugged based on a preset debugging tool.

[0005] In an optional implementation, the memory state information includes stack information, and the method further includes: Acquire the stack information according to the captured crash signal, wherein the stack information includes a function call path when the application crashes; Parsing the stack information to obtain a function call hierarchy, wherein each level of the function call hierarchy includes a function name, a file name, and a corresponding source code line number; The call stack is traced back layer by layer according to the function call level and the function call path until the function name and source code line number of the source code where the current application crashes are located are determined.

[0006] In an optional embodiment, the method further comprises: Matching the error log in the log information according to preset keywords; The crash occurrence time of the current application is determined according to the timestamp of the error log.

[0007] In an optional implementation, before debugging the memory state information, the operating environment information, and the log information based on a preset debugging tool, the method further includes: Performing data format conversion on the memory status information, the operating environment information and the log information according to a preset data format; Packing the memory status information, the operating environment information and the log information after data format conversion into a debugging snapshot file; The debugging snapshot file is stored in a preset designated path.

[0008] In an optional embodiment, the method further comprises: Obtaining the running thread status of the current application; When it is determined that the running thread state of the target application is in an uninterruptible sleep state or a zombie state, it is determined that the current application has a thread blocking situation; the target application is any multiple application in the current application.

[0009] A second aspect of the present application provides a Linux error status capturing device, the device comprising: The shortcut key monitoring module is used to monitor whether the developer triggers the preset shortcut key to generate an error state capture instruction; An error status capture module, used to capture the memory status information, operating environment information and log information of the current application when the error status capture instruction is detected; The debugging tool integration module is used to debug the memory status information, the operating environment information and the log information based on a preset debugging tool.

[0010] A third aspect of the present application provides a Linux system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the Linux error status capture method when executing the computer program.

[0011] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned Linux error status capture method when the computer program is executed by a processor.

[0012] In summary, the Linux error status capture method, device, Linux system and storage medium provided by the present application generate error status capture instructions by monitoring preset shortcut keys, avoiding the complexity of manually starting debugging tools and setting breakpoints. Developers only need to press the preset shortcut keys when the program is running to automatically capture the error status, simplifying the debugging steps; when the error status capture instruction is detected, the memory status information, operating environment information and log information of the current application are immediately captured, ensuring that the key information at the moment of crash can be captured in time, reducing the risk of information loss, and further debugging and analyzing the captured information through the preset debugging tool, thereby improving the accuracy of capturing key information when sporadic crashes, thereby effectively solving the problem of difficult crash handling. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of a Linux error state capture method shown in an embodiment of the present application; Figure 2 It is a flowchart of a method for capturing an error state triggered by a shortcut key shown in an embodiment of the present application; Figure 3 is another process flow chart of a method for capturing an error state triggered by a shortcut key shown in an embodiment of the present application; Figure 4 It is a flowchart of an information storage packaging method shown in an embodiment of the present application; Figure 5 A functional module diagram of a Linux error status capture device shown in an embodiment of the present application; Figure 6 It is a structural diagram of a Linux system shown in an embodiment of the present application. DETAILED DESCRIPTION

[0014] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0015] The following will clearly and completely describe the concept, specific structure and technical effects of the present invention in combination with the embodiments and drawings, so as to fully understand the purpose, characteristics and effects of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, other embodiments obtained by technicians in this field without creative work are all within the scope of protection of the present invention. In addition, all the connection / connection relationships involved in the patent do not refer to the direct connection of components, but refer to the formation of a better connection structure by adding or reducing connection accessories according to the specific implementation situation. The various technical features in the invention can be combined interchangeably without conflicting with each other.

[0016] Reference Figure 1 , which is a flow chart of a Linux error status capturing method shown in an embodiment of the present application, and the Linux error status capturing method includes the following steps.

[0017] S11, monitoring whether the developer triggers a preset shortcut key to generate an error state capture instruction.

[0018] In some embodiments, the developer can pre-set a shortcut key (for example, key F1, key combination Tab+Esc) to trigger the error state capture instruction. Specifically, the developer can press the pre-set shortcut key during the application running process to trigger the error state capture instruction through the preset shortcut key. When the Linux system detects the error state capture instruction triggered by the developer, it immediately starts the error state capture process.

[0019] Among them, application crash refers to an unrecoverable error that occurs during the execution of the application, such as access to invalid memory, deadlock, division by zero error or other serious exceptions. In order to capture crash events, the Linux system is equipped with a crash detection mechanism. In the Linux environment, when the application crashes, the Linux system will send a specific signal, such as SIGSEGV (segment fault), SIGABRT (Abort signal), SIGFPE (arithmetic exception), or SIGILL (illegal instruction). The specific signal means that the application has an exception and cannot continue to execute. The Linux system can further display it to the developer through the developer interaction interface, and the developer can trigger the preset shortcut key. In other embodiments, the Linux system can capture whether a specific signal is generated during the operation of the application through programming code (such as signal function or sigaction function). Once the crash signal is captured, the application will execute the corresponding exception handling process. In particular, the developer can insert some custom error handling logic when capturing the signal, such as calling memory dump (core dump) generation, logging and other operations. Once the application crash is detected, the programming program of the Linux system detects the crash event by capturing signals (such as SIGSEGV, SIGABRT, etc.). At this time, the Linux system will respond and start the error state capture process. By setting a signal processing function or an exception handling mechanism, when an application crashes, the Linux system does not terminate the application directly, but instead executes a custom capture task.

[0020] In order to facilitate developers to quickly debug when the program is abnormal or crashes, the Linux system can continuously run the shortcut key monitoring module in the background to monitor specific shortcut keys. When it is detected that the developer presses the preset shortcut key, the system will immediately start the capture process. Even if the program crashes, the developer can actively trigger the capture process by pressing the shortcut key. Specifically, the shortcut keys pre-set by the developer can be monitored at the Linux system level, which can be a single shortcut key (such as F1) or a shortcut key combination (such as Tab+Esc). Among them, the shortcut key monitoring module runs in the background of the Linux system, waiting for the developer's input. During the running of the application, the Linux system continuously monitors whether the shortcut key is triggered through the shortcut key monitoring module. When it is detected that the shortcut key is triggered, the error state capture module is notified to start working. That is, when the developer presses the preset shortcut key during the running of the application, the shortcut key monitoring module can detect the shortcut key signal and send a trigger notification to the error state capture module. When the shortcut key signal, that is, the error state capture instruction, is detected, the Linux system first suspends other interactions of the application to prevent interference during the error capture process.

[0021] S12, when the error status capture instruction is detected, the memory status information, operating environment information and log information of the current application are captured.

[0022] After receiving the error status capture instruction triggered by the shortcut key, the error status capture module performs the error status capture operation, including the tasks of capturing memory status information, collecting runtime environment information (called runtime environment information) and collecting log information, to help developers understand the running status of the application when it is abnormal or crashes.

[0023] Refer to Figure 2 and Figure 3 , the error status capture module performs the following steps in sequence: Step S21, capturing the memory status information of the current application.

[0024] Among them, the memory status information includes, but is not limited to, stack information and key variable values. In some embodiments, the Linux system can obtain the stack information in the memory status information by accessing the memory image of the application or through a debugging tool. Specifically, when the error status capture module receives the error status capture instruction, the execution of the application is suspended, and a preset debugging tool (such as GDB, strace, Valgrind, etc.) is called to access the memory of the application. In addition, when the application crashes (such as due to a segmentation error, illegal memory access, etc.), the Linux system can generate corresponding crash information (such as SIGSEGV), and the Linux system can capture the crash signal provided by the Linux system through backtrace() (in C / C++) and debugging tools such as GDB to obtain the current stack information.

[0025] In order to more clearly understand the inventive concept of the present application, the debugging tool in the embodiment of the present application is explained using GDB as an example.

[0026] When using GDB to obtain stack information, use the bt command to obtain the call stack (backtrace). The stack information contains the function call path when the application crashes. Further, the stack information is parsed to determine the function call level (i.e., call chain) it contains. Each level of the function call level will display the function name, file name, and corresponding source code line number.

[0027] For example, GDB output might be: #0foo() at example.c:35 #1bar() at example.c:23 #2main() at example.c:10 Next, the Linux system traces back the call stack layer by layer according to the function call path and the function call level to find the code location where the crash occurred (for example, line 35 of the foo() function), or uses debugging symbols (such as the -g compilation option) to obtain the specific source code line number, that is, the crash point. When the crash point is determined, the call stack is traced back layer by layer starting from the function level where the crash point is located. In each level, the function name, file name, and source code line number are checked to confirm whether the current level is the target crash point. If not, continue to trace back to the previous caller. When tracing back to a certain level, if it is confirmed that the function name, file name, and source code line number of the level match the expected crash point, then stop tracing back, and the function name and source code line number of the level are the source code location that caused the application to crash. In other embodiments, the Linux system can also check the context of the function call to confirm whether there are problems such as null pointer dereference, out-of-bounds access, or resource release errors. In addition, if the stack information contains calls to library functions or system functions, the Linux system can further analyze the internal implementation of functions such as library functions or system functions to determine whether the application crash was caused by functions such as library functions or system functions.

[0028] During the backtracing process, the Linux system can also obtain the key variable values ​​in the application through GDB. In GDB, you can use info locals and info args to view the key variable values ​​of the current function, where the key variable values ​​can include the status of local variables, global variables, and registers, etc., to determine the specific running status of the application when it crashes or is abnormal. Through the key variable values, the abnormal cause of the application crash can be determined. For example, the program may crash due to null pointers, invalid memory references, inconsistent data structures, etc. Specifically, the Linux system can determine whether the key variable values ​​meet the preset parameter conditions, that is, whether the key variable values ​​are legal, including null pointer checks, array out-of-bounds checks, uninitialized variable checks, and combined code logic checks.

[0029] (1) Null pointer check.

[0030] During program execution, a null pointer check mechanism is implemented for all code segments involving pointer operations. Specifically, when the program attempts to access memory through a pointer or perform a pointer dereference operation (for example, *ptr=10), it first determines whether the pointer is null (i.e., equal to NULL) or points to an illegal address (such as 0xFFFFFFFF). If the check finds that the pointer is null or points to an illegal address, it is determined that the current application has crashed.

[0031] (2) Array out-of-bounds check.

[0032] For all code segments in the program that involve array or buffer access, implement an array out-of-bounds check mechanism. Before accessing an array element, first determine whether the array index exceeds the valid range of the array. That is, if the program crashes when accessing an array or buffer, check whether the array index is out of bounds. If the check finds that the index is out of bounds, immediately terminate the current access and trigger the error handling process to determine that the current application has crashed.

[0033] (3) Check for uninitialized variables.

[0034] When the program is compiled and run, all local variables are checked for uninitialization. Make sure that any local variable has been correctly assigned a value before using it. Uninitialized variables may cause undefined behavior in some cases, leading to crashes. Therefore, when obtaining the value of a key variable, you can determine whether the key variable value has been correctly assigned before use. (4) Combine code logic checking.

[0035] During program execution, continuously verify the value of the variable in combination with the code logic. For each key variable, set reasonable checkpoints and verification rules based on its expected scope and purpose. When the key variable value is obtained, combine the key variable value with the code logic to confirm whether the key variable value meets expectations. For example, if a variable value is greater than the expected range (such as a negative number) when the program crashes, it may be that some boundary conditions are not handled correctly, causing the application to crash.

[0036] In other embodiments, the Linux system can obtain memory usage to analyze the abnormal cause of the application crash based on the memory usage. Among them, the memory usage reflects the memory allocation and usage status when the application crashes, and can reveal problems such as memory overflow and memory leak. Specifically, the Linux system can use debugging tools such as Valgrind (for memory leak detection) and GDB (for checking memory) to capture the memory usage of the application. Through the debugging tool, the Linux system can obtain the memory stack information at the time of the crash and view the allocation and release of each object.

[0037] For example, valgrind can detect if your application has memory leaks or out-of-bounds accesses: ==12345== ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from0) ==12345== malloc / free: in use at exit: 20 bytes in 1 blocks. If the application crashes due to heap overflow or stack overflow, it can be confirmed by checking the size of the stack or the memory allocation. The memory overflow analysis can include stack overflow and stack overflow, while the memory leak analysis can include tool-supported memory check and manual memory check. Specifically, when the application crashes due to stack overflow, first check the depth of the stack, and analyze the stack call chain at the time of the crash based on the stack information to confirm whether there is a recursive call that is too deep or insufficient stack space allocation. If the recursive call is found to be too deep, the recursive algorithm should be optimized to reduce the recursive depth, or use an iterative algorithm instead. If the stack space allocation is insufficient, the stack space size should be adjusted or the function call structure should be optimized. When the application crashes due to heap overflow, use a memory analysis tool (such as valgrind) to check the heap allocation of the application, and analyze the memory allocation record before the crash to confirm whether there is over-allocation of memory or improper memory management. If over-allocation of memory is found, the memory usage strategy should be optimized to reduce unnecessary memory allocation. If the memory is found to be improperly managed, the matching of memory allocation and release operations should be checked to ensure that each malloc() or new operation is followed by a corresponding free() or delete operation. For tool support analysis, use memory leak detection tools (such as valgrind, AddressSanitizer) to perform static or dynamic analysis on the application to check whether there is a memory leak problem. The memory leak detection tool can detect whether the application does not release memory correctly during operation, resulting in memory overflow. Then, based on the tool's detection results, locate the source of the memory leak, that is, the unreleased memory block and its allocation location. For manual memory detection, manually check the matching of memory management operations at the critical path of the application and the memory allocation operation before the crash to ensure that each malloc() or new operation is followed by a corresponding free() or delete operation, and the operation order is correct. For complex memory management logic, such as dynamic arrays, linked lists and other data structures, special attention can also be paid to their memory allocation and release strategies to ensure that memory leaks do not occur.

[0038] Through the above optional implementation, by capturing the memory status information of the application, including stack information and key variable values, and using debugging tools such as GDB for backtrace analysis, the source code location of the application crash and the cause of the exception are determined. Through mechanisms such as null pointer checks, array out-of-bounds checks, uninitialized variable checks, and combined code logic checks, potential crash problems can be effectively identified. At the same time, memory usage is analyzed to reveal problems such as memory overflow and memory leaks. This solution improves the stability and reliability of the application, provides developers with detailed error diagnostic information, and helps to quickly locate and fix crash problems.

[0039] Step S22, collecting operating environment information.

[0040] The operating environment information includes system resource usage (which may include CPU usage, memory occupancy, disk I / O, etc.) and active thread information to determine whether the application has caused an exception due to excessive consumption of system resources. System resource usage refers to the occupation of system resources (such as CPU, memory, disk, and network, etc.) by the application during operation. Specifically, the Linux system can obtain the CPU usage (that is, CPU usage) by accessing the / proc / stat file, where the / proc / stat file contains statistical information on various CPU states since the system was started (such as developer space, system space, idle time, etc.), and calculates the CPU load based on the CPU statistical information, and then determines whether the program crashed due to excessive consumption of CPU resources. At the same time, the Linux system can access the / proc / meminfo file, which provides information such as the total amount of system memory, free memory, used memory, cache memory, etc. By reading the memory usage data in the / proc / meminfo file, including MemTotal (total system memory), MemFree (current free memory), MemAvailable (system available memory), Buffers (system buffer memory) and Cached (system cache memory), and analyzing the read data, the memory usage of the application can be determined, and whether there are problems such as insufficient memory, memory leaks, and excessive memory allocation. At the same time, the Linux system can also use the df command or read the / proc / diskstats file to obtain disk usage data, by analyzing the disk read and write rate, disk space usage, and whether the disk is full. At the same time, the Linux system can also use the netstat command or read the / proc / net / dev file to obtain the current network connection and traffic usage data, and by analyzing the traffic statistics of the network interface, determine whether the program crash or abnormality is caused by a network bottleneck.

[0041] In a multi-threaded program, the Linux system can also obtain the status information of each thread by accessing the / proc / [pid] / task / [tid] / status file, or directly use the ps command to list the current active threads, and analyze the abnormal reasons for the application crash based on the thread information, such as thread synchronization errors (such as deadlocks and race conditions). Specifically, when the application crashes, you can use GDB's info threads command to view all threads of the current program and their status, and use threadapply all bt to view the stack information of all threads. If the thread is in the waiting state, there may be a deadlock in the application. You can further analyze whether there is resource competition by checking the thread holding the lock, and confirm whether there are multiple threads modifying the same shared resource at the same time through multi-thread stack information, causing the crash.

[0042] In an optional embodiment, the method further comprises: Obtaining the running thread status of the current application; When it is determined that the running thread state of the target application is in an uninterruptible sleep state or a zombie state, it is determined that the current application has a thread blocking situation; the target application is any multiple application in the current application.

[0043] In some embodiments, the Linux system can read the running thread state of the corresponding thread of the current application and analyze whether the thread is in the S (sleep), R (run), Z (zombie), D (uninterruptible sleep) and other states, where the R state indicates that the thread is running, the S state indicates that the thread is sleeping, the Z state indicates that the thread is a zombie thread, which has been terminated but not recycled by the parent process, and the D state indicates that the thread is performing operations such as disk I / O, which may cause blocking. If there are multiple threads, that is, the threads corresponding to multiple applications (referred to as target applications) are in the D state or Z state, it may indicate that the application is deadlocked or the thread is blocked.

[0044] When the CPU usage, memory usage, thread status and other information are obtained, they are aggregated into a structured data file (such as JSON, XML or custom format). For example: { "cpu_usage": 75.5, "mem_usage": { "total": "16GB", "used": "8GB", "free": "4GB" }, "threads": [ {"tid": 1234, "status": "R", "cpu_usage": 10}, {"tid": 5678, "status": "D", "cpu_usage": 20} ] } In some embodiments, the Linux system can also provide developers with a graphical interface or command line tool to display the operating environment data. For example, using graphical tools (such as Grafana, Prometheus) to display the real-time changes of CPU, memory, threads and other data to help developers better understand the resource usage of the application.

[0045] Through the above optional implementation, by comprehensively collecting operating environment information, including CPU usage, memory usage, disk I / O, network conditions, and active thread status, key data is provided for diagnosing application anomalies. By monitoring and analyzing this data, crashes caused by problems such as excessive system resource consumption or thread blocking can be quickly located. At the same time, the information is aggregated into structured data files and displayed in a graphical interface or command line tool, which enhances the readability and ease of use of the data, helps developers intuitively understand the operating status of the application, and thus take effective measures to optimize resource usage and thread management, and improve the stability and performance of the application.

[0046] Step S23, extracting the latest log information of the application.

[0047] In some embodiments, when an application crashes, the Linux system usually generates a core dump file or records the crash information in the application's log file. Through logging, the stack trace, error message, memory status and other information of the crash can be obtained. In order to ensure that the exception of the application can be captured, the developer usually integrates a log module inside the program to capture key log information during the execution of the program, especially when an error occurs. In some embodiments, when the application is abnormal or crashes, the latest log information (error log, stack information, warning log, etc.) is obtained to quickly locate the abnormal location of the application. Among them, the log information exists in the form of a file, recording important information of the application during operation, including errors, exceptions, warnings, debugging information, etc. Specifically, after obtaining the memory status information and system resource information, the Linux system can automatically extract the latest log information from the application's log file (for example, the default is to obtain the latest 100 lines, which can be configured). The log information includes events before and after the abnormal or crash event occurs, so as to determine the root cause of the application abnormality or crash based on the log information (such as program logic errors, external service call failures, etc.). It should be noted that the developer can customize the log range, such as collecting the logs of the latest 500 lines or a specified time period by configuration. And the collection of log information can be performed after the collection of memory status information and operating environment information is completed. Because log files are usually part of the file system, reading log files is relatively lightweight. It is only necessary to ensure that the log files are not occupied or being modified by other processes.

[0048] In an optional embodiment, the method further comprises: Matching the error log in the log information according to preset keywords; The crash occurrence time of the current application is determined according to the timestamp of the error log.

[0049] In some embodiments, the Linux system can pre-configure log sources and automatically extract log information from the configured log sources. Then, it uses keyword matching (such as "ERROR", "CRITICAL", etc.) to identify error logs in the stored log information. Based on the error codes corresponding to the error logs, the types of application crashes or exceptions can be determined, such as uncaught exceptions, resource exceptions, and logical errors. Among them, uncaught exceptions refer to situations like null pointer references, array out-of-bounds, database connection failures, etc.; resource exceptions refer to insufficient memory, file permission errors, network connection failures, etc.; and logical errors refer to crashes caused by incorrect business logic. Further, the timestamp in the error log is determined, and based on the timestamp, the exact time of the crash or exception occurrence can be determined, that is, the crash time of the current application. Subsequently, the Linux system can continuously monitor the log information to analyze whether there are new application crashes or exceptions. Through the above optional implementation methods, by extracting the latest log information of the application, especially error logs, stack information, and warning logs, it provides an important basis for quickly locating the root cause of application exceptions or crashes. Through keyword matching and timestamp analysis, the error type can be accurately identified and the crash time can be determined, effectively improving the efficiency and accuracy of problem diagnosis. In addition, it also supports customizing the log range and configuring log sources, enhancing flexibility and applicability, helping developers take timely measures to fix problems and ensuring the stable operation of the application.

[0050] It should be noted that the execution order of the above steps is carried out step by step in a specific order. Since capturing memory status information requires direct interaction with the application's memory space, especially when the application crashes, it may involve locking or pausing part of the program's resources (such as threads). If multiple tasks are performed at the same time, that is, the above steps are processed in parallel, it may cause resource competition (such as memory, CPU, I / O, etc.), resulting in inaccurate capture results or system instability. The log information is continuously updated during the writing process, so when capturing log information in real time, it is necessary to ensure the correctness of file reading and writing. If multiple tasks are performed at the same time, the log information may not be fully recorded or lost. In addition, when capturing memory status information, it is necessary to rely on specific tools or memory images. If the log information is captured while capturing the memory status information, part of the memory content may not be saved in the log information in time, or there may be a time misalignment between the log information and the memory status information. The call of GDB is usually executed after all key information is captured. If the debugging tool is called in advance, it may not be possible to fully load all the required data, thus affecting the debugging effect. Tasks such as capturing memory status information, collecting operating environment information, and extracting log information need to be coordinated at the specific moment when the application crashes or anomalies occur. If these tasks are performed simultaneously, the system will become complex and difficult to coordinate, which may lead to missing, duplication or inconsistency of information.

[0051] In an optional implementation, before debugging the memory state information, the operating environment information, and the log information based on a preset debugging tool, the method further includes: Performing data format conversion on the memory status information, the operating environment information and the log information according to a preset data format; Packing the memory status information, the operating environment information and the log information after data format conversion into a debugging snapshot file; The debugging snapshot file is stored in a preset designated path.

[0052] There is a lack of seamless integration between existing tools. For example, the captured memory status information needs to be manually exported and imported into other tools for further analysis, which increases the complexity of debugging and the possibility of errors. Figure 4In an embodiment of the present application, the Linux system first creates a temporary file structure containing memory status information, operating environment information, and log information. After the error status capture module completes the corresponding error status capture task, that is, after capturing the memory status information, runtime environment information, and log information, the captured information is converted into a preset data format (such as JSON, XML, or a custom binary format) using a data packaging module, and written into a temporary document corresponding to the pre-created temporary file structure. The temporary file is then compressed to save storage space and renamed to a .debugsnap file, that is, integrated and packaged into a debug snapshot file. Finally, the .debugsnap file is saved in a specified directory, or automatically stored according to a preset naming rule. Through the standardized .debugsnap file format, compatibility with different debugging tools is ensured, which facilitates sharing of debugging information within a team and improves collaboration efficiency.

[0053] When the data is packaged, the debugging tool integration module automatically starts the debugging tool according to the configuration to load the .debugsnap file, parse the stack information, memory status information, etc., to help developers analyze the cause of the program crash. In GDB, the Linux system can directly load and analyze the program status through the generated script or manual command to quickly locate the problem.

[0054] Through the above optional implementation, the standardized integration of memory status information, operating environment information and log information is achieved through data format conversion and packaging, and a debug snapshot file (.debugsnap) compatible with different debugging tools is generated, which reduces manual operations and error rates, while improving the sharing and collaboration efficiency of debugging information. Among them, the standardized file format ensures the consistency and accuracy of debugging information, allowing developers to use debugging tools to quickly locate and analyze the cause of program crashes, accelerate the problem repair process, and improve the stability and reliability of the application.

[0055] S13, debugging the memory status information, the operating environment information and the log information based on a preset debugging tool.

[0056] Furthermore, the captured information is integrated with GDB using the debugging tool integration module, including functions such as generating debugging scripts and automatically starting debugging sessions. Among them, the debugging tool integration module automatically calls the existing debugging tool to load the .debugsnap file according to the current task requirements, and assists developers in further debugging. Specifically, the memory status information, operating environment information, and log information in the .debugsnap file are obtained through GDB, and the memory status information, operating environment information, and log information are analyzed to determine whether the application has crashed or an exception has occurred. When a crash or an exception is determined, the debugging process, that is, the exception repair process, is entered. Among them, the exception repair process includes: (1) Code fixes.

[0057] In some embodiments, the Linux system can locate the specific code line where the crash and exception occurred according to the stack information and the misalignment log according to the above steps, and can fix common problems such as logic errors, null pointer references, array out-of-bounds, etc. in the code. In addition, an exception handling mechanism can be added to the key parts of the application, such as try-catch statements, to avoid program crashes, while optimizing resource management, ensuring that no longer needed resources are released in a timely manner, and using performance analysis tools to detect and fix memory leaks.

[0058] (2) System configuration and environment repair.

[0059] In some embodiments, the Linux system adjusts the system resource configuration according to the insufficient resources displayed by the log information, such as increasing the server memory, optimizing the application memory usage, or increasing the thread pool size; at the same time, adjusts the resource limits of the operating system, such as the maximum number of file descriptors, the maximum memory usage, etc., to support the high-load operation of the program, and checks and optimizes the external service configurations that the program operation depends on, such as the database connection pool configuration, the timeout setting of the API call, the network bandwidth limit, etc.; (3) Redesign the program architecture.

[0060] In some embodiments, the Linux system can reduce the dependency between modules and reduce the complexity of the system by modularizing the program with complex business logic or tightly coupled modules. In addition, the Linux system can also enhance the logging function of the program, record more key information, and help locate problems quickly. In other embodiments, the Linux system can also introduce a real-time monitoring system to monitor the health of applications and systems, and respond to potential problems in a timely manner through a real-time alarm mechanism, and design a more robust fault-tolerant mechanism, such as a retry strategy, downgrade processing, automatic recovery, etc., to ensure that the program can continue to run stably when an exception occurs.

[0061] Through the above optional implementation methods, through integrated debugging tools, the captured information can be efficiently used to locate and repair application crashes or exceptions, covering comprehensive optimization of code, system configuration, environment and program architecture, thereby improving problem repair speed, enhancing system stability, reducing maintenance costs, and ensuring high availability of the application and user experience.

[0062] In some embodiments, the Linux system provides a concise user interface, allowing developers to customize shortcut key combinations, select the log range to be captured, set the data storage path, etc. At the same time, the interface can display the capture progress and status to ensure that users understand the capture process.

[0063] In order to facilitate understanding of the inventive concept of the present application, the present application provides a specific embodiment for illustration. Suppose a developer is developing a Linux application and encounters an occasional crash during the program operation. The traditional debugging method requires manually starting GDB, setting breakpoints, and reproducing the crash environment. However, using the "Quick Debug Snapshot" of the present invention, the developer only needs to press a preset shortcut key when the program crashes, and the system immediately captures the current memory status, operating environment information, and the latest log information, and generates a .debugsnap file. At the same time, the system automatically generates a GDB debugging script and starts GDB, loads the debugging snapshot, and the developer can start debugging immediately without manual settings, which greatly improves the debugging efficiency.

[0064] Reference Figure 5 , which is a functional module diagram of a Linux error status capture device shown in an embodiment of the present application.

[0065] In some embodiments, the Linux error state capture device 50 may include a plurality of functional modules composed of computer program segments. The computer programs of the various program segments of the Linux error state capture device 50 may be stored in a memory of the Linux system and executed by at least one processor to perform (see Figure 1 1. The function of capturing the Linux error status is described in detail. According to the functions performed by the module, it can be divided into multiple functional modules. The functional modules may include: a shortcut key monitoring module 501, an error status capturing module 502, a debugging tool integration module 503 and a data packaging module 504. The module referred to in this application refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0066] The shortcut key monitoring module 501 is used to monitor whether the developer triggers the preset shortcut key to generate an error state capture instruction. The error status capturing module 502 is used to capture the memory status information, operating environment information and log information of the current application program when the error status capturing instruction is detected.

[0067] The debugging tool integration module 503 is used to debug the memory state information, the operating environment information and the log information based on a preset debugging tool.

[0068] The error status capture module 502 is also used to: obtain the stack information according to the captured crash signal, the stack information includes the function call path when the application crashes; parse the stack information to obtain the function call hierarchy, wherein each level of the function call hierarchy includes the function name, file name and corresponding source code line number; and trace back the call stack layer by layer according to the function call hierarchy and the function call path until the function name and source code line number of the source code where the current application crashes are determined.

[0069] The error status capturing module 502 is further used to: match the error log in the log information according to preset keywords; and determine the crash occurrence time of the current application according to the timestamp of the error log.

[0070] The data packaging module 504 is used to: convert the data format of the memory status information, the operating environment information and the log information according to a preset data format; package the memory status information, the operating environment information and the log information after the data format conversion into a debugging snapshot file; and store the debugging snapshot file to a preset designated path.

[0071] The error status capture module 502 is also used to: obtain the running thread status of the current application; when it is determined that the running thread status of the target application is in an uninterruptible sleep state or a zombie state, determine that the current application has a thread blocking situation; the target application is any multiple applications in the current application.

[0072] It should be understood that the various variations and specific embodiments of the Linux error status capture method provided in the above embodiments are also applicable to the Linux error status capture device of the present embodiment. Through the above detailed description of the Linux error status capture method, those skilled in the art can clearly know the implementation method of the Linux error status capture device in the present embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0073] See also Figure 6FIG. 1 is a schematic diagram of the structure of a Linux system according to an embodiment of the present application. In a preferred embodiment of the present application, the Linux system 6 includes a memory 61 , at least one processor 62 and at least one communication bus 63 .

[0074] Those skilled in the art should understand that Figure 6 The structure of the Linux system shown does not constitute a limitation of the embodiments of the present application, and can be either a bus structure or a star structure. The Linux system 6 can also include more or less other hardware or software than shown in the figure, or a different component arrangement.

[0075] In some embodiments, the Linux system 6 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices. The Linux system 6 may also include developer equipment, which includes but is not limited to any electronic product that can interact with developers through keyboards, mice, remote controls, touch pads, or voice-controlled devices, such as personal computers, tablet computers, smart phones, digital cameras, etc.

[0076] In the above embodiments provided in the present application, it should be understood that the disclosed methods, devices, computer-readable storage media and Linux systems can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple components or modules can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or components or modules, which can be electrical, mechanical or other forms.

[0077] The components described as separate components may or may not be physically separated, and the components shown as components may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the components may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0078] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each component may exist physically separately, or two or more modules may be integrated into one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0079] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0080] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0081] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0082] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A Linux error state capture method, characterized in that: The method comprises: Monitor whether the developer triggers the preset shortcut key to generate an error status capture instruction; When the error state capture instruction is detected, the memory state information, operating environment information and log information of the current application are captured; The memory status information, the operating environment information and the log information are debugged based on a preset debugging tool.

2. The Linux error state capturing method according to claim 1, characterized in that: The memory status information includes stack information, and the method further includes: Acquire the stack information according to the captured crash signal, wherein the stack information includes a function call path when the application crashes; Parsing the stack information to obtain a function call hierarchy, wherein each level of the function call hierarchy includes a function name, a file name, and a corresponding source code line number; The call stack is traced back layer by layer according to the function call level and the function call path until the function name and source code line number of the source code where the current application crashes are located are determined.

3. The Linux error state capturing method according to claim 1, characterized in that: The method further comprises: Matching the error log in the log information according to preset keywords; The crash occurrence time of the current application is determined according to the timestamp of the error log.

4. The Linux error state capturing method according to claim 1, characterized in that: Before debugging the memory state information, the operating environment information, and the log information based on the preset debugging tool, the method further includes: Performing data format conversion on the memory status information, the operating environment information and the log information according to a preset data format; Packing the memory status information, the operating environment information and the log information after data format conversion into a debugging snapshot file; The debugging snapshot file is stored in a preset designated path.

5. The Linux error state capturing method according to claim 1, characterized in that: The method further comprises: Obtaining the running thread status of the current application; When it is determined that the running thread state of the target application is in an uninterruptible sleep state or a zombie state, it is determined that the current application has a thread blocking situation; the target application is any multiple application in the current application.

6. A Linux error status capture device, characterized in that: The device comprises: The shortcut key monitoring module is used to monitor whether the developer triggers the preset shortcut key to generate an error state capture instruction; An error status capture module, used to capture the memory status information, operating environment information and log information of the current application when the error status capture instruction is detected; The debugging tool integration module is used to debug the memory status information, the operating environment information and the log information based on a preset debugging tool.

7. A Linux system, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the Linux error state capturing method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the Linux error status capturing method according to any one of claims 1 to 5 are implemented.