Methods, apparatus, equipment and storage media for detecting abnormal dump files

By determining the dependencies of tasks in the storage system and comparing dump files, abnormal dump files can be identified and located, solving the problem that existing testing strategies struggle to discover defects in complex scenarios and enabling rapid fault location and diagnosis.

CN120610843BActive Publication Date: 2025-10-31INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511065361.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-31
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing automated testing strategies for single-type operations in storage systems are insufficient to fully uncover potential defects under complex loads or edge scenarios, and running faulty tasks may cause node crashes.

Method used

By determining the dependencies of tasks to be processed, separating basic tasks from faulty tasks, obtaining and comparing dump files during the operation, and identifying abnormal dump files to locate the fault point.

Benefits of technology

Quickly locate the fault point, save the complete context of the fault occurrence, facilitate problem reproduction and diagnosis, and verify the system's ability to handle anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610843B_ABST
    Figure CN120610843B_ABST
Patent Text Reader

Abstract

This application provides a method for detecting abnormal dump files, which can be applied to the field of big data technology. The method includes: determining the execution order of multiple tasks to be processed based on their dependencies, wherein the multiple tasks include at least one basic task and at least one faulty task; obtaining multiple first dump files and multiple second dump files, wherein the first dump files are generated during the execution of at least one basic task in the execution order, and the second dump files are generated during the execution of at least one faulty task in the execution order; comparing the multiple second dump files and the multiple first dump files according to the fault type of the at least one faulty task, and identifying the abnormal dump file from the multiple second dump files. This application also provides an apparatus, device, and storage medium for detecting abnormal dump files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, specifically to a method, apparatus, device, and storage medium for detecting abnormal dump files. Background Technology

[0002] As storage device systems continue to develop and improve, their architecture and functions become increasingly complex, and automated testing strategies for single-type operations can no longer meet the needs of defect detection.

[0003] Automated testing strategies for single-type operations of storage systems can be broadly categorized into two types: one is running a single basic service, which has limited test coverage and makes it difficult to fully discover potential defects in the system under complex loads or edge scenarios; the other is running faulty tasks, which have a relatively simple operation method and perform faulty operations on nodes multiple times in a short period of time, which can cause the nodes to enter an abnormal state and crash. Summary of the Invention

[0004] In view of the above problems, this application provides a method, apparatus, device and storage medium for improving the detection efficiency of abnormal dump files.

[0005] The first aspect of this application provides an anomaly dump file method, comprising: determining the execution order of multiple tasks to be processed based on dependencies between them, wherein the multiple tasks to be processed include at least one basic task and at least one faulty task; obtaining multiple first dump files and multiple second dump files, wherein the first dump files are dump files generated during the execution of at least one basic task in the execution order, and the second dump files are dump files generated during the execution of at least one faulty task in the execution order; and comparing the multiple second dump files and the multiple first dump files based on the fault type of the at least one faulty task to determine an anomaly dump file from the multiple second dump files.

[0006] A second aspect of this application provides an abnormal dump file detection device, comprising: a task sorting module, configured to determine the running order of each task to be processed according to the dependency relationship between multiple tasks to be processed, wherein the multiple tasks to be processed include at least one basic task and at least one faulty task; a dump acquisition module, configured to acquire multiple first dump files and multiple second dump files, wherein the first dump files are multiple dump files generated during the running of at least one basic task in the running order, and the second dump files are multiple dump files generated during the running of at least one faulty task in the running order; and an anomaly detection module, configured to compare the multiple second dump files and the multiple first dump files according to the fault type of the at least one faulty task, and determine an abnormal dump file from the multiple second dump files.

[0007] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0008] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0009] The anomaly dump file detection method based on the embodiments of this application can quickly locate the fault point by comparing multiple dump files triggered during the operation of the basic task and the fault task. The first dump file and the second dump file save the complete context (such as memory and stack) when the fault occurs, which is convenient for reproducing and diagnosing the problem. By actively running the fault task and collecting the dump files, the system's ability to handle anomalies can be verified. Attached Figure Description

[0010] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0011] Figure 1 The illustration schematically depicts application scenarios of the abnormal dump file detection method, apparatus, device, medium, and program product according to embodiments of this application;

[0012] Figure 2 A flowchart illustrating an abnormal dump file detection method according to an embodiment of this application is shown schematically.

[0013] Figure 3 The schematic diagram illustrates a test flowchart in which the basic business and fault tasks according to an embodiment of this application are not parallel.

[0014] Figure 4 The diagram illustrates a test flowchart in which basic business and fault tasks are performed in parallel according to an embodiment of this application.

[0015] Figure 5 This schematically illustrates a structural block diagram of an abnormal dump file detection device according to an embodiment of this application; and

[0016] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing an abnormal dump file detection method according to an embodiment of this application. Detailed Implementation

[0017] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0020] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0021] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0022] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0023] Figure 1 The illustration shows an application scenario of the abnormal dump file detection method, apparatus, device, and storage medium according to embodiments of this application.

[0024] like Figure 1 As shown, a server may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a microprocessor (MCU) or a field-programmable gate array (FPGA)) and a memory 104 for storing data are also shown. The server may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server described above. For example, the server may also include components that are more complex than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0025] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the link congestion control method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0027] The abnormal dump file detection method in this embodiment can be executed by a server. Here, the server refers to the entire server, including the relevant components and processors within the server that need to execute the abnormal dump file detection method. Alternatively, the abnormal dump file detection method in this embodiment can be executed by the processor 102. The dump file acquisition process in the abnormal dump file detection method in this embodiment can be executed by the memory 104. In some examples of this embodiment, the abnormal dump file detection method is described using the execution by the server as an example.

[0028] The abnormal dump file detection method in the embodiments of this application can be applied to scenarios involving the detection of abnormal dump files. By analyzing the error codes, exception types (such as access violations, division by zero errors) and stack traces in the dump file, the abnormal dump file can quickly identify the module or line of code that caused the crash.

[0029] The following will be based on Figure 1 The described scene, through Figures 2-4 The abnormal dump file detection method of the disclosed embodiments is described in detail.

[0030] Figure 2 A flowchart illustrating an abnormal dump file detection method according to an embodiment of this application is shown schematically.

[0031] like Figure 2 As shown, the abnormal dump file detection method of this embodiment includes operations S210 to S230.

[0032] In operation S210, the execution order of each task to be processed is determined according to the dependency relationship between multiple tasks to be processed. The multiple tasks to be processed include at least one basic task and at least one fault task.

[0033] For example, multiple tasks to be processed can be obtained in ways including but not limited to: obtaining them from the test framework's configuration file, dynamically creating a task list through algorithms or rules, retrieving them from the message queue of a distributed system, or extracting them from a database table.

[0034] In the embodiments of this application, a task list containing all pending tasks is dynamically created by rules, and multiple pending tasks are obtained from the task list.

[0035] The system receives multiple pending tasks and categorizes each task into at least one basic task and at least one fault task based on its task type. Basic tasks represent normal business process tasks, while fault tasks represent tasks related to fault handling.

[0036] For example, basic tasks are the core functions or basic operations that must be performed when the system is running normally, used to build a stable initial state as a benchmark for subsequent fault testing, including starting Controller Area Network (CAN) bus communication, initializing sensors, loading drivers, allocating memory pools, starting virtual machines, and deploying containers.

[0037] For example, fault scenarios are designed to simulate abnormal operations or failure scenarios that the system may encounter, and to verify its fault tolerance or abnormal handling mechanisms. These include simulating sensor signal loss, injecting error controller LAN messages, forcibly triggering kernel panic, randomly terminating node processes, and simulating network partitioning.

[0038] For example, dependencies include a basic task depending on a faulty task.

[0039] Multiple tasks to be processed are divided into basic tasks and fault tasks to ensure that the detection results of fault tests are caused only by fault tasks, rather than by the instability of the basic tasks themselves.

[0040] In operation S220, multiple first dump files and multiple second dump files are obtained. The first dump files are dump files generated during the execution of at least one basic task in the execution order, and the second dump files are dump files generated during the execution of at least one faulty task in the execution order.

[0041] For example, the first dump file is triggered at a critical point in the basic task's execution (such as the start or end of the task) and is used to save information such as memory, variables, stack, and logs during the normal operation of the task.

[0042] For example, the second dump file is triggered when a fault is introduced during the execution of the faulty task or when a fault (such as a crash or exception capture) occurs in a node. It is used to save the context information at the time of the fault, such as the error stack and memory snapshot.

[0043] For example, a node can be a host, router, server, mobile terminal, computer terminal, or other device.

[0044] For example, the execution order of each task to be processed includes: prioritizing the execution of basic tasks that depend on the faulty task, followed by executing faulty tasks that can be run in parallel. By properly sequencing these tasks, the impact of independent sensor failures can be verified, and the system-level fault tolerance can also be tested.

[0045] In operation S230, based on the fault type of at least one faulty task, multiple second dump files and multiple first dump files are compared, and an abnormal dump file is determined from the multiple second dump files.

[0046] For example, the fault type of a faulty task can be determined by identifying the dump file directory, the location where the dump file was generated, or the time when the dump file was generated.

[0047] The anomaly dump file detection method based on the embodiments of this application can quickly locate the fault point by comparing multiple dump files triggered during the operation of the basic task and the fault task. The first dump file and the second dump file save the complete context (such as memory and stack) when the fault occurs, which is convenient for reproducing and diagnosing the problem. By actively running the fault task and collecting the dump files, the system's ability to handle anomalies can be verified.

[0048] In embodiments of this application, determining the running order of each task to be processed based on the dependencies between multiple tasks to be processed may include: performing compatibility detection on at least one faulty task and at least one basic task when the running order of at least one faulty task and at least one basic task is parallel; and updating the running order between at least one faulty task and at least one basic task in response to incompatibility between at least one faulty task and at least one basic task.

[0049] For example, each task to be processed has three states: waiting, running, and completed.

[0050] For example, non-dependencies are also included between multiple pending tasks.

[0051] Non-dependency means that at least two tasks to be processed will not trigger each other or be in a waiting state during operation, do not share runtime necessary resources, and do not transfer data between tasks.

[0052] For example, non-dependency relationships include: multiple fault tasks that can run in parallel, where a single fault does not affect other parallel fault tasks, the execution order of each fault task can be arbitrarily changed, and there is no resource sharing between fault tasks; multiple basic tasks that can run in parallel (and can be stacked), where each basic task handles different business dimensions, can be linearly stacked at runtime, and there are no state or resource conflicts; multiple non-parallel fault tasks, where each fault task has a specific execution order, and the failure of a preceding fault task does not affect the scheduling of subsequent fault tasks; and multiple non-parallel basic tasks, where the system processes only one basic task of the same type at a time, and each basic task is exclusive during runtime.

[0053] For example, it can represent multiple fault tasks in parallel without shared resource conflicts or state dependencies, and can simulate signal anomalies from multiple sensors simultaneously, with the execution order not affecting the final state.

[0054] Compatibility checks are performed between the basic tasks and faulty tasks that have dependencies to determine whether the basic tasks and faulty tasks cause race conditions. If a race condition is caused, it indicates that the basic tasks and faulty tasks are incompatible.

[0055] For example, compatibility detection includes identifying competition for shared resources (such as locks, memory regions) between basic tasks and faulty tasks that meet dependency conditions.

[0056] When a faulty task and a basic task are incompatible, the execution order between the faulty task and the basic task is changed from parallel to non-parallel, and the execution order of all pending tasks with dependencies is updated.

[0057] For example, the execution order includes: prioritizing the execution of basic tasks that depend on faulty tasks, which can avoid inconsistent states of basic tasks due to unmet preconditions; and secondly, running multiple faulty tasks that can be run in parallel, which can shorten the test cycle.

[0058] The abnormal dump file detection method based on the embodiments of this application can maximize resource utilization and reduce the completion time of all pending tasks by running independent pending tasks in parallel.

[0059] In the embodiments of this application, the abnormal dump file detection method may further include: after performing compatibility testing, determining the parallelizable task slots, wherein the parallelizable task slots are the maximum number of at least two basic tasks or at least two faulty tasks that can be parallelized during operation.

[0060] For example, after performing compatibility testing, the remaining parallelizable task slots are calculated. The remaining parallelizable task slots represent the maximum number of tasks that can be run in parallel. If the number of faulty tasks or basic tasks that can be run in parallel exceeds the maximum number of tasks that can be run in parallel, a basic task or faulty task with a number less than or equal to the maximum number of tasks is selected from among the multiple parallelizable basic tasks or multiple parallelizable faulty tasks to run.

[0061] For example, the system continues to run until the number of waiting or running basic tasks and faulty tasks are both 0. It then iterates through the running basic tasks or faulty tasks, removes tasks with a completed status from the running queue, increases the remaining parallel task slots, schedules the basic tasks or faulty tasks that need to run according to the running order, removes completed tasks and releases resources in real time through dynamic queue management, and maximizes resource utilization based on slots and task types.

[0062] The abnormal dump file detection method based on the embodiments of this application can ensure that there is no circular dependency between the basic task and the faulty task by the dependency and non-dependency relationship of multiple pending tasks. That is, there is no basic task depending on the faulty task and the faulty task depending on the basic task, thus avoiding deadlock. Furthermore, by running the non-dependent pending tasks in parallel, the resource utilization can be maximized and the completion time of all pending tasks can be reduced.

[0063] Based on the dependencies between multiple pending tasks, the execution order of basic business and fault tasks is determined. This execution order includes both parallel and non-parallel execution. Furthermore, through... Figures 3-4 The process of determining the abnormal dump file based on the different execution order of basic business and fault tasks in the disclosed embodiments is described in detail.

[0064] In embodiments of this application, obtaining multiple first dump files includes: when the execution order of at least one faulty task and at least one basic task is not parallel, obtaining multiple dump files obtained after the execution of each basic task in parallel execution order, as multiple first dump files.

[0065] Figure 3 The schematic diagram illustrates a test flowchart in which the basic business and fault tasks according to an embodiment of this application are non-parallel.

[0066] like Figure 3 As shown, in some embodiments, when the execution order of at least one faulty task and at least one basic task is not parallel, the acquisition of multiple first dump files in the above operation S220 may further include steps S321 to S327.

[0067] In step S321, initialize all test environments.

[0068] For example, clear all historical dump files to avoid old data interfering with the test, preset the number of loops for fault tasks, control the upper limit of test iterations to prevent infinite loops, and perform tasks that must be executed sequentially, such as initializing hardware and loading core drivers.

[0069] For example, the testing environment includes, but is not limited to, verifying the driver or firmware in abnormal conditions (such as register errors, interrupt conflicts), detecting problems such as task scheduling, memory management or priority inversion, verifying the fault tolerance of the electronic control unit in the event of sensor failure or communication interruption, detecting the impact of host failure on the virtual machine state, detecting memory leaks or race conditions in kernel modules (such as device drivers), and discovering potential vulnerabilities through fault injection.

[0070] In step S322, the non-parallel basic task begins to run.

[0071] In step S323, after the non-parallel basic tasks have been completed, the parallel basic tasks begin to run.

[0072] In this embodiment, "parallelizable" is also known as "overlayable." By running parallelizable basic tasks, the concurrent execution of multiple tasks in real-world scenarios is simulated, ensuring that the test covers high-load or race conditions.

[0073] In step S324, all dump files obtained after the completion of the overlayable basic tasks are used as the first dump files.

[0074] In step S325, after all the basic tasks that can be run in parallel have been completed, the faulty task begins to run.

[0075] When basic tasks and faulty tasks are not parallel, basic tasks usually have higher business priority, while faulty tasks, as safeguard tasks, have system priority. Therefore, it is necessary to run all basic tasks that are not parallel to this faulty task first, and then run this faulty task to eliminate the uncertainty caused by the competition between faulty tasks and basic tasks. Furthermore, errors in basic tasks will not affect the operation of faulty tasks.

[0076] In step S326, after the fault task in step S324 is completed, all the dump files obtained after the completion of the task are acquired as all the second dump files.

[0077] Compare all the second dump files with all the first dump files. If no dump file that is different from any of the first dump files appears in any of the second dump files, repeat steps S324 to S326, update the fault task in step S324, until the preset number of iterations is reached. If a dump file that is different from any of the first dump files appears in any of the second dump files, the different dump file is regarded as an abnormal dump file.

[0078] In step S327, after an abnormal dump file is generated, the running process ends.

[0079] The abnormal dump file detection method based on the embodiments of this application can accurately identify system state changes caused by faulty tasks by comparing the first dump file and the second dump file, avoid misjudgment, and automatically repeat the process of obtaining the first dump file and the second dump file by pre-setting the number of loops, thereby reducing manual intervention and improving testing efficiency.

[0080] In embodiments of this application, obtaining multiple first dump files includes: when the execution order of at least one faulty task and at least one basic task is parallel, obtaining multiple dump files obtained after the execution of each basic task in a non-parallel order, as multiple first dump files.

[0081] Figure 4 The diagram illustrates a test flowchart in which basic business and fault tasks are performed in parallel according to an embodiment of this application.

[0082] like Figure 4 As shown, in some embodiments, when at least one faulty task and at least one basic task are executed in parallel, the acquisition of multiple first dump files in the above operation S220 may further include steps S421 to S427.

[0083] In step S421, initialize all test environments.

[0084] For example, clear all historical dump files to avoid old data interfering with the test, preset the number of loops for fault tasks, control the upper limit of test iterations to prevent infinite loops, and perform tasks that must be executed sequentially, such as initializing hardware and loading core drivers.

[0085] In step S422, the non-parallel basic task is started to run. At the same time, in step S424, multiple dump files generated during the running of the non-parallel basic task are obtained as the first dump files for all tasks.

[0086] These independent basic tasks depend on the faulty tasks, so multiple first dump files need to be obtained when the basic tasks start running to avoid interference from multiple second dump files generated by the faulty tasks running at the same time.

[0087] In step S423, after all non-parallel basic tasks have been run, parallel basic tasks are started.

[0088] After step S425 and step S424, the fault task begins to run.

[0089] In step S426, after the fault task is completed, all dump files obtained after the fault task is completed are acquired as all the second dump files.

[0090] Compare all the second dump files with all the first dump files. If no dump file that is different from any of the first dump files appears in any of the second dump files, repeat steps S424 to S426, update the fault task in step S424, until the preset number of iterations is reached. If a dump file that is different from any of the first dump files appears in any of the second dump files, the different dump file is regarded as an abnormal dump file.

[0091] In step S427, after an abnormal dump file is generated, the running process ends.

[0092] For example, by using structured exception handling or signal processing mechanisms, multiple exception dump files generated by errors occurring during severe operation can be captured. These exception dump files include complete diagnostic information when the error is discovered. By terminating the operation, the abnormal state can be prevented from spreading to other operating systems when an exception occurs, and the abnormal operation process can be prevented from continuing to write erroneous data. Timely termination can prevent the continuous growth of leaked memory.

[0093] The abnormal dump file detection method based on the embodiments of this application selects different test processes by running multiple pending tasks in different order, obtains multiple first dump files, and compares them with multiple second dump files. This method can accurately identify system state changes caused by faulty tasks and avoid deadlock when multiple running processes compete for resources because basic tasks and faulty tasks are waiting for each other to release resources.

[0094] In the embodiments of this application, when at least one faulty task generates multiple dump files and the number N of at least one faulty task is greater than or equal to 3, the fault start time and fault end time of N faulty tasks are obtained; the difference between the fault start time of the Nth faulty task and the fault end time of the second faulty task is determined; and the fault start time of the Nth faulty task is updated based on the difference.

[0095] For example, situations that may generate dump files include code defects, cold restarts during operation, and warm restarts. Therefore, when running a faulty task that may generate a dump file, it is necessary to record the start time and end time of the faulty operation. The start time and end time of each faulty task are stored in a time list according to the running order.

[0096] For example, when the number of faulty tasks in the time list is greater than or equal to 3, before running the next faulty task, record the current fault start time, compare the fault start time with the fault end time of the second faulty task in the list, and if the difference between the two times is greater than 45 minutes, the faulty task can be run and the current fault start time is recorded in the time list; if the difference is less than or equal to 45 minutes, the faulty task that is about to be run will wait until the difference is greater than 45 minutes, record the current fault start time in the time list, and then run the current faulty task.

[0097] The abnormal dump file detection method based on the embodiments of this application records the fault start time and fault end time of the faulty task, which facilitates the reproduction and diagnosis of problems existing in the system when abnormal dump files occur, and can verify the system's ability to handle anomalies.

[0098] For example, abnormal dump file detection is divided into two cases. The first case is when the running faulty task itself does not generate a dump file. If a dump file is generated during the running process, all dump files in the path where the dump file was generated before the faulty task was run need to be recorded as all first dump files. After the faulty task is completed, all dump files generated after the task is completed are recorded as all second dump files.

[0099] The second scenario is when the faulty task itself generates a dump file. In this case, it is necessary to determine whether an abnormal dump file was generated in addition to the normal dump file. Before running these faulty tasks, mark them as dump files, and use all marked dump files as the first dump files. Then, record the dump file marks that the faulty task has after it finishes running, and use all the dump file marks as the second dump files.

[0100] In embodiments of this application, when at least one faulty task generates multiple dump files, the multiple second dump files are compared with the multiple first dump files, and a target second dump file that is different from the multiple first dump files is determined from the multiple second dump files; if the number of target second dump files is greater than or equal to 2, the target second dump file is determined to be an abnormal dump file.

[0101] In cases where at least one faulty task generates multiple dump files, the faulty task is marked with dump files before execution. All marked dump files are considered as first dump files. The dump file markings after the faulty task finishes execution are recorded, and all those marked are considered as second dump files. From the multiple second dump files, it is determined whether there are two or more target second dump files that are different from the multiple first dump files. If there are two or more target second dump files, it can be considered that abnormal dump files have been generated. The fault area where the abnormality occurred during the execution of the faulty task can be located through these abnormal dump files.

[0102] The abnormal dump file detection method based on the embodiments of this application can quickly locate the variable, memory region or code segment that causes the abnormality by comparing the key differences between the first dump file and the second dump file, exclude irrelevant dump files, and improve the accuracy of fault diagnosis.

[0103] In embodiments of this application, when multiple dump files are generated in the path of running at least one faulty task, the multiple second dump files are compared with the multiple first dump files, and a target second dump file that is different from the multiple first dump files is determined from the multiple second dump files; if the number of target second dump files is greater than or equal to 1, the target second dump file is determined to be an abnormal dump file.

[0104] If multiple dump files are generated in the path where at least one faulty task is running, record all dump files in the path where the dump files were generated before the faulty task was running, as all first dump files. After the faulty task is completed, record all dump files generated after the task is completed, as all second dump files.

[0105] The abnormal dump file detection method based on the embodiments of this application compares the key differences between the first dump file and the second dump file, and combines the fault type of the fault task to determine the abnormal dump file that best matches the characteristics of the fault task from multiple second dump files. This can quickly locate the variable, memory region or code segment that causes the abnormality, exclude irrelevant dump files, and improve the accuracy of fault diagnosis.

[0106] For example, dump file detection can include: in a computer blue screen crash, analyzing memory dump files to locate driver or hardware failures; in an operating system kernel crash, using the operating system's kernel crash dump mechanism to generate a kernel crash dump file, and combining it with interactive debugging tools to check for memory leaks or null pointer accesses; in a graphics processor driver crash, analyzing error report dump files output by enabling the debug or verification layer to locate shader compilation errors; and in simulated sensor failures, dumping intermediate results of the sensing algorithm to verify the failure mode.

[0107] The above testing process automatically identifies abnormal dump files, reduces the workload of manually checking multiple dump files, collects dump files for parallel basic tasks, covers high-concurrency scenarios, and avoids complex issues such as missing race conditions.

[0108] Based on the above-mentioned abnormal dump file detection method, this application also provides an abnormal dump file detection device. The following will be combined with... Figure 5 The device is described in detail.

[0109] Figure 5 A schematic block diagram of an abnormal dump file detection device according to an embodiment of this application is shown.

[0110] like Figure 5 As shown, the abnormal dump file detection device 500 of this embodiment includes a task sorting module 510, a dump acquisition module 520, and an abnormal detection module 530.

[0111] The task sorting module 510 is used to determine the execution order of each task to be processed based on the dependencies between multiple tasks to be processed. The multiple tasks to be processed include at least one basic task and at least one faulty task. In one embodiment, the task sorting module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0112] The dump acquisition module 520 is used to acquire multiple first dump files and multiple second dump files. The first dump files are multiple dump files generated during the execution of at least one basic task in the execution order, and the second dump files are multiple dump files generated during the execution of at least one faulty task in the execution order. In one embodiment, the dump acquisition module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0113] The anomaly detection module 530 is used to compare multiple second dump files and multiple first dump files according to the fault type of at least one faulty task, and determine the abnormal dump file from the multiple second dump files. In one embodiment, the anomaly detection module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0114] According to embodiments of this application, any multiple modules of the task sequencing module 510, dump acquisition module 520, and anomaly detection module 530 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the task sequencing module 510, dump acquisition module 520, and anomaly detection module 530 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the task sequencing module 510, dump acquisition module 520, and anomaly detection module 530 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0115] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing an abnormal dump file detection method according to an embodiment of this application.

[0116] like Figure 6 As shown, an electronic device 600 according to an embodiment of this application includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0117] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0118] According to embodiments of this application, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.

[0119] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0120] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.

[0121] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0123] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0124] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A method for detecting abnormal dump files, characterized in that, The method includes: Based on the dependencies between multiple tasks to be processed, the execution order of each task to be processed is determined. The multiple tasks to be processed include at least one basic task and at least one fault task. The execution order includes: first running the basic task that depends on the fault task, and then running the fault task that can be run in parallel. Obtain multiple first dump files and multiple second dump files, wherein the first dump files are dump files generated during the execution of the at least one basic task in the order of execution, and the second dump files are dump files generated during the execution of the at least one faulty task in the order of execution. Based on the fault type of the at least one faulty task, multiple second dump files and multiple first dump files are compared to determine an abnormal dump file from the multiple second dump files. The step of comparing multiple second dump files and multiple first dump files based on the fault type of the at least one faulty task to determine an abnormal dump file from the multiple second dump files includes: In the case where at least one faulty task itself generates multiple dump files The plurality of second dump files are compared with the plurality of first dump files, and a target second dump file that is different from the plurality of first dump files is determined from the plurality of second dump files; If the number of the target second dump files is greater than or equal to 2, the target second dump files are identified as the abnormal dump files; In the case where multiple dump files are generated in the path where the at least one faulty task is running, The plurality of second dump files are compared with the plurality of first dump files, and a target second dump file that is different from the plurality of first dump files is determined from the plurality of second dump files; If the number of the target second dump files is greater than or equal to 1, the target second dump file is identified as the abnormal dump file.

2. The method according to claim 1, characterized in that, The step of determining the execution order of each task based on the dependencies between multiple tasks includes: When the execution order of the at least one faulty task and the at least one basic task is parallel, Perform compatibility testing between the at least one faulty task and the at least one basic task; In response to incompatibility between the at least one faulty task and the at least one basic task, the execution order between the at least one faulty task and the at least one basic task is updated.

3. The method according to claim 1, characterized in that, The acquisition of multiple first dump files includes: When the execution order of the at least one faulty task and the at least one basic task is not parallel, Multiple dump files obtained after the basic tasks, which are executed in parallel order, are used as the multiple first dump files.

4. The method according to claim 2, characterized in that, The process of obtaining multiple first dump files also includes: When the execution order of the at least one faulty task and the at least one basic task is parallel, Multiple dump files obtained after the basic tasks in non-parallel execution order are completed are used as the multiple first dump files.

5. The method according to claim 1, characterized in that, The method further includes: When at least one faulty task generates multiple dump files and the number N of the at least one faulty task is greater than or equal to 3. Obtain the fault start time and fault end time for N faulty tasks; Determine the difference between the fault start time of the Nth faulty task and the fault end time of the second faulty task; Based on the difference, update the fault start time of the Nth faulty task.

6. An abnormal dump file detection device, characterized in that, The device includes: The task sorting module is used to determine the running order of each task to be processed based on the dependency relationship between multiple tasks to be processed. The multiple tasks to be processed include at least one basic task and at least one fault task. The running order includes: first running the basic task that depends on the fault task, and then running the fault task that can be run in parallel. The dump acquisition module is used to acquire multiple first dump files and multiple second dump files. The first dump files are multiple dump files generated during the execution of the at least one basic task in the order of execution. The second dump files are multiple dump files generated during the execution of the at least one faulty task in the order of execution. An anomaly detection module is used to compare multiple second dump files and multiple first dump files according to the fault type of the at least one faulty task, and determine an abnormal dump file from the multiple second dump files. The step of comparing multiple second dump files and multiple first dump files according to the fault type of the at least one faulty task and determining an abnormal dump file from the multiple second dump files includes: when the at least one faulty task itself generates multiple dump files, The plurality of second dump files are compared with the plurality of first dump files, and a target second dump file that is different from the plurality of first dump files is determined from the plurality of second dump files; If the number of the target second dump files is greater than or equal to 2, the target second dump files are identified as the abnormal dump files; In the case where multiple dump files are generated in the path where the at least one faulty task is running, The plurality of second dump files are compared with the plurality of first dump files, and a target second dump file that is different from the plurality of first dump files is determined from the plurality of second dump files; If the number of the target second dump files is greater than or equal to 1, the target second dump file is identified as the abnormal dump file.

7. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Fault alarm method and device, computer equipment and storage medium

    CN119814529A