Dynamic registration of software components for diagnosing root cause of fault
By dynamically registering software components in computer systems and creating dynamic registries, the problem of difficulty in collecting specific fault data when diagnosing the root cause of failure in complex systems is solved, and automated diagnostic data collection is realized, improving fault recovery efficiency.
Patent Information
- Application Number
- CN202380078018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2023-10-17
- Publication Date
- 2025-06-20
AI Technical Summary
In complex computer systems, when diagnosing the root cause of a failure, prior art has difficulty collecting the required specific first-time failure data without compromising system availability, resulting in a time-consuming and potentially unavailable for complete diagnosis.
By dynamically registering software components, create a dynamic registry associated with threads, recording the identifiers and diagnostic content indicators of the software components to automatically collect diagnostic data when a failure occurs.
It realizes automated collection of data required for fault diagnosis without affecting system availability, reducing the time and complexity of the diagnostic process and improving the efficiency of fault recovery.
Smart Images

Figure CN120188147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more particularly, to a method, apparatus, and product for dynamically registering software components for diagnosing the root cause of a fault. Background Art
[0002] The EDVAC computer system developed in 1948 is generally considered to be the beginning of the computer era. Since then, computer systems have evolved into extremely complex devices. Today's computers are much more complex than the early EDVAC systems. A computer system typically includes a combination of hardware and software components, applications, operating systems, processors, buses, memories, input / output devices, and the like. As advancements in semiconductor processing and computer architecture continue to drive improvements in computer performance, more complex computer software has also evolved to take full advantage of the high performance of the hardware, making today's computer systems much more powerful than they were a few years ago.
[0003] As computer systems and their associated computer programs become increasingly complex, diagnosing faults and other errors in computer programs has become more challenging. To understand the challenges of fault isolation and diagnosis in large, complex application environments, some background and definitions related to terms such as program, process, thread, and dump are provided below. A program is a set of instructions designed to accomplish a particular task. A process is an instance of a program that is executing, along with any resources required for the program instance to run. The operating system (OS) is responsible for managing the resource tasks required to turn a program into a running process. A program can have multiple instances, and each instance of a running program is a process. Typically, each process has a separate memory address space in which it runs independently. Generally, each process is isolated from other processes and cannot directly access data in other processes. A thread is a unit of work or control flow that runs within a process. A single-threaded process contains one thread, such that the process and the thread are combined into one, and thus only one task is executed at a time. A multi-threaded process contains multiple threads, and the process executes multiple tasks simultaneously. A dump is a snapshot of the memory associated with one or more processes, typically including program code, system-related control blocks, process-related control blocks, and other storage areas.
[0004] When software problems occur, it is usually dependent on software that collects first-failure data to collect the data required to diagnose the problem. The specific data required for a particular failure scenario can vary widely. For example, a computer system may be executing hundreds or thousands of processes and threads simultaneously. The failure is usually limited to a subset of processes and threads, while the rest of the system remains operational. In the worst-case scenario, the system or production application may just be running slowly or unresponsive. Collecting "all" data on a computer system at the time of failure (e.g., using a full memory dump) is impractical and may require system downtime. For clients running mission-critical applications with high availability requirements, a full system outage is usually not an acceptable option. Therefore, it is very important to be able to identify and limit the specific first-failure data that needs to be collected at the time of failure without affecting system availability. System data dumps are usually captured, but they do not always contain all the data required to debug and diagnose the root cause of the failure. In such cases, the client usually needs to recreate the environment of the failure to collect additional information, which can be time-consuming and may result in multiple outages. In some cases, the recreation may take months or even years to complete. In other cases, the client may not be able to recreate the failure scenario, resulting in the problem being undiagnosable. Therefore, a new solution is needed to collect the necessary documentation required to diagnose the root cause of the failure. Summary of the Invention
[0005] According to one aspect of the present invention, a method for dynamic registration of software components for diagnosing the root cause of a failure is provided. The method includes: determining the start of the execution of a thread associated with a process, creating a dynamic registry associated with the thread, and determining one or more software components associated with the execution of the thread. The method further includes creating an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component.
[0006] According to another aspect of the present invention, an apparatus for dynamic registration of software components for diagnosing the root cause of a failure is provided. The apparatus includes a computer processor and a computer memory operatively coupled to the computer processor. The computer memory stores computer program instructions that, when executed by the computer processor, cause the apparatus to perform the following steps: determining the start of the execution of a thread associated with a process; creating a dynamic registry associated with the thread; determining one or more software components associated with the execution of the thread; and creating an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component, the diagnostic content indicator indicating component-specific diagnostic information associated with the software component.
[0007] According to another aspect of the present invention, there is provided a computer program product for dynamic registration of software components for diagnosing the root cause of a fault. The computer program product is disposed on a computer-readable medium and includes computer program instructions that, when executed, cause the computer to perform the following steps: determine the start of the execution of a thread associated with a process; create a dynamic registry associated with the thread; determine one or more software components associated with the execution of the thread; and create an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component, the diagnostic content indicator indicating component-specific diagnostic information associated with the software component.
[0008] In some examples, the method further includes: determining that the execution of the thread has been completed and deleting the dynamic registry associated with the thread. In some examples, the method further includes: determining that a trigger event associated with the thread has occurred; collecting diagnostic data for at least one of the one or more software components based on the diagnostic content indicator associated with the software component; and storing dump information including the collected diagnostic data in a central repository.
[0009] In some examples, collecting diagnostic data for at least one of the one or more software components includes collecting diagnostic data for a subset of the one or more software components. In some examples, the subset is determined according to a priority scheme. In some examples, the trigger event includes a fault during the execution of the thread. In some examples, the diagnostic data further includes one or more of the following: hardware diagnostic information or console logs. In some examples, the trigger event includes an exception during the execution of the thread. In some examples, the exception is determined based on artificial intelligence analysis and prediction.
[0010] In some examples, the entry in the dynamic registry for each of the one or more software components further includes the active or inactive state of the software component. In some examples, diagnostic information is collected only for software components having an active state. In some examples, the active state of the software component is set based on a function call of the thread to the software component. In some examples, the inactive state of the software component is set based on the return of a function call of the thread to the software component.
[0011] Other objects, features, and advantages of the present invention will become apparent from the following more particular description of exemplary embodiments of the invention, which are illustrated in the accompanying drawings, in which like reference numerals generally represent like parts in the exemplary embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A block diagram of an example computing system for the dynamic registration of software components configured to diagnose the root cause of a fault.
[0013] Figure 2 A flowchart showing an example method for the dynamic registration of software components in the thread life cycle for tracking the root cause of fault diagnosis.
[0014] Figure 3 An example program for diagnosing the root cause of a fault is illustrated.
[0015] Figure 4 A specific process example of utilizing the dynamic registration of software components for the root cause of fault diagnosis is illustrated.
[0016] Figure 5 A flowchart showing another example method for the dynamic registration of root cause software components for fault diagnosis. Detailed Description
[0017] Exemplary methods, apparatuses, and products of the present invention are for the dynamic registration of software components for diagnosing the root cause of a fault. The following will be described in conjunction with the accompanying drawings, starting from Figure 1 begin. Figure 1 A network diagram of a system configured for the dynamic registration of software components for the root cause of fault diagnosis is shown. Figure 1 A block diagram of an example computing system 100 is shown, which includes at least one computer processor 110 or "CPU" and a random access memory (RAM) 120 connected to the processor 110 and to other components of the computing system 100 through a high-speed memory bus 113 and a bus adapter 112.
[0018] Stored in the RAM 120 is an operating system 122. Operating systems suitable for the dynamic registration of software components for diagnosing the root cause of faults in embodiments of the present invention include UNIX TM 、Linux TM 、Microsoft Windows TM 、AIX TM 、etc., as well as other operating systems that will be associated with those skilled in the art. Figure 1 The operating system 122 in is shown in the RAM 120, but many components of such software are usually also stored in non-volatile memory, such as, for example, on a data repository 132, such as a disk drive. Also stored in the RAM 120 are a dynamic diagnostic data manager component 124 and a dump collector component 126, which will be further described below.
[0019] Figure 1The computing system 100 therein includes a disk drive adapter 130 and other components of the computing system 100 connected to the processor 110 via an expansion bus 117 and a bus adapter 112. The disk drive adapter 130 connects non-volatile data storage in the form of a data repository 132 to the computing system 100. Disk drive adapters configured to insert serial numbers into editable tables useful in computers according to embodiments of the present invention include integrated drive electronics (IDE) adapters, small computer system interface (SCSI) adapters, etc., and other adapters that will come to the mind of those skilled in the art. Non-volatile computer memory can also be implemented in the form of optical disk drives, electrically erasable programmable read-only memories (or EEPROMs or "flash memories"), RAM drives, etc., and other memories that will come to the mind of those skilled in the art. The dynamic diagnostic data manager component 124 is configured to create, store, and maintain a dynamic data registry 138 in the RAM 120 for each thread. The dynamic data registry 138 contains entries for each component indicating which diagnostic data is to be collected for that thread. One or more dynamic data registries will be used by the dump collector component 126 for root cause failure diagnosis, as further described herein.
[0020] Figure 1 The exemplary computing system 100 therein includes one or more input / output (I / O) adapters 116. The I / O adapters implement user-oriented input / output through software drivers and computer hardware for controlling output to a display device (such as a computer display screen) and receiving user input from a user input device 118 (such as a keyboard and a mouse). Figure 1 The exemplary computing system 100 therein includes a video adapter 134, which is a type of I / O adapter specifically designed for graphic output to a display device 136 (such as a display screen or a computer monitor). The video adapter 134 is connected to the processor 110 via a high-speed video bus 115, a bus adapter 112, and a front-side bus 111 (which is also a high-speed bus).
[0021] Figure 1The exemplary computing system 100 therein includes a communication adapter 114 for data communication with other computers and for data communication with a data communication network. Such data communication can be carried out serially via an RS-232 connection, via an external bus (such as the Universal Serial Bus USB), via an IP data communication network, etc., and other ways will be associated by those skilled in the art. The communication adapter implements data communication at the hardware level, by which a computer can send data communication to another computer directly or via a data communication network. Examples of communication adapters useful in a computer configured to insert a serial number into an editable table according to an embodiment of the present invention include a modem for wired dial-up communication, an Ethernet (IEEE 802.3) adapter for wired data communication, and an 802.11 adapter for wireless data communication, etc.
[0022] Figure 1 The communication adapter therein is communicatively coupled to a wide area network (WAN) 140, which also includes other computing devices, such as Figure 1 the computing devices 141 and 142 shown therein. In a particular embodiment, the computing system 100 includes a server, and the computing devices 141 and 142 are client devices of the server.
[0023] Figure 1 The exemplary system of the arrangement of the server and other devices shown therein is for illustration and not limitation. According to various embodiments of the present invention, a data processing system may include additional servers, routers, other devices, and a peer-to-peer architecture, as will be associated by those skilled in the art. The networks in these data processing systems can support many data communication protocols, including but not limited to the Transmission Control Protocol (TCP), Internet Protocol (IP), Hypertext Transfer Protocol (HTTP), Wireless Access Protocol (WAP), Handheld Device Transfer Protocol (HDTP), etc. Various embodiments of the present invention can be implemented on a variety of hardware platforms other than Figure 1 those shown.
[0024] Existing solutions for the root cause of fault diagnosis require a large number of manual processes and have limited automatic collection of complete diagnostic data in complex problem debugging. This is usually an error-prone and time-consuming process, in which diagnostic data is collected and analyzed relative to the code flow. Usually, subject matter experts need to be consulted. If the source of the problem is not clear, additional problem determination operations need to be performed at the client side, including collecting additional diagnostic data and re-creating the problem. This process usually needs to be repeated until the problem is solved. In addition, collecting an independent dump of all information requires restarting the system, which is usually not an option for high-availability clients and is a time-consuming process.
[0025] System dumps are typically collected by recovery routines, service level indication processing (SLIP) traps, and manually through the console. However, these existing dump collection methods pose challenges. Recovery dumps may not contain all the necessary data because the recovery routine has limited visibility into component interactions. The recovery routine typically does not know all the other components in the process and the interactions between components. Additionally, in complex scenarios involving multiple components, critical information may be missing from the recovery dump, resulting in the inability to resolve the problem. The recovery routine may be able to identify the data that its own components need to dump, but it does not always identify the data that other components unknown to the recovery routine in the failure flow need to dump. SLIP dumps may not contain all the necessary data because of limited user knowledge. SLIP dumps are driven by a set of keywords specified by humans (usually complex) that specify the data content, which is typically based on an incomplete understanding of the problem process and interactions. In complex scenarios involving multiple components, critical information may be missing from the SLIP dump. Therefore, an additional recreation of the failure is necessary to understand the root cause of the failure. For ongoing process-related conditions such as loops, high CPU usage, or hang conditions, the operator can manually collect console dumps. Similar to SLIP dumps, console dumps are operator-specified and are typically based on an incomplete understanding of the problem process and interactions.
[0026] One or more embodiments of the present invention provide for the dynamic registration of software components for the root cause of fault diagnosis. In an embodiment, an intelligent diagnostic unified management protocol (SDUMP) is provided to collect the documentation required for the problem process of one or more software components including those used to diagnose the root cause of a fault. Various embodiments provide a dynamic registration facility capable of collecting all the necessary documentation for the problem process of one or more application, middleware, and operating system components involved in diagnosing the root cause of a fault. In various embodiments, when a thread executes a workflow, the components of that workflow (e.g., application, middleware, or operating system components) are dynamically registered. Each registry entry includes an identifier of the component and a diagnostic content indicator associated with it, which describes the data to be collected if a problem occurs. In one or more embodiments, the registry entry includes an identifier of the software component, a diagnostic content indicator associated with the software component, and a status that reflects whether the component is active (registered) or inactive (unregistered) in the execution flow. In one or more embodiments, the diagnostic content indicator points to an array of component-specific regions that will provide relevant diagnostic data.
[0027] According to various embodiments, the Dynamic Diagnostic Data Manager component 124 maintains a Dynamic Data Registry 138 for tracking all software components in the problem flow. Initially, when a thread is created, the software component responsible for thread creation anchors a dynamic registry for the specific thread via the Dynamic Diagnostic Data Manager component 124. The dynamic registry for a thread is used to track all software components in the execution flow with a separate entry for each component in the execution flow. During normal processing, when a software component is added to the execution flow, the entry for the corresponding component in the thread's registry is marked as active. Once a software component is no longer active in the execution flow, the entry associated with that software component is marked as inactive in its registry. When the thread completes, the registry for that thread is deleted.
[0028] In a particular embodiment, component registration is performed based on the execution of a system call or service (e.g., a program call or a recovery routine).
[0029] One or more embodiments provide a dynamic registry that allows real-time viewing of software components related to a thread flow at any diagnostic time of interest. In various embodiments, the dynamic registry helps in collecting First Failure Data Capture (FFDC) related to the software components involved during an exceptional event, such as a thread failure or a persistent process condition (e.g., a loop, a hang, or other anomalies), to help a debugger determine the problem without having to recreate the conditions when the exception occurs (i.e., reproduce).
[0030] According to various embodiments, the Dynamic Diagnostic Data Manager component 124 provides facilities for each active thread of a process to activate, maintain, and ultimately deactivate its own unique dynamic registry entry. Each software component encountered along the thread flow will register itself. In one or more embodiments, the dynamic registry entry also includes a diagnostic content indicator (or diagnostic level) for indicating which component-specific data is most helpful in problem determination between specific parts of the flow. In one or more embodiments, the dynamic registry entry includes an indication of whether the software component is active or inactive. When a triggering event (e.g., a failure or a failure request dump during the execution of a thread) occurs, the Dump Collector component 126 queries the dynamic registry of the thread to collect information on all active component entries and their associated diagnostic content indicators, and adds the corresponding data to the data dump. In other embodiments, the triggering event includes an exception during the execution of a thread. In a particular embodiment, the (multiple) processes identified as causing the exception are determined based on artificial intelligence (or knowledge-based) analysis and / or prediction. In various embodiments, an exception includes not only a specific hardware or software failure detected by the server hardware or software, but also, for example, an incongruous, inconsistent, strange condition, or situation that may cause abnormal behavior of a computer system, such as latency, contention, resource exhaustion, or an outright failure.
[0031] In another embodiment, diagnostic information is collected only for software components that are in an active state. In another embodiment, an active state is set for a software component based on a function call by a thread to the software component. In another embodiment, an inactive state is set for a software component based on a return from a function call by a thread to the software component. In one embodiment, when the current active component is deactivated upon a return from a function call, or otherwise, the deactivation is logged in a dynamic registry.
[0032] In an embodiment, the dynamic registry is used to collect diagnostic data in an exception scenario where the triggering event is not an explicit failure but is identified as a result of a higher-level (e.g., artificial intelligence or machine learning-based) exception prediction and / or analysis.
[0033] In one or more embodiments, software components activated by a function call or other means associated with a particular thread are automatically identified and tracked. In these embodiments, when activated, the first software component is registered in a thread-specific dynamic software registry. Each other software component accessed or activated by the first software component is registered in the thread-specific dynamic software registry when first accessed or activated. A triggering event generates and captures a snapshot of all active software components in the dynamic software registry. The snapshot is used to capture and save all dump information for all active components associated with the event trigger. In a particular embodiment, a snapshot of all active software components is saved in a central repository upon the triggering event. In other embodiments, in addition to only the dump information, additional diagnostic data is collected, such as additional hardware diagnostic information or console logs.
[0034] In a particular embodiment, a continuous process condition may require a diagnostic view of all threads in the process. In such a case, the dump collector component 126 can query the dynamic registry for each thread, as described above, and then merge all the corresponding requested data into the dump.
[0035] In addition to process or thread failures, exceptions are those situations where normal operating behavior that persists can lead to latency, contention, or resource exhaustion. In these cases, the system appears to be running but is plagued by latency or other issues that may ultimately lead to system failure. Typically, an operations or system programmer analyzes such situations. If a software component is thought to be at fault, FFDC may also be required in such cases. Before corrective action is taken, an operator may request a "console dump", and the dynamic data registry 138 may help ensure that the correct diagnostic data is automatically captured for the exception.
[0036] In one or more embodiments, the Dynamic Diagnostic Data Manager component 124 is active from system initialization until system termination. The Dynamic Diagnostic Data Manager component 124 is responsible for managing the content of the dynamic registry structure. In a particular embodiment, the Dynamic Diagnostic Data Manager component 124 provides an Application Programming Interface (API) for lifecycle management, including activation, deactivation, and dynamic modification of entries for each participant component (or product). According to various embodiments, a "participant" refers to the representation of a subsystem / system component. In a particular embodiment, each participant is provided with an index to view and modify its dynamic registry entry. At system initialization, the Dynamic Diagnostic Data Manager component 124 allocates sufficient space for each existing participant to have an entry in the dynamic registry structure of the dynamic data registry 138, and the dynamic registry 138 determines the total size of the dynamic registry structure to be added as the dynamic data registry 138 for each thread.
[0037] In a particular embodiment, the Dynamic Diagnostic Data Manager component 124 implements the following APIs and their related functions:
[0038] ● NewPlayer: Dynamically adds a new participant to the dynamic registry structure and allocates a larger dynamic registry structure for the new thread.
[0039] ● Activate: Enables the dumping of its diagnostic data when the participant participates and includes the diagnostic content indicator of the participant.
[0040] ● Deactivate: Disables its entry when the participant detaches from the thread flow.
[0041] ● ModifyIndicator: Updates the diagnostic content indicator of the participant.
[0042] ● QueryThreadDiagnosticData: Called by the Dump Collector component 126 to query what data is to be dumped from each participating participant based on the diagnostic content indicator under the target thread.
[0043] In one or more embodiments, the Dynamic Diagnostic Data Manager component 124 implements the following rules to manage the behavior of the interaction of the dynamic registry:
[0044] LIFECYCLE RULES: Used during the lifecycle of a thread
[0045] 1. When a thread starts execution, a dynamic registry structure is created and anchored to the thread.
[0046] 2. When a participant (e.g., a component) registers (e.g., issues an Active command), the dynamic registry structure entry is marked as active.
[0047] 3. When a participant (e.g., a component) deregisters (e.g., issues a Deactivate), the dynamic registry structure entry is marked as inactive and is efficiently managed for performance reasons.
[0048] 4. When a thread terminates, the dynamic registry structure is deleted.
[0049] REGISTRATION RULES: Used during participant registration 1. The called participant will self-register as active in the following cases:
[0050] a) Executes work on a separate thread on behalf of the current thread.
[0051] b) Begins to execute code in another memory space (e.g., after a PC routine in z / OS).
[0052] c) Requests an internal serialization resource (e.g., a lock, latch, or ENQ in z / OS).
[0053] d) Sets a failure return code to pass back to its caller and marks itself as permanent (e.g., for the lifetime of a thread). This is to prevent cascading conditions due to error return codes such as hangs, loops, sick but not dead (SBND), or exception codes, which would ultimately result in a request dump.
[0054] e) Establishes a recovery routine (e.g., ESTAE or FRR in z / OS, or CDT in AIX).
[0055] 2. During subsystem creation, the creator of the subsystem (e.g., z / OS UNIX, Db2, IMS, CICS) self-registers.
[0056] DEREGISTRATION RULES: Make itself inactive during participant deregistration
[0057] 1. During subsystem deletion:
[0058] a) When disconnecting from the primary subsystem, the subsystem realizes the completed / terminated connection and deregisters itself.
[0059] REGISTRY MODIFICATION RULES: Used during registry modification
[0060] 1. Update the diagnostic content indicator when a participant wishes to add additional diagnostic data or remove existing diagnostic data to be included in the dump.
[0061] a) Add data when relevant in the workflow (e.g., when working in z / OS with different address spaces and associated data spaces).
[0062] b) Remove data from the workflow that is no longer relevant.
[0063] In a particular embodiment, the participant is responsible for defining its list of diagnostic content indicators and the associated diagnostic data to be collected.
[0064] Example embodiment of a Dynamic Registry Structure (DynRegStructure):
[0065] DynRegStructure
[0066] 1Header information
[0067] 2ThreadPtr / * Thread pointer * /
[0068] 2ProcessPtr / * Process pointer * /
[0069] 1DynRegPlayer(nnn) / * Array of entries for each component / subsystem * /
[0070] 2DynRegEntryId / * Unique identifier for the participant * /
[0071] 2DynRegEntryFlagActive bit(1) / * The participant is involved and needs to be dumped * /
[0072] 2DynRegEntryIndicatorValue bit(4) / * Diagnostic content indicator (string / integer) * /
[0073] 2DynRegEntryIndicatorPtr / * Ptr to the list of the participant's data to be collected at dump time * /
[0074] Figure 2A flowchart of an example method 200 is shown for dynamically registering software components for diagnosing the root cause of a fault during the lifecycle of a thread. The method 200 includes the start of a thread (202), where the component responsible for thread creation anchors (204) a dynamic data registry 138 for the thread in the RAM 120. In a particular embodiment, the entries in the dynamic data registry 138 include an identifier of the software component, a diagnostic content indicator associated with the software component, and a status reflecting whether the component is active (registered) or inactive (unregistered) in the execution flow. The thread then continues normal processing (206).
[0075] During normal processing of the thread (206), new components may be added to or removed from the execution flow (208). When a component is executing, the component decides whether to activate or deactivate (210) its component entry in the thread's dynamic registry. If the thread is not complete, the method continues normal processing of the thread (206). If the thread completes (212), the component responsible for thread termination deletes (214) the dynamic data registry 138 associated with the thread. The thread then ends (216).
[0076] Figure 3 An example program 300 for diagnosing the root cause of a fault in accordance with an embodiment of the present invention is illustrated. Process 301 includes a plurality of active threads (Thread 1 to Thread n) 302. Each of the active threads 302 in the active threads creates, anchors, maintains, and ultimately deletes its own unique dynamic registry 304. In a particular embodiment, program 300 is facilitated by a dynamic diagnostic data manager component 124, which provides an interface between a particular thread 302 and the thread's dynamic registry 304. Each of the software components encountered along the flow of a thread 302 activates or deactivates its entry and adjusts the associated diagnostic content indicator in its particular dynamic registry 304. The diagnostic content indicator points to an array of component-specific data that will be most helpful in problem determination at a particular part of the flow. Thus, each dynamic registry 304 may include registry entries corresponding to a plurality of software components (e.g., Component 1 to Component n) encountered in the active thread. When registering during the dynamic registry 304, a software component may change its diagnostic content indicator to reflect a changed diagnostic level associated with the software component, or deactivate the entry when the component is no longer relevant to the flow.
[0077] At the time point 306 of a dump of a thread in the request thread 302, for example, requested by a user or automatically, the dump collector component 126 queries the dynamic registry 304 of the thread to read and / or archive (308) the registry to collect information about the component entries registered in the dynamic registry 304 of the thread and their associated diagnostic content indicators, and adds the corresponding data to the dump. The dump collector component adds diagnostic data for the components listed from the (multiple) dynamic registries 304 to the dump data (310). In a particular embodiment, the dump collector component 126 archives a copy of the dynamic registry 304 to a repository for subsequent inspection. Continuous process conditions may require a diagnostic view of all threads within the process: in such a case, the dump collector component 126 queries the dynamic registry 304 of each thread and merges the corresponding requested data into the dump.
[0078] Figure 4 An example of a specific process 400 for using dynamically registered software components for root cause fault diagnosis according to an embodiment of the present invention is shown. In a first example scenario, using the dynamic registry as described above, the running thread 402 attempts to perform an operation on an encrypted file system. When the encryption software component is waiting for a response from the encryption hardware, an asynchronous timeout fault occurs on the thread, which triggers a request for a dump (436). The dump collector component 126 queries the dynamic registry 410 of the faulty thread to collect the data indicated by the diagnostic level (or diagnostic content indicator) of each registered component in the dynamic registry 410, and adds this data to the dump.
[0079] When a timeout occurs, checking the dynamic registry 410 of the thread shows that a request by the UNIX component 404 to read a file causes the UNIX component 404 to self-register (406) in the first entry 408 of the dynamic registry 410, with a diagnostic level (or diagnostic content indicator) of 1, indicating to dump all storage associated with UNIX-specific data. Next, the UNIX component 404 calls the file system component 412, which self-registers (414) in the entry 416 in the dynamic registry 410, with a diagnostic content indicator of 1, indicating to dump all data associated with all file system-specific data. The file system component 412 initiates an I / O operation as part of the read request. The I / O component 418 self-registers (420) in the entry 422 in the dynamic registry 410, with a diagnostic content indicator of 1. The I / O component 418 notices that the file system of the target data is encrypted and calls the encryption component 424 for decryption. The encryption component 424 self-registers (426) in the entry 428 in the dynamic registry 410, with a diagnostic content indicator of 1, indicating to dump the storage associated with the requested operation. When the encryption component 424 is about to send a decryption request to the encryption hardware component 430 to access the encryption key, the encryption component 424 modifies its diagnostic content indicator to 2 to ensure that data related to the hardware request will also be dumped. Additionally, the encryption component 424 registers the encryption hardware component 430 (432) in the fifth entry 434 in the dynamic registry 410, with a diagnostic content indicator of 1, indicating to collect encryption hardware logs. The dump obtained by the dump collector component 126 contains the requested FFDC and specific data required for the recovery routine.
[0080] In a second example scenario, without using the dynamic registry 410, as described above, the thread 402 attempts to perform a read operation on an encrypted file again, and again an asynchronous timeout occurs while the encryption component 424 is waiting for a response from the encryption hardware component 430. The dump collector component 126 has no dynamic registry to query and thus relies on the recovery routine of the encryption component to identify the data to be included in the dump. The encryption component 424 will direct the dump collector component 126 to collect relevant data of the encryption component in the dump, and the dump collector component 126 may also capture thread-specific data and data of its caller (e.g., I / O component data). However, the dump collector component 126 will not know to collect any UNIX or file system data and will not be able to collect encryption hardware logs. Since the encryption hardware component caused the delay, the user or operator needs to manually request the hardware colleague to provide the hardware logs and hopes that the logs are still available. It may be necessary to recreate the program, request a dump with additional data, and manually request the hardware logs more promptly to understand the root cause of the failure.
[0081] In a third example scenario, some components participate in the dynamic registry 410, but other components do not. In such a scenario, data is collected based on a combination of the diagnostic levels of the participating components and the data requested by the recovery routines.
[0082] For further illustration, Figure 5 FIG. 500 is a flow chart of an example method 500 in accordance with an embodiment of the present disclosure for dynamic registration of software components for root cause of fault diagnosis. The method 500 includes: determining the start of execution of a thread associated with a process (502); creating a dynamic registry associated with the thread (504); determining one or more software components associated with the execution of the thread (506); and creating an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component, the diagnostic content indicator describing component-specific diagnostic information associated with the software component (508).
[0083] In another embodiment, the method 500 further includes: determining that the thread execution has completed and deleting the dynamic registry associated with the thread. In another embodiment, the method further includes: determining that a trigger event associated with the thread has occurred; collecting diagnostic data for at least one of the one or more software components based on the diagnostic content indicator associated with the software component; and storing dump information including the collected diagnostic data in a central repository.
[0084] In another embodiment, collecting diagnostic data for at least one of the one or more software components includes collecting diagnostic data for a subset of the one or more software components. In another embodiment, the subset is determined according to a priority scheme. For example, in an embodiment, the priority scheme gives preference to diagnostic data that is considered more important than other diagnostic data.
[0085] In another embodiment, the trigger event includes a fault during the execution of the thread. In another embodiment, the diagnostic data further includes one or more of the following: hardware diagnostic information or console logs. In another embodiment, the trigger event includes an exception during the execution of the thread. In another embodiment, the exception is determined based on artificial intelligence analysis and prediction. In yet another embodiment, the exception is determined based on machine learning.
[0086] In another embodiment, the entry in the dynamic registry for each of one or more software components further includes the active or inactive state of the software component. In another embodiment, diagnostic information is collected only for software components with an active state. In another embodiment, the active state of a software component is set based on a function call to the software component by a thread. In another embodiment, the inactive state of a software component is set based on the return of a function call to the software component by a thread.
[0087] In view of the foregoing explanation, the reader will recognize that the advantages of the dynamic registration of software components for diagnosing the root cause of a failure according to an embodiment of the present invention include:
[0088] ● Improve first failure data capture (FFDC) so that the system supports collecting more information without returning to the client or customer for program recreation, thereby preventing customer downtime. Thus, the number of recreations required is reduced and serviceability is improved.
[0089] ● Improve FFDC by providing a complete profile of the software components and their interactions to determine the diagnostic data required.
[0090] ● Automate the collection of diagnostic data required for multi-component problems, eliminating the need for subject matter experts to analyze and determine what data needs to be collected to solve the problem.
[0091] ● Reduce the debugging time of the test and development cycles.
[0092] ● Greatly improve the integrity of the diagnostic data collected for failure recovery processing.
[0093] ● Eliminate the need for the client to decide which subsystems / components need to be dumped for debugging operations. Thus, the speed and accuracy of debugging operations are improved.
[0094] Exemplary embodiments of the present invention have been mainly described in the context of a fully functional computer system for the dynamic registration of software components for the root cause of fault diagnosis. However, those skilled in the art will recognize that the present invention can also be embodied in a computer program product placed on a computer-readable storage medium for use with any suitable data processing system. Such a computer-readable storage medium can be any storage medium for machine-readable information, including magnetic media, optical media, or other suitable media. Examples of such media include magnetic disks or floppy disks in a hard disk drive, optical discs for an optical disc drive, magnetic tapes, and other media that will occur to those skilled in the art. Those skilled in the art will immediately recognize that any computer system with appropriate programming means will be able to execute the steps of the method of the present invention as embodied in a computer program product. Those skilled in the art will also recognize that although some exemplary embodiments described in this specification are directed to software installed and executed on computer hardware, alternative embodiments implemented as firmware or hardware are also within the scope of the present invention.
[0095] The present invention can be a system, a method, and / or a computer program product. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0096] A computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of a computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0097] The computer-readable program instructions described herein can be downloaded to a corresponding computing / processing device from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0098] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, in order to carry out aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit.
[0099] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0100] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions that implement aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0101] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0102] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0103] It will be understood from the foregoing description that various modifications and changes may be made to the various embodiments of the present invention without departing from the true spirit of the invention. The description in this specification is for illustrative purposes only and should not be construed as restrictive. The scope of the invention is defined only by the language of the appended claims.
Claims
1. A method for dynamic registration of software components for diagnosing the root cause of a fault, the method comprising: Determine the start of the execution of a thread associated with a process; Create a dynamic registry associated with the thread; Determine one or more software components associated with the execution of the thread; And Create an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component, the diagnostic content indicator indicating component-specific diagnostic information associated with the software component.
2. The method according to claim 1, further comprising: Determine that the execution of the thread has been completed; And Delete the dynamic registry associated with the thread.
3. The method according to claim 1 or claim 2, the method further comprising: Determine that a trigger event associated with the thread has occurred; Collect diagnostic data for at least one of the one or more software components based on the diagnostic content indicator associated with the software component; And Store dump information including the collected diagnostic data in a central repository.
4. The method according to claim 3, wherein collecting diagnostic data for at least one of the one or more software components includes collecting diagnostic data for a subset of the one or more software components.
5. The method according to claim 4, wherein the subset is determined based on a priority scheme.
6. The method according to any one of claims 3 to 5, wherein the trigger event includes a fault during the execution of the thread.
7. The method according to any one of claims 3 to 5, wherein the diagnostic data further includes one or more of the following: hardware diagnostic information or console logs.
8. The method according to any one of claims 3 to 5, wherein the trigger event includes an exception during the execution of the thread.
9. The method according to claim 8, wherein the software component associated with the exception is determined based on artificial intelligence analysis and prediction.
10. The method according to any of the preceding claims, wherein the entry in the dynamic registry for each of the one or more software components further includes the active or inactive state of the software component.
11. The method according to claim 10, wherein the diagnostic information is collected only for the software components including the active state.
12. The method according to claim 10 or claim 11, wherein the active state is set for the software component based on a function call of the thread to the software component.
13. The method according to any one of claims 10 to 12, wherein an inactive state is set for a software component based on a return of a function call of the software component by the thread.
14. An apparatus for dynamic registration of software components for diagnosing a root cause of a fault, the apparatus comprising a computer processor and a computer memory operatively coupled to the computer processor, the computer memory having computer program instructions therein, which when executed by the computer processor, cause the apparatus to perform the following steps: Determine a start of execution of a thread associated with a process; Create a dynamic registry associated with the thread; Determine one or more software components associated with the execution of the thread; and Create an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component, the diagnostic content indicator indicating component-specific diagnostic information associated with the software component.
15. The apparatus according to claim 14, wherein the computer instructions further cause the apparatus to perform the following steps: Determine that the execution of the thread has been completed; and Delete the dynamic registry associated with the thread.
16. The apparatus according to claim 14 or claim 15, wherein the computer instructions further cause the apparatus to perform the following steps: Determine that a trigger event associated with the thread has occurred; Collect diagnostic data for at least one of the one or more software components based on the diagnostic content indicator associated with the software component; and Store dump information including the collected diagnostic data in a central repository.
17. The apparatus according to claim 16, wherein collecting diagnostic data for at least one of the one or more software components includes collecting diagnostic data for a subset of the one or more software components.
18. A computer program product for dynamic registration of software components for diagnosing a root cause of a fault, the computer program product being disposed on a computer-readable medium, the computer program product including computer program instructions which, when executed, cause a computer to perform the following steps: Determine a start of execution of a thread associated with a process; Create a dynamic registry associated with the thread; Identify one or more software components associated with the execution of the thread; and Create an entry in the dynamic registry for each of the one or more software components, the entry including an identifier of the software component and a diagnostic content indicator associated with the software component, the diagnostic content indicator indicating component-specific diagnostic information associated with the software component.
19. The computer program product according to claim 18, wherein the computer-readable medium comprises a signal medium.
20. The computer program product according to claim 18, wherein the computer-readable medium comprises a storage medium.
Citation Information
Cited By
Dynamic registration of software components for diagnosis of root cause of failure
US12602276B2