Programmatic performance anomaly detection

By receiving and analyzing the speed data of the workload manager, comparing the current speed with the expected speed factors, and generating remediation actions, the problem of difficulty in collecting evidence in the existing technology is solved, and more accurate and efficient performance abnormality detection is achieved.

CN114253751BActive Publication Date: 2025-06-06INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111120560.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-25
Filing Date
2021-09-24
Publication Date
2025-06-06
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reduce false positive information in performance abnormality detection, and lacks the ability to collect supporting evidence to operate in abnormal mode, which affects the accurate identification of performance abnormalities.

Method used

By periodically receiving velocity data in the address space from the workload manager, creating an expected velocity value, and comparing the current velocity value with a factor of the expected velocity value, a remediation action indicating an exception is generated based on the current velocity value below the factor.

Benefits of technology

More accurate performance abnormality detection is achieved, reducing the possibility of false positive information, and improving the efficiency of performance abnormality handling through automated remediation actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253751B_ABST
    Figure CN114253751B_ABST
Patent Text Reader

Abstract

A method, system, and computer program product for performance anomaly detection are provided. Velocity data for one or more address spaces is periodically received from a workload manager. An expected velocity value is created for each of the one or more address spaces. A factor of the expected velocity value is compared to a current velocity value from the velocity data. Based on the current velocity value being below the factor, a remedial action is generated indicating an anomaly.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Embodiments of the present invention relate generally to computer systems, and more particularly to performance anomaly detection.

[0002] Programmatic performance anomaly detection involves the analysis of system behavior to determine the range of metrics that indicate normal behavior vs. the range that indicates abnormal behavior. To reduce the likelihood of false positive information, the collection of supporting evidence of abnormal behavior helps further narrow down the associated problem symptoms. However, the identification of such evidence usually requires the system to operate in an abnormal mode to collect valuable data. Summary of the invention

[0003] Among other things, a method for performance anomaly detection is provided. Velocity data for one or more address spaces is periodically received from a workload manager. An expected velocity value is created for each of the one or more address spaces. A factor of the expected velocity value is compared to a current velocity value from the velocity data. Based on the current velocity value being below the factor, a remedial action is generated indicating an anomaly.

[0004] Embodiments are further directed to computer systems and computer program products having substantially the same features as the above-described computer-implemented methods.

[0005] Additional features and advantages are achieved through the techniques described herein. Other embodiments and aspects are described in detail herein. For a better understanding, please refer to the specification and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other features and advantages will be apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0007] Figure 1 is a functional block diagram of an exemplary system according to an embodiment of the present invention;

[0008] Figure 2 depicts a predictive failure analysis system according to an embodiment of the present invention;

[0009] Figure 3 Describes the workflow of the predictive failure analysis system; and

[0010] Figure 4 is a schematic functional block diagram of a computing device for implementing various aspects of the present invention according to an embodiment of the present invention. DETAILED DESCRIPTION

[0011] The present disclosure relates generally to the field of programmatic performance anomaly detection. Program and system anomaly detection analyzes normal program and system behavior and discovers anomalous execution caused by attacks, misconfigurations, program errors, and unusual usage patterns.

[0012] Anomaly detection involves identifying unexpected items or events in a dataset that differ from the norm. Anomaly detection assumes that anomalies occur rarely in the data and that the characteristics of anomalies differ significantly from normal instances.

[0013] A common approach used by IT operations staff is to assume that everything is operating well until a performance problem occurs. In current practice, several siloed management tools are used that monitor system behavior and provide drill-down mechanisms to determine potential symptoms. The nature and complexity of problem determination can vary based on user background and experience. For example, an experienced administrator may know to execute one tool instead of another, or to execute a specific series of commands, while a less experienced administrator may not. Operator commands can be used to look for unusual behavior. However, in very high-speed computing environments, it is advantageous to integrate performance degradation detection into the process to automatically initiate further analysis of possible potential anomalies.

[0014] The workload manager (WLM) component of the operating system currently enables system administrators to define performance goals in service classes. A service class is a named workgroup within a workload that has performance goals, resource requirements, and similar performance characteristics of business importance to the enterprise.

[0015] This includes metrics that indicate average response time, response time within percentiles, speed targets, and targets for arbitrary workloads. Speed ​​is a measure of how fast work should run when ready without being delayed by system resources. It is defined as a measure of the processor activity used to process the workload over time, along with the latency introduced to support processing the workload. Latency includes operating system processing related to the processor, storage, and I / O, including memory paging, page swapping, job creation and initialization delays, etc.

[0016] Predictive Failure Analysis (PFA) is an operating system component that collects data, models the collected data to create expected values ​​or rates, and compares current metric usage to factors of the expected values ​​or rates to determine if anomalous behavior is occurring. PFA's functionality proactively detects corruption in address space that could cause system outages.

[0017] In current practice, the output of WLM and the output of PFA are separate. PFA can collect historical data based on a single address space, a group of address spaces, or the entire system. However, PFA does not collect performance data, nor does it collect data from WLM to monitor performance.

[0018] Embodiments of the present invention combine the processing of WLM and PFA by allowing PFA to collect WLM velocity data on an address space granularity basis, model expected values ​​based on historical data, and compare current velocity to a factor of the expected value. The modeled data is used to determine whether the address space is operating normally, or is below normal behavior and thus degraded. The resulting assessment is then used to determine whether to declare that a performance anomaly is occurring on the system. The determination of a performance anomaly is used to initiate a process that can directly alert the facility's automation products and / or system administrators to immediately resolve the performance anomaly. For example, the facility's automation products can generate reports and / or problem tickets and initiate the collection of relevant diagnostic data to further determine the problem symptoms.

[0019] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings.

[0020] Figure 1 1 is a functional block diagram of a computer system 100. The computer system includes a computer system / server (server) 12 according to an embodiment of the present invention. The computer system 100 may include more than one server 12. The server 12 may include any computer capable of performing the functions of hosting and executing WLM and PFA; receiving large amounts of log and similar data (e.g., terabytes or more) from hardware, operating systems, and applications; performing statistical analysis on the log and similar data; and modeling the collected data to determine whether anomalies occur on one or more workloads.

[0021] The functions and processes of the server 12 may be described in the context of computer system executable instructions (such as program modules, routines, objects, data structures and logic, etc.) that perform specific tasks or implement specific abstract data types. The server 12 may be part of a distributed cloud computing environment in which one or more servers 12 perform tasks linked through a communication network (such as network 13).

[0022] like Figure 1 As shown, server 12 may include one or more processors or processing units 16 , a system memory 28 , and a bus 18 that couples various system components including system memory 28 to processing unit 16 .

[0023] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures.

[0024] The server 12 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by the computer system / server 12, and it includes volatile and nonvolatile media, removable and non-removable media.

[0025] Memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Server 12 may also include other removable / non-removable, volatile / non-volatile computer system storage media. For example, storage system 34 may include non-removable non-volatile magnetic media, such as a "hard drive" and an optical drive for reading from or writing to a removable non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media. Each device in storage system 34 may be connected to bus 18 via one or more data media interfaces (such as I / O interface 22).

[0026] Each program 40 represents one of a plurality of programs stored in the storage system 34 and loaded into the memory 28 for execution. The programs 40 include instances of operating systems, applications, system utilities, or the like. Each program 40 includes one or more modules 42. In the present invention, both WLM and PFA are instances of programs 40. Several configurations of WLM and PFA are possible. For example, WLM and PFA may all reside on the same server 12.

[0027] The server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, display 24, etc.; one or more devices that enable a user to interact with the server 12; and / or any device that enables the server 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communications may occur via an input / output (I / O) interface 22. The server 12 may communicate with one or more networks (such as network 13) via a network adapter 20. As shown, the network adapter 20 communicates with other components of the server 12 via a bus 18. Although not shown, other hardware and / or software components may be used in conjunction with the server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0028] Figure 2 Describes an embodiment of the present invention that can be used in Figure 1 A predictive failure analysis system (PFA system) 200 is implemented on a computer system 100.

[0029] The predictive failure analysis address space (PFA address space) 215 of the PFA system 200 receives raw performance data 250 from the WLM in real time or near real time, or in batches. The frequency of the raw performance data 250 collection can be configured. For example, by default, the collection occurs every minute, but can be configured differently. The received performance data 250 is stored in the data set 220 for further processing.

[0030] Additional configurable parameters include the minimum required number of minutes (uptime) that an address space must be active before PFA collects historical data. This avoids collecting data for transient or short-running address spaces. The default is 60 minutes. If an address space ends and is restarted, the address space is considered a new job and must meet the minimum uptime. Data from address spaces with the same name is not used to model newly activated address spaces. Multiple address spaces with the same name will be collected separately using a key of name / address space identifier / start time. Address spaces started within the first hour after the IPL of the server do not need to wait before being collected. However, the address space will need to be active for one full collection interval before being collected.

[0031] The category configurable parameter can be used to define which categories of address space will be collected. Specifying a lower category automatically includes the higher category. For example, if IMPORTANT is specified, then the CRITICAL and IMPORTANT categories are collected.

[0032] The CRITICAL address space is an address space identified with critical system work and infrastructure (e.g., system tasks). The IMPORTANT address space includes the CRITICAL address space plus critical middleware servers that are defined as very important. The NORMAL address space includes the CRITICAL and IMPORTANT address spaces plus normal work. Normal work includes non-server applications and services. By using the default of IMPORTANT, server-type address spaces will be included in the collection as long as they meet uptime requirements and are not specifically excluded from the collection through configuration parameters. Discretionary work is not an allowed category.

[0033] The PFA address space 215 may provide one or more interfaces, such as a GUI, a command line, and a parameter file, to receive management commands to perform actions on the data set 220. The action may specify which workloads, address spaces, and / or job data include or exclude from the WLM data set. Different actions may further specify which of the data sets 220 are to be included in the predictive failure analysis prediction modeling (PFA modeling) 225. Additional parameters for controlling the operation of the PFA address space 215 include parameters for stopping / starting / modifying the collection of certain categories of data, for adding / deleting workloads and address spaces for collection, and for excluding specific jobs from the collection. Additional parameters may specify the frequency of analyzing and modeling the data set 220. The data set 220 may be sorted by address space source, date, record type, or other criteria. The data set 220 is input to the PFA modeling 225 and becomes the historical data 230 to update the model. The PFA address space 215 stores the original data sets of the previous hour, twenty-four hours, and seven days as the historical data 230. These periods are configurable. The previous model may be stored in the historical data 230. PFA modeling 225 may use machine learning including custom algorithms developed by the business implementing PFA system 200. PFA modeling 225 may utilize algorithms from one or more statistical modeling software packages such as IBM Machine learning) output application programming interface (API) to create models.

[0034] Figure 3 The workflow of a PFA system 200 according to an embodiment of the present invention is depicted.

[0035] At 310, the PFA address space 215 receives address space speed data from the WLM. Speed ​​can be calculated as (use samples * 100) / (use samples + delay samples), where the use samples include all types of processors (e.g., CPU, memory, cache) and I / Os of the use samples. Delay samples include all types of processor delays, I / O delays, storage delays, and queue delays. Based on these so-called "use" and "delay" samples, the WLM address space speed is calculated, which is a measurement of how fast the work should run when ready without being delayed due to resources managed by the WLM. The speed is a percentage from "0" to "100". A low speed value indicates that the address space has very few resources it needs and is competing for resources with other address spaces. A high speed value indicates that the address space has all the resources it needs to execute. For example, "100" indicates that the sampled address space has not encountered any delays due to processor or I / O resources managed by the WLM.

[0036] At 320, the PFA address space 215 notifies the PFA modeling 225 to model the velocity data. The modeling results in the expected velocity values ​​for each address space being monitored. The velocity values ​​for each address space are calculated every 12 hours by default. The expected velocity values ​​are calculated for one hour of historical data, twenty-four hours of historical data, and seven days of historical data. These time periods may be configurable.

[0037] At 330 , the current speed is compared to a factor (ie, a percentage) of the expected speed value.

[0038] If at 340, the comparison indicates that the factor of the expected speed value is too low compared to the current speed, then at 350, the PFA address space 215 reports the exception and impact based on the WLM importance level setting. An alert is generated, which can be input to automated systems and IT personnel for generating problem tickets. The exception may also be reported to operating system components that perform runtime diagnostics. The alert may include an application identifier (such as a name or job number), a server identifier, an indicator of the nature of the problem, including any system messages. The importance level indicates how important it is for the workload to meet its performance goals. For example, after the data modeling period used to establish the normality boundary, if an address space (even one whose WLM service class meets its goals) is experiencing performance problems, it will be detected and warned before it can be noticed by an administrator.

[0039] Figure 4 Shows the application of Figure 3 The computing device 400 may include a corresponding set of internal components 800 and external components 900, which together may provide an environment for software applications. Each internal component in the set of internal components 800 includes: one or more processors 820; one or more computer-readable RAMs 822; one or more computer-readable ROMs 824 on one or more buses 826; and an execution Figure 3 One or more operating systems 828 for algorithms; and one or more computer-readable tangible storage devices 830. The one or more operating systems 828 are stored on one or more corresponding computer-readable tangible storage devices 830 for execution by one or more corresponding processors 820 via one or more corresponding RAMs 822 (which typically include cache memory). Figure 4 In the illustrated embodiment, each of the computer-readable tangible storage devices 830 is a magnetic disk storage device of an internal disk drive. Alternatively, each of the computer-readable tangible storage devices 830 is a semiconductor memory device such as ROM 824, EPROM, flash memory, or any other computer-readable tangible storage device that can store computer programs and digital information.

[0040] Each set of internal components 800 also includes a R / W drive or interface 832 to read from and write to one or more computer-readable tangible storage devices 936, such as CD-ROMs, DVDs, SSDs, USB memory sticks, and disks.

[0041] Each set of internal components 800 may also include a network adapter (or switch port card) or interface 836, such as a TCP / IP adapter card, a wireless WI-FI interface card, or a 3G or 4G wireless interface card or other wired or wireless communication link. The operating system 828 associated with the computing device 400 may be downloaded to the computing device 400 from an external computer (e.g., a server) via a network (e.g., the Internet, a local area network, or other wide area network) and a corresponding network adapter or interface 836. The network adapter (or switch port adapter) or interface 836 and the operating system 828 associated with the computing device 400 are loaded into the corresponding disk drive 830 and network adapter 836.

[0042] External components 900 may also include a touch screen 920, a keyboard 930, and a pointing device 934. Device driver 840, R / W driver or interface 832, and network adapter or interface 836 include hardware and software (stored in storage device 830 and / or ROM 824).

[0043] Different embodiments of the present invention may be implemented in a data processing system suitable for storing and / or executing program code, the data processing system including at least one processor coupled directly or indirectly to a memory element via a system bus. The memory element includes, for example, a local memory employed during actual execution of the program code, a block storage device, and a cache memory that provides temporary storage of at least some program code in order to reduce the number of times the code must be retrieved from the block storage device during execution.

[0044] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, DASD, tapes, CDs, DVDs, thumb drives and other storage media, etc.) may be coupled to the system either directly or through intervening I / O controllers. Network adapters may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems and Ethernet cards are just a few of the available types of network adapters.

[0045] The present invention may be a system, method and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or medium) having computer-readable program instructions thereon, and the computer-readable program instructions are used to cause a processor to perform various aspects of the present invention.

[0046] Computer readable storage medium can be a tangible device capable of retaining and storing instructions for use by instruction execution devices. Computer readable storage medium can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer readable storage medium includes the following: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device (such as punched card or a convex structure in a groove with instructions recorded thereon), and any suitable combination of the above. Computer readable storage medium as used herein should not be interpreted as transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (e.g., light pulses by fiber optic cables) or electrical signals transmitted by wires.

[0047] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0048] The computer-readable program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, configuration data of an integrated circuit, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and procedural programming languages ​​such as "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit (including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA)) may execute computer-readable program instructions to personalize the electronic circuit by utilizing the state information of the computer-readable program instructions, so as to perform various aspects of the present invention.

[0049] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each box of the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0050] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing device to produce a machine, such that instructions executed via a processor of the computer or other programmable data processing device create units for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may direct a computer, a programmable data processing device, and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0051] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operating steps are executed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0052] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to different embodiments of the present invention.To this end, each box in the flow chart or block diagram can represent a part of a module, segment or instruction, which includes one or more executable instructions for realizing the logical function of the specification.In some alternative implementations, the function marked in the box may not occur in the order marked in the figure.For example, depending on the function involved, the two boxes shown in succession can actually be completed as a step, at the same time, substantially at the same time, in a partially or completely overlapping manner, or these boxes can sometimes be executed in the opposite order.It will also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a system based on special-purpose hardware, which performs a specified function or action or performs a combination of special-purpose hardware and computer instructions based on the system based on special-purpose hardware.

[0053] Although preferred embodiments have been described and illustrated herein in detail, it will be apparent to those skilled in the relevant art that various modifications, additions, substitutions and the like are possible without departing from the spirit of the present disclosure, and therefore these are considered to be within the scope of the present disclosure as defined in the following claims.

Claims

1. A method for programmatic performance anomaly detection, include: periodically receiving velocity data for one or more address spaces from a workload manager; modeling an expected velocity value for each of the one or more address spaces based on historical data for predictive failure analysis; comparing a factor of the expected speed value to a current speed value from the speed data; and Based on the current speed value being below the factor, remedial action is taken indicating an abnormality.

2. The method according to claim 1, in, The velocity data is received in near real time, real time, or in batches.

3. The method according to claim 1, in, The current speed value is calculated as the usage samples multiplied by one hundred divided by the sum of the usage samples and the delay samples.

4. The method according to claim 1, in, Modeling the expected speed value further includes: inputting the received speed data and historical data into a statistical modeling software package; and The expected speed value is output.

5. The method according to claim 3, in, The usage samples include all types of processor usage, and wherein the latency samples include all types of processor latency, I / O latency, storage latency, and queue latency.

6. The method of claim 1, wherein the remedial action comprises generating an alert to an automated problem reporting system, wherein the alert comprises an application identifier such as a name or job number, a server identifier, an indicator of the nature of the problem, and any system messages.

7. The method according to claim 1, in, The factor of the expected speed and the periodicity of collection of the speed data are configurable.

8. A computer program product for programmed performance anomaly detection, the computer program product comprising a storage device having program code embodied therein, the program code being executable by a processor of a computer to perform the method according to any one of claims 1 to 7.

9. A computer system for programmed performance anomaly detection, include: one or more processing units; one or more computer readable memories; as well as One or more computer-readable storage media having program instructions stored thereon, the program instructions being used to be executed by at least one of the one or more processing units via at least one of the one or more computer-readable memories, wherein the computer system is capable of executing the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for detecting abnegation service aggression

    CN101465760A

  • User energy-level anomaly detection

    IN201641013912A