System and method for automatically calculating recovery metrics with real-time telemetry and proposing recovery plans

The smart simulation recovery module uses machine learning to predict and automate recovery processes, addressing the challenge of estimating recovery times and data loss from attacks, enhancing business resilience through accurate and timely restoration.

JP7738678B2Active Publication Date: 2025-09-12JPMORGAN CHASE BANK NA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023571577
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-18
Filing Date
2022-04-28
Publication Date
2025-09-12
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

Existing systems lack the ability to predict and automate recovery times and data loss in the event of a destructive malware attack or system failure, impacting business resilience and recovery strategies.

Method used

A smart simulation recovery module using machine learning to estimate recovery time and data loss, incorporating real-time telemetry and historical data to identify risks and initiate autonomous recovery processes.

Benefits of technology

Enables accurate prediction of recovery times and data loss, allowing for proactive resilience enhancement and automated recovery strategies, ensuring timely restoration of critical business processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007738678000001
    Figure 0007738678000001
  • Figure 0007738678000002
    Figure 0007738678000002
  • Figure 0007738678000003
    Figure 0007738678000003
Patent Text Reader

Abstract

Various methods, apparatus / systems and media are provided for understanding the recovery of business services due to a loss of availability occurring in an information technology infrastructure. The system and method of the present invention automatically predicts or detects the probability of an availability incident and improves the determination of the severity of the incident based on technology component attribute data, incident history data or other metadata by using machine learning models to determine the associated risk and impact, and the machine learning models go into an alert state to determine the capacity requirements / availability of alternative affected infrastructure and initiate the orchestration of recovery, total recovery time, and potential data loss.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 189,881, filed May 18, 2021, the entire contents of which are incorporated herein by reference. [Technical Field]

[0002] The present disclosure relates generally to information technology (IT), and more particularly to a method and apparatus for implementing or realizing a smart simulation recovery module that automatically estimates the total recovery time and amount of data loss that may occur to recover an application, infrastructure, or business process in the event of a destructive malware attack or attack, or in the event of a system maintenance delay or unexpected system failure, when an application or group of applications supporting a business process must be rebuilt and restored. [Background technology]

[0003] The developments described in this section are known to the inventors. However, unless otherwise indicated, any development described in this section should not be deemed to constitute prior art or to be known to those skilled in the art merely by virtue of its inclusion in this section.

[0004] As today's IT environment rapidly grows, so too does the threat and regulatory landscape. Organizationally critical entities may need to understand the time and cost impact of restoring critical business services and processes, including IT. Business processes and services typically involve intertwined technology, people, and processes across business units, support teams, and management. Ensuring that organizationally critical entities can predict the time it will take to restore a business service or process or other organizationally significant unit of measure to a reliable state and compare it against a recovery time objective (RTO) and / or industry-standard recovery metrics is essential to ensure that technology recovery issues can be remediated before recovery times exceed the RTO and become incidents. Summary of the Invention [Problem to be solved by the invention]

[0005] Business organizations within organizationally critical institutions may view technology recovery through various attributes, such as business product recovery, technology product recovery, business process recovery, and technology component recovery. Using telemetry data and metadata provided by organizational systems of record, it may be possible to predict recovery times, such as analyzed return to business operations, recovery point objectives, MTBF (mean time before failure), MTTR (mean time to recover, repair, respond, or resolve), MTTA (mean time to confirm), and MTTF (mean time to failure), to build new resilience capabilities and take action to enhance the resilience of organizationally critical institutions. Therefore, a need exists for systems and methods for predicting technology recovery times based on business requirements, determining potential risks and impacts, and enhancing business resilience, and proposing or implementing automated recovery strategies. [Means for solving the problem]

[0006] The present disclosure, through one or more aspects and / or embodiments and / or specific configurations or subcomponents, provides, among other things, but is not limited to, various systems, servers, devices, methods, media, programs, and platforms for implementing a smart simulation recovery module that automatically estimates the total recovery time and amount of data loss required to recover an application, infrastructure, or business process that may occur when an application or group of applications supporting a business process, including IT, must be rebuilt and restored due to a destructive malware attack or the event of such an attack. In an exemplary embodiment, the amount of data loss may be an index that requires calculation. One example of the purpose of the present disclosure is to determine the impact on an application, infrastructure, or service.

[0007] For example, the present disclosure, through one or more various aspects and / or embodiments and / or specific configurations or subcomponents, also provides various systems, servers, devices, methods, media, programs, and platforms for implementing a smart simulation recovery module that implements an active machine learning module configured to receive and maintain current and historical application data, asset data, data recovery data, business process data, technical product data, application transaction data, incident and change management data, and threat intelligence data to calculate recoverability / recovery metrics for critical business processes and services. The active machine learning module may be configured to monitor critical business processes and services to identify potential operational risks where the expected recovery time does not match the actual recovery time due to regulations, business goals, technical designs, or suboptimal business processes. The active machine learning module may be further configured, but is not limited to, to identify the business processes or services and calculate the potential operational risk for the entire company or organizational units when a potential availability incident is detected, and to launch an orchestration engine to prepare for recovery.

[0008] The present disclosure, through one or more various aspects and / or embodiments and / or specific configurations or subcomponents, also provides various systems, servers, devices, methods, media, programs, and platforms for implementing a smart simulation recovery module that can be configured to provide automatic adjustments based on, among other things, critical telemetry data, such as backup information including last known good state (trusted state), changes in technology capabilities, and the time period of an event. Additionally, the smart simulation recovery module can be further configured to automatically schedule recovery tests to validate simulation results based on key criteria. Similarly, the smart simulation recovery module can be further configured to allow users to select various time periods for disruptive events and to adjust the resiliency of an application based on future solutions. In this way, the smart simulation recovery module can provide a platform for application and / or business operations teams to review the potential impact of migrating to new technologies.

[0009] In one aspect of the present disclosure, a method for automatically predicting and analyzing the recovery time to a reliable state of an application, infrastructure, or business process using real-time telemetry using one or more processors and one or more memories is disclosed. The method may include: establishing a communication link between a plurality of data sources and an event bus; implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data related to an application or applications supporting an infrastructure or business process from the plurality of data sources via the event bus; using the machine learning model to automatically predict a probability of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determining, based on the probability data, an associated risk and impact of a potential amount of data loss in the event of a destructive attack, requiring the application or applications supporting the infrastructure or business process to be rebuilt and / or recovered; and dynamically providing an estimate of a total recovery time and / or a total rebuild time for the application or applications based on determining the associated risk and impact of the amount of data loss.

[0010] In other aspects of the present disclosure, the plurality of real-time telemetry data may include, but is not limited to, application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence data for calculating resilience and / or recovery metrics for critical infrastructure or business processes and services.

[0011] In a further aspect of the present disclosure, the method may further comprise monitoring the critical infrastructure or the business processes and services to identify potential business risks where expected recovery times do not match actual recovery times due to regulations, business objectives, technology designs, or suboptimal business processes.

[0012] In yet another aspect of the present disclosure, the method may further include using the machine learning model to automatically detect the availability incident based on the received plurality of real-time telemetry data and the historical event data, and performing automatic functionality repair and / or functionality modification based on the inferred data in response to the destructive attack event.

[0013] In a further aspect of the present disclosure, the method may further comprise initiating a fully autonomous recovery process for the affected environment in response to the destructive attack event.

[0014] In other aspects of the present disclosure, the event bus may be configured to analyze and execute remediation activities in real time by providing one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources, although the present disclosure is not limited thereto.

[0015] In yet another aspect of the present disclosure, the method may further include, upon receiving the plurality of real-time telemetry data, dynamically tracking information asset data within the organization; dynamically tracking the application or applications within the organization; and dynamically tracking resilience and recovery attribute data of technology products used in the organization.

[0016] In a further aspect of the present disclosure, the method may further comprise recording pending infrastructure or business-related changes, affected assets, implementation timeframes, and an assessment of human-induced risks and impacts when determining the associated risk and impact of the amount of data loss.

[0017] In other aspects of the present disclosure, the machine learning model may be configured to go into an alert state, determine the capacity requirements and / or availability of alternative affected infrastructure, and initiate orchestration of recovery, the total recovery time, and the amount of data loss, although the present disclosure is not limited thereto.

[0018] In yet another aspect of the present disclosure, the method may further include automatically providing simulation results of recovery and automatically scheduling recovery tests to verify the simulation results.

[0019] In a further aspect of the present disclosure, the machine learning model may be further configured to identify business processes or services to calculate potential business risks to the entire organization or organizational units when a potential availability incident is detected, and to trigger an orchestration engine to perform recovery.

[0020] In another aspect of the present disclosure, the method may further comprise providing a centralized repository of past recovery incident data as an input for the probabilistic data used in predictive recovery analysis, the centralized repository mapping application inventory to business operations modules to determine upstream and / or downstream availability impacts to improve decision making, the business operations modules configured to identify business processes, services, products, and their importance to the organization's business operations.

[0021] In yet another aspect of the present disclosure, the method may further comprise performing one or more of a recovery process to rebuild the application or group of applications, a recovery process to a last known trusted state of the application or group of applications, and a simulation / test of the recovery process.

[0022] In one aspect of the present disclosure, a system for automatically predicting and analyzing the recovery time to a reliable state of an application, infrastructure, or business process using real-time telemetry is disclosed. The system may include a processor and a memory operatively connected to the processor via a communications interface and storing computer-readable instructions that, when executed, cause the processor to perform the following steps: establish a communication link between a plurality of data sources and an event bus; implement a machine learning model configured to receive a plurality of real-time telemetry data and historical event data related to an application or applications supporting an infrastructure or business process from the plurality of data sources via the event bus; use the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determine, based on the probability data, an associated risk and impact of a potential amount of data loss if the application or applications supporting the infrastructure or business process must be rebuilt and / or recovered in the event of a destructive attack; and dynamically provide an estimate of a total recovery time and / or a total rebuild time for the application or applications based on determining the associated risk and impact of the amount of data loss.

[0023] In a further aspect of the present disclosure, the processor is further configured to monitor the critical infrastructure or the business processes and services to identify potential business risks where expected recovery times do not match actual recovery times due to regulations, business objectives, technology designs, or suboptimal business processes.

[0024] In yet another aspect of the present disclosure, the processor is further configured to automatically detect the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model, and to perform automatic functionality repair and / or functionality modification based on the inferred data in response to the destructive attack event.

[0025] In a further aspect of the present disclosure, the processor is further configured to initiate a fully autonomous remediation process for the affected environment in response to the destructive attack event.

[0026] In yet another aspect of the present disclosure, the processor, upon receiving the plurality of real-time telemetry data, is further configured to dynamically track information asset data within an organization, dynamically track the application or applications within the organization, and dynamically track resilience and recovery attribute data of technology products used in the organization.

[0027] In a further aspect of the present disclosure, the processor, in determining the associated risk and the impact of the amount of data loss, is further configured to record pending infrastructure or business-related changes, affected assets, implementation timeframes, and an assessment of human-induced risks and impacts.

[0028] In yet another aspect of the present disclosure, the processor is further configured to automatically generate simulation results of recovery and automatically schedule recovery tests to verify the simulation results.

[0029] In another aspect of the present disclosure, the processor is further configured to provide a centralized repository of past recovery incident data as an input for the probabilistic data used in predictive recovery analysis, the centralized repository mapping application inventory to business operations modules to determine upstream and / or downstream availability impacts to improve decision making, the business operations modules configured to identify business processes, services, products, and their importance to the organization's business operations.

[0030] In yet another aspect of the present disclosure, the processor is further configured to execute one or more of a recovery process to rebuild the application or the group of applications, a recovery process to a last known trusted state of the application or the group of applications, and a simulation / test of a recovery process.

[0031] In one aspect of the present disclosure, a non-transitory computer-readable medium configured to store instructions for automatically predictively analyzing the recovery time to a reliable state of an application, infrastructure, or business process using real-time telemetry may be disclosed, which, when executed, may cause a processor to: establish a communication link between a plurality of data sources and an event bus; implement a machine learning model configured to receive, from the plurality of data sources via the event bus, a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure or business process; use the machine learning model to automatically predict a probability of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determine, based on the probability data, an associated risk and impact of a potential amount of data loss in the event of a destructive attack, in which the application or applications supporting the infrastructure or business process must be rebuilt and / or recovered; and dynamically provide an estimate of a total recovery time and / or a total rebuild time for the application or applications based on determining the associated risk and impact of the amount of data loss.

[0032] In a further aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform a procedure to monitor the critical infrastructure or the business processes and services to identify potential business risks where expected recovery times do not match actual recovery times due to regulations, business objectives, technology designs, or suboptimal business processes.

[0033] In yet another aspect of the present disclosure, the instructions, when executed, may further cause the processor to automatically detect the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model, and to perform automatic functionality repair and / or functionality modification based on the inferred data in response to the destructive attack event.

[0034] In a further aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform a procedure to initiate a fully autonomous recovery process for the affected environment in response to the destructive attack event.

[0035] In yet another aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform, upon receiving the plurality of real-time telemetry data, dynamically tracking information asset data within an organization, dynamically tracking the application or applications within the organization, and dynamically tracking resilience and recovery attribute data of technology products used in the organization.

[0036] In a further aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform a procedure to record pending infrastructure or business-related changes, affected assets, implementation timeframes, and an assessment of human-induced risks and impacts when determining the associated risk and impact of the amount of data loss.

[0037] In yet another aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform steps of automatically providing simulation results of recovery and automatically scheduling recovery tests to verify the simulation results.

[0038] In another aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform steps of providing a centralized repository of past recovery incident data as an input for the probabilistic data used in predictive recovery analysis, the centralized repository mapping application inventory to business operations modules to determine upstream and / or downstream availability impacts for improved decision-making, the business operations modules configured to identify business processes, services, products, and their importance to an organization's business operations.

[0039] In yet another aspect of the present disclosure, the instructions, when executed, may further cause the processor to perform one or more of a recovery process to rebuild the application or the group of applications, a recovery process to a last known trusted state of the application or the group of applications, and a simulation / test of a recovery process.

[0040] The present disclosure will be further explained by way of non-limiting examples of preferred embodiments of the present disclosure in the following detailed description, with reference to the following figures, in which like reference numerals refer to like structures / components throughout the various views: [Brief explanation of the drawings]

[0041] [Figure 1] FIG. 2 illustrates components of a computer system that implements a smart simulation recovery module in an exemplary embodiment. [Figure 2] 1 illustrates an example of a network environment including a smart simulation recovery device in an exemplary embodiment. [Figure 3] FIG. 1 is a system diagram illustrating an architecture for implementing a smart simulation recovery device with a smart simulation recovery module in an exemplary embodiment. [Figure 4]FIG. 4 is a system diagram illustrating the major components that implement the smart simulation recovery module of FIG. 3 in an exemplary embodiment. [Figure 5] FIG. 5 illustrates an example architecture for automated predictive analysis of information technology (IT) process recovery through real-time telemetry, as implemented by the smart simulation recovery module of FIG. 4 in an exemplary embodiment. [Figure 6] FIG. 1 is an end-to-end flow diagram illustrating the implementation of a smart simulation recovery module in an exemplary embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0042] One or more of the various aspects and / or embodiments and / or particular configurations or subcomponents of the present disclosure are intended to provide one or more of the advantages detailed above and / or below.

[0043] The examples may also be embodied as one or more non-transitory computer-readable media having stored thereon instructions for one or more aspects of the technology illustrated and described in the examples herein, in some examples including executable code that, when executed by one or more processors, causes the processors to perform the steps necessary to implement the method of each example of the technology illustrated and described herein.

[0044] The illustrated embodiments are shown and described in terms of functional blocks and / or units and / or modules, as is customary in the art. Those skilled in the art will appreciate that such blocks and / or units and / or modules may be physically realized using electronic (or optical) circuits, such as logic circuits, discrete components, microprocessors, hardwired circuits, memory elements, and wiring connections, which may be formed using semiconductor-based or other fabrication technologies. When blocks and / or units and / or modules are realized using a microprocessor or the like, the microprocessor or the like may be programmed with software (e.g., microcode) to perform the functions described herein and, in some cases, may be driven by firmware and / or software. Alternatively, each block and / or unit and / or module may be realized as dedicated hardware, or a combination of dedicated hardware for some functions and a processor (e.g., one or more programmed microprocessors and associated circuitry) for other functions. Furthermore, each block and / or unit and / or module of each illustrated embodiment may be physically divided into two or more interacting separate blocks and / or units and / or modules without departing from the scope of the concept of the present invention. Furthermore, the blocks and / or units and / or modules of each illustrated embodiment may be physically combined into a more complex block and / or unit and / or module without departing from the scope of the present disclosure.

[0045] 1 illustrates an example system of hardware and software components with embedded firmware that may be used to implement a smart simulation recovery module for automated, predictive analysis of information technology (IT) process recovery to a reliable state using real-time telemetry, in accordance with embodiments described herein. The figure generally depicts system 100, which may include a diagrammatically illustrated computer system 102.

[0046] Computer system 102 may include instructions that are executable to cause computer system 102, alone or in conjunction with other described devices, to perform one or more of the methods or computer-based functions disclosed herein. Computer system 102 may operate as a standalone device or may be connected to other systems or peripheral devices. For example, computer system 102 may include or be included within one or more of any computer, server, system, communication network, or cloud environment. Moreover, the instructions may operate in such a cloud-based computing environment.

[0047] When deployed in a network, computer system 102 may function as a server or client user in a server-client user network environment, as a client user in a cloud computing environment, or as a peer computer system in a peer-to-peer (or distributed) network environment. Computer system 102, or portions thereof, may be embodied as or incorporated into a variety of devices, such as a personal computer, tablet computer, set-top box, personal digital assistant, mobile device, palmtop computer, laptop computer, stationary computer, communication device, wireless smartphone, personal trusted device, wearable device, global positioning satellite (GPS) device, web appliance, or any machine capable of executing instructions (either sequential or otherwise) that define the operations to be performed by the computer system. Also, while only one computer system 102 is shown, alternative embodiments may include any collection of systems or subsystems that individually or collectively execute instructions or perform functions. Throughout this disclosure, the term "system" should be interpreted to encompass any collection of systems or subsystems that individually or collectively execute one or more instructions to perform one or more computer functions.

[0048] As shown in FIG. 1, the computer system 102 may include one or more processors 104. The processors 104 are tangible and non-transient. As used herein, the term "non-transient" should be interpreted as referring to a state that persists for a certain period of time, rather than a state that is permanent. Specifically, the term "non-transient" negates a temporary nature, such as a particular carrier wave or signal, which exists only transiently at a given time and place. The processor 104 is a component of an article of manufacture and / or machine. The processor 104 is configured to execute software instructions to perform functions described in the embodiments herein. The processor 104 may be a general-purpose processor or part of an application-specific integrated circuit (ASIC). The processor 104 may also be a microprocessor, microcomputer, processor chip, controller, microcontroller, digital signal processor (DSP), state machine, or programmable logic device. Processor 104 may also be a logic circuit, such as a programmable gate array (PGA) like a field programmable gate array (FPGA), or other type of circuitry comprised of discrete gate and / or transistor logic. Processor 104 may also be a central processing unit (CPU), a graphics processing unit (GPU), or both. Any processor described herein may also comprise multiple processors, parallel processors, or both. Multiple processors may be contained within or connected to a single device, or may be contained within or connected to multiple devices.

[0049] The computer system 102 may further include computer memory 106. The computer memory 106 may communicatively comprise static memory, dynamic memory, or both. Memory, as described herein, is a tangible storage medium capable of storing data and executable instructions and is non-transient while the instructions are stored. Again, the term "non-transient" as used herein should be interpreted as referring to the nature of a state that persists for a period of time, rather than the nature of a state that is permanent. Specifically, the term "non-transient" negates the notion of a temporary nature, such as the nature of a particular carrier wave or signal, which exists only transiently at any time and in any place. The memory is a component of an article of manufacture and / or machine. Memory, as described herein, is a computer-readable medium from which data and executable instructions can be read by a computer. Memory as described herein may be any form of storage medium known in the art, such as random access memory (RAM), read-only memory (ROM), flash memory, electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, cache, removable disk, tape, compact disk read-only memory (CD-ROM), digital versatile disk (DVD), floppy disk, Blu-ray disk, etc. Memory may be volatile or non-volatile memory, and may be secure and / or encrypted, or non-secure and / or unencrypted. Of course, computer memory 106 may be any combination of memories or may consist of a single storage.

[0050] The computer system 102 may further include a display unit 108, such as any known display device, such as a liquid crystal display (LCD), an organic light emitting diode (OLED), a flat panel display, a solid state display, a cathode ray tube (CRT), or a plasma display.

[0051] The computer system 102 may further include one or more input devices 110, such as a keyboard, a touch input screen or pad, voice input, a mouse, a remote control device consisting of a wireless keypad, a microphone coupled to a voice recognition engine, a camera such as a video camera or a still camera, a cursor control device, a global positioning system (GPS) device, an altimeter, a gyroscope, an accelerometer, a proximity sensor, or any combination thereof. Those skilled in the art will appreciate that embodiments of the computer system 102 may include more than one input device 110. Furthermore, those skilled in the art will appreciate that the above examples of input devices 110 are not exhaustive and that the computer system 102 may include any additional or alternative input device 110.

[0052] Computer system 102 may further include a media reader 112 configured to read any one or more instructions (e.g., software, etc.) from any memory described herein. The instructions may be executed by a processor to implement one or more of the methods or processes described herein. In a particular embodiment, the instructions may reside, in whole or at least in part, in memory 106 and / or in media reader 112 and / or in processor 110 while executing on computer system 102.

[0053] Computer system 102 may further include any additional devices, components, parts, peripherals, hardware, software, or any combination thereof, commonly understood and known to be included with or within a computer system, such as, but not limited to, a network interface 114, output device(s) 116, etc. Output device(s) 116 may be, but is not limited to, speakers, audio output, video output, remote control output, printer, or any combination thereof.

[0054] The components of the computer system 102 may be interconnected and communicate with each other via a communication link such as a bus 118. As shown in FIG. 1, the components may be interconnected and communicate with each other via an internal bus. However, one skilled in the art will appreciate that any of the components may be connected via an expansion bus. Furthermore, the bus 118 may be capable of communication according to any commonly understood and known standard, such as, but not limited to, a peripheral component interconnect, a peripheral component interconnect express, a parallel advanced technology attachment, or a serial advanced technology attachment, or other standard.

[0055] The computer system 102 may communicate with one or more additional computer devices 120 via a network 122. The network 122 may be any network commonly understood and known in the art, such as, but not limited to, a local area network, a wide area network, the Internet, a telephone network, a Wi-Fi network, or a short-range network. The short-range network may comprise, for example, Bluetooth, Zigbee, infrared, near-field communication, ultraband, or any combination thereof. Those skilled in the art will appreciate that additional understood and known networks 122 may be used in addition or instead, and further, that the above examples of the network 122 are not limited to or exhaustive. Also, while FIG. 1 depicts the network 122 as a wireless network, those skilled in the art will appreciate that the network 122 may also be a wired network.

[0056] Further computing device 120 is depicted in FIG. 1 as a personal computer. However, those skilled in the art will appreciate that in alternative embodiments of the present application, computing device 120 may be any device capable of executing instructions (either sequential or otherwise) that define the operations it performs, such as a laptop computer, a tablet PC, a personal digital assistant, a mobile device, a palmtop computer, a desktop computer, a communications device, a wireless telephone, a personal trusted device, a web appliance, or a server. Of course, those skilled in the art will appreciate that the above-described devices are merely exemplary, and that device 120 may be any additional device or arrangement commonly understood and known in the art without departing from the scope of the present application. For example, computing device 120 may be the same as or similar to computer system 102. Those skilled in the art will also appreciate that the device may be any combination of devices.

[0057] Of course, those skilled in the art will recognize that the above-described components of computer system 102 are merely exemplary and not intended to be all-inclusive and / or comprehensive, and that the above examples of such components are also merely exemplary and not intended to be all-inclusive and / or comprehensive.

[0058] In various embodiments of the present disclosure, the methods described herein may be implemented by a hardware computer system executing a software program. Also, as a non-limiting example embodiment, implementations may include modes of operation including distributed processing, component / object distributed processing, and parallel processing capabilities. Virtual computer system processing may be constructed to implement one or more of the methods or functions described herein, and the processors described herein may be used to support a virtual processing environment.

[0059] As described herein, the embodiments provide an optimal method for implementing a smart simulation recovery module that automatically estimates the total recovery time and amount of data loss required to recover an application, infrastructure, or business process when an application or group of applications supporting a business process, including IT, must be rebuilt and restored [due to a destructive malware attack or (hereinafter the same)] in the event of such an attack, but the present disclosure is not limited thereto.

[0060] Please refer to Figure 2, which is a schematic diagram illustrating an example network environment 200 for implementing the Smart Simulation Recovery Device (SSRD) of the present disclosure.

[0061] In exemplary embodiments, the above-described problems associated with conventional methods and systems may be addressed by, but are not limited to, implementing or realizing the SSRD 202 shown in Figure 2 and a smart simulation and recovery module that automatically estimates the total recovery time and data loss required to recover an application, infrastructure, or business process in the event of a destructive malware attack that requires the rebuilding and recovery of an application or applications supporting the business process, including IT. For example, but not limited to, the above-described problems associated with conventional methods and systems may be addressed by, but are not limited to, implementing or realizing the SSRD 202 shown in Figure 2 and a smart simulation and recovery module that implements an active machine learning module configured to receive application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence and calculate resilience / recovery metrics for critical business processes and services.

[0062] The SSRD 202 may be the same as or similar to the computer system 102 described in connection with FIG.

[0063] The SSRD 202 may store one or more applications, which may consist of executable instructions that, when executed by the SSRD 202, cause the SSRD 202 to perform operations such as sending and receiving network messages, as well as other operations shown and described below with reference to the figures. The applications may be implemented as modules or components of other applications. The applications may also be implemented as operating system extensions, modules, plug-ins, etc.

[0064] Moreover, the application may operate in a cloud-based computing environment. The application may run within or as one or more virtual machines or one or more virtual servers that may be operated in the cloud-based computing environment. The application, and even the SSRD 202 itself, may be provided on one or more virtual servers operating in the cloud-based computing environment rather than being tied to one or more specific physical network computing devices. The application may also run on one or more virtual machines (VMs) running on the SSRD 202. In one or more embodiments of the present technology, the one or more virtual machines running on the SSRD 202 may be maintained or managed by a hypervisor.

[0065] 2, the network environment 200 includes an SSRD 202 connected to multiple server devices 204(1)-204(n), which host multiple data stores 206(1)-206(n), as well as multiple client devices 208(1)-208(n), via one or more communications networks 210. A communications interface (e.g., network interface 114 of computer system 102 of FIG. 1) of the SSRD 202 operatively connects to and communicates with the SSRD 202 and / or the server devices 204(1)-204(n) and / or the client devices 208(1)-208(n), all of which are interconnected via communications network 210. However, other types and / or numbers of communications networks or systems, with other types and / or numbers of connections and / or configurations between other devices and / or elements, may be used.

[0066] Communications network 210 may be the same as or similar to network 122 described in connection with Figure 1. However, SSRD 202 and / or server devices 204(1)-204(n) and / or client devices 208(1)-208(n) may be connected to one another in a different topology. Network environment 200 may also include other network devices known in the art (and thus not described herein), such as, for example, one or more routers and / or switches.

[0067] By way of example only, communication network 210 may include one or more local area networks (LANs) or one or more wide area networks (WANs) and may utilize Ethernet or the industry standard protocol TCP / IP, although other types and / or numbers of protocols and / or communication networks may also be utilized. Communication network 202 of this example may utilize any suitable interface mechanism or network communication technology, such as, for example, any suitable form of communication traffic (e.g., voice, modem, etc.), a public switched telephone network (PSTN), an Ethernet-based packet data network (PDN), or a combination thereof.

[0068] SSRD 202 may be a standalone device or may be integrated with one or more other devices or apparatuses (e.g., one or more of server apparatuses 204(1)-204(n)). In one particular example, SSRD 202 may be hosted on one of server apparatuses 204(1)-204(n), although other configurations are also possible. Additionally, one or more of SSRD 202 may reside within the same or different communications networks, such as one or more public, private, or cloud networks.

[0069] The multiple server devices 204(1)-204(n) may be the same as or similar to the computer system 102 or computer device 120 described in connection with FIG. 1 and may have any configuration or combination of configurations described in connection with that figure. For example, any of the server devices 204(1)-204(n) may have one or more processors, memory, communication interfaces, etc., which may be connected to each other by a communication link such as a bus. However, other numbers and / or types of network devices may be used. In this example, the server devices 204(1)-204(n) may process requests received from the SSRD 202 via the communication network 210, for example, in accordance with HTTP-based and / or JavaScript Object Notation (JSON) protocols. However, other protocols may be used.

[0070] Server devices 204(1)-204(n) may be hardware or software, or may form a system with multiple servers within a pool that may consist of an internal network or an external network. Server devices 204(1)-204(n) host data stores 206(1)-206(n) configured to store metadata, data quality rules, and newly generated data.

[0071] Although server devices 204(1)-204(n) are depicted as individual devices, one or more of the processes of each server device 204(1)-204(n) may be distributed to one or more separate network computing devices that collectively constitute one or more of server devices 204(1)-204(n). Furthermore, server devices 204(1)-204(n) are not limited to any particular configuration. That is, server devices 204(1)-204(n) may be comprised of multiple network computing devices operating in a master / slave approach, where any one of server devices 204(1)-204(n) operates to manage and / or coordinate the operation of the remaining network computing devices.

[0072] The server devices 204(1)-204(n) may operate as multiple network computing devices within, for example, a cluster architecture, a peer-to-peer architecture, a virtual machine, or a cloud architecture. In other words, the technology disclosed herein should not be construed as being limited to a single environment, and other configurations and architectures are also contemplated.

[0073] The plurality of client devices 208(1)-208(n) may also be the same as or similar to computer system 102 or computer device 120 described in connection with Figure 1, and may have any configuration or combination of configurations described in connection with that figure. As used herein, a "client device" refers to any computing device that interfaces with communications network 210 to obtain resources from one or more server devices 204(1)-204(n) or other client devices 208(1)-208(n).

[0074] In exemplary embodiments, the example client devices 208(1)-208(n) may comprise any type of computing device capable of supporting implementation of SSRD 202, which may be configured to implement a smart simulation recovery module that automatically estimates the total recovery time and amount of data loss required to recover an application, infrastructure, or business process in the event of a destructive malware attack that requires the rebuilding and recovery of an application or group of applications supporting a business process, including IT, but the present disclosure is not limited thereto.

[0075] Thus, client devices 208(1)-208(n) may be mobile computing devices, stationary computing devices, laptop computing devices, tablet computing devices, virtual machines (including cloud-based computers), etc., that host document collaborative software such as, for example, chat applications, email applications, voice-to-text applications, etc.

[0076] Client devices 208(1)-208(n) may execute an interface application, such as a standard web browser or a standalone client application, that may communicate with SSRD 202 over communications network 210 to provide an interface for communicating user requests. Client devices 208(1)-208(n) may further include a display device (e.g., a display screen, a touch screen, etc.) and / or an input device (e.g., a keyboard, etc.), among other components.

[0077] Although an exemplary network environment 200 is shown and described herein, including SSRD 202, server devices 204(1)-204(n), client devices 208(1)-208(n), and communication network 210, other topologies and other types and / or numbers of systems and / or devices and / or components and / or elements may be used. Those skilled in the relevant art will recognize that the exemplary systems described herein are merely exemplary, as numerous variations are possible in the specific hardware and software used to implement the exemplary systems.

[0078] One or more of the devices depicted in network environment 200, such as SSRD 202, server devices 204(1)-204(n), and client devices 208(1)-208(n), may be configured to operate as virtual instances on the same physical machine. For example, one or more of SSRD 202, server devices 204(1)-204(n), and client devices 208(1)-208(n) may operate on the same physical device rather than as separate devices communicating over communications network 210. Additionally, there may be more or fewer SSRDs 202, server devices 204(1)-204(n), or client devices 208(1)-208(n) than depicted in FIG. 2 .

[0079] Additionally, any system or device in any example may be replaced by two or more computing systems or devices, where appropriate, thereby implementing distributed processing principles and advantages, such as redundancy and replication, to enhance the robustness and performance of the device or system in each example. By way of example only, each example may be implemented by one or more computer systems extending across any suitable network using any suitable interface mechanism or traffic technology, such as any suitable form of communications traffic (e.g., voice, modem, etc.) alone, a wireless traffic network, a cellular traffic network, a packet data network (PDN), the Internet, an intranet, or any combination thereof.

[0080] FIG. 3 is a system diagram 300 for implementing an SSRD with a Smart Simulation Recovery Module (SSRM) in one exemplary embodiment.

[0081] As shown in FIG. 3 , an SSRD 302 including an SSRM 306 may be connected to a server 304 and one or more data stores 312 (i.e., multiple data sources) via a communications network 310. The SSRD 302 may also be connected to multiple client devices 308(1)-308(n) via the communications network 310, although the disclosure is not limited thereto. In exemplary embodiments, the SSRM 306 may be implemented within the client devices 308(1)-308(n), although the disclosure is not limited thereto. In exemplary embodiments, the client devices 308(1)-308(n) may implement the SSRM 306 to automatically estimate the total recovery time and amount of data loss required to recover an application, infrastructure, or business process in the event of a destructive malware attack that requires the rebuilding and recovery of an application or group of applications supporting a business process, including IT, but the disclosure is not limited thereto.

[0082] While FIG. 3 illustrates an exemplary embodiment in which SSRD 302 includes SSRM 306, SSRD 302 may include other rules, policies, modules, data stores, applications, and the like. In exemplary embodiments, data store 312 may be embedded within SSRD 302. While FIG. 3 illustrates a single data store 312, exemplary embodiments may include multiple data stores 312. Data store 312 may comprise one or more data storage devices configured to store information data corresponding to application data, asset data, data recovery data, business process data, technical product data, application transaction data, threat intelligence, historical event data, and the like, but the disclosure is not limited in this regard. For example, data store 312 may comprise one or more memories configured to store information such as rules, programs, production requirements, testing requirements, control requirements, regulatory requirements, business requirements, and other general organizational policies. In exemplary embodiments, SSRM 306 may be storage platform agnostic and configured to be deployed across multiple storage tiers with both structured and unstructured data.

[0083] In exemplary embodiments, the SSRM 306 may be configured to receive a continuous feed of data in real time from the data store 312 and the server 304 via the communications network 310 .

[0084] Additionally, in each exemplary embodiment, data store 312 may be one or more public, private, or hybrid cloud-based data stores, or a combination thereof, that support user authentication, data store security, integration with existing data stores and developments, and a store with an open API specification definition file (i.e., JSON format) corresponding to the application, although the present disclosure is not limited thereto.

[0085] In exemplary embodiments, SSRM 406 may be implemented through a user interface, such as a web user interface, a build automation tool primarily used for Java projects, a private Jenkins, etc. (although the disclosure is not limited thereto), and may be integrated with public, private, or hybrid cloud platforms and distributed file system platforms through SSRM 406 and authentication services (although the disclosure is not limited thereto).

[0086] As described below, the SSRM 306 may be configured to: establish communication links between a plurality of data sources (i.e., data store 312) and an event bus; implement a machine learning model configured to receive, from the plurality of data sources via the event bus, a plurality of real-time telemetry data and historical event data associated with an application or group of applications supporting infrastructure or a business process; use the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determine, based on the probability data, an associated risk and impact of a potential amount of data loss in the event of a destructive attack, if the application or group of applications supporting the business process must be rebuilt and / or recovered; and dynamically provide estimated data of a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss.

[0087] A plurality of client devices 308(1)-308(n) are depicted as being in communication with the SSRD 302. In this regard, the plurality of client devices 308(1)-308(n) may be "clients" of the SSRD 302 and will be described as such herein. However, it is known and understood that the plurality of client devices 308(1)-308(n) are not necessarily "clients" of the SSRD 302 or any of the entities described herein in connection with the SSRD 302. Any additional or alternative relationship, or no relationship at all, may exist between one or more of the plurality of client devices 308(1)-308(n) and the SSRD 302.

[0088] Any of the client devices in the plurality of client devices 308(1)-308(n) may be, for example, a smartphone, a personal computer, etc. Of course, the plurality of client devices 308(1)-308(n) may also be any other device described herein. In exemplary embodiments, the server 304 may be the same as or equivalent to the server device 204 shown in FIG. 2.

[0089] The method may be performed over a communications network 310, which may consist of multiple networks such as those described above. For example, in one exemplary embodiment, one or more of multiple client devices 308(1)-308(n) may communicate with the SSRD 302 via broadband or cellular communications. Of course, such embodiments are exemplary only and are not limiting or exhaustive.

[0090] FIG. 4 is a system diagram that implements the smart simulation recovery module of FIG. 3 in one exemplary embodiment.

[0091] 4, system 400 may include an SSRD 402 in which a smart simulation recovery module (SSRM) 406 may be embedded, one or more data stores (i.e., multiple data sources) 412, a server 404, client devices 408(1)-408(n), and a communication network 410. In exemplary embodiments, SSRD 402, SSRM 406, data store 412, server 404, client devices 408(1)-408(n), and communication network 410 shown in FIG. 4 may be the same as or similar to SSRD 302, SSRM 306, data store 312, server 304, client devices 308(1)-308(n), and communication network 310 shown in FIG. 3, respectively.

[0092] 4, SSRM 406 may include a communications module 414, an implementation module 416, a prediction module 418, a calculation module 420, an execution module 422, a monitoring module 424, a detection module 426, a remediation module 428, an activation module 430, a tracking module 432, a recording module 434, and a scheduling module 436. In exemplary embodiments, data store 412 may be external to SSRD 402 and may consist of various systems maintained and operated by the organization. Alternatively, in exemplary embodiments, data store 412 may be embedded within SSRD 402 and / or SSRM 406.

[0093] The method may be performed via a communications module 414 and a communications network 410. The communications network 410 may consist of multiple networks such as those described above. For example, in one exemplary embodiment, each component of the SSRM 406 may communicate with the server 404 and the data store 412 via the communications module 414 and the communications network 410. Of course, such embodiments are exemplary only and are not limiting or exhaustive.

[0094] In each example embodiment, the communications network 410 and communications module 414 may be configured to establish links between the data store 412, the client devices 408(1)-408(n), and the SSRM 406.

[0095] In exemplary embodiments, the communication module 414, the implementation module 416, the prediction module 418, the calculation module 420, the execution module 422, the monitoring module 424, the detection module 426, the repair module 428, the activation module 430, the tracking module 432, the recording module 434, and the scheduling module 436 may each be implemented by a microprocessor or the like, which may be programmed with software (e.g., microcode) to perform the various functions described herein, and in some cases may be driven by firmware and / or software. Alternatively, the communication module 414, the implementation module 416, the prediction module 418, the calculation module 420, the execution module 422, the monitoring module 424, the detection module 426, the repair module 428, the activation module 430, the tracking module 432, the recording module 434, and the scheduling module 436 may each be implemented as dedicated hardware or a combination of dedicated hardware for some functions and a processor (e.g., one or more programmed microprocessors and associated circuitry) for other functions. Additionally, in each exemplary embodiment, communication module 414, implementation module 416, prediction module 418, calculation module 420, execution module 422, monitoring module 424, detection module 426, repair module 428, activation module 430, tracking module 432, recording module 434, and scheduling module 436 may each be physically divided into two or more separate interacting blocks and / or units and / or devices and / or modules without departing from the scope of the inventive concept.

[0096] In each example embodiment, the communication module 414, implementation module 416, prediction module 418, calculation module 420, execution module 422, monitoring module 424, detection module 426, repair module 428, activation module 430, tracking module 432, recording module 434, and scheduling module 436 of SSRM 406 may each be invoked by a corresponding API, although the disclosure is not limited in this respect.

[0097] FIG. 5 illustrates an example architecture for automated predictive analysis of business process recovery using real-time telemetry, as implemented by SSRM 406 of FIG. 4 in one exemplary embodiment.

[0098] As shown in FIG. 5, an example architecture diagram 500 can include multiple data sources 512 connected to an event bus 503, and an artificial intelligence (AI) / machine learning (ML) module 502 connected to the event bus 503. Data can flow from the data sources 512 to the event bus 503. Data can flow from the event bus 503 to the AI / ML 502. Data can flow from the AI / ML 502 to an incident / event change register 504, to a trust data register 506, and to an orchestration engine 508, which can trigger a process recovery / test simulation process 510. Resulting data from the process recovery / test simulation process is fed back to the event bus 503 for consumption by the AI / ML module 502.

[0099] In example embodiments, data can flow bidirectionally between AI / ML module 502 and data lake 513. In SSRM 406, data from data lake 513 and data from event bus 503 can be used to trigger process recovery / test reporting process 514.

[0100] In example embodiments, the AI / ML module 502 may be configured to receive application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence data from corresponding data sources 512 and calculate recoverability / recovery metrics for critical business processes and services. The SSRM 406 may monitor critical business processes and services to identify potential business risks where the expected recovery time does not match the actual recovery time due to regulations, business objectives, technical design, or suboptimal business processes.

[0101] In example embodiments, when a potential availability incident is detected, AI / ML module 502 may be further configured to identify the business processes or services, calculate the potential business risk to the entire company or organizational unit, and trigger orchestration engine 508 to prepare for recovery.

[0102] In example embodiments, SSRM 406 may include an asset inventory module that may be configured to dynamically track information assets within an organization, although the disclosure is not limited in this respect.

[0103] In example embodiments, SSRM 406 may have an application inventory module that may be configured to dynamically track applications within an organization, although the disclosure is not limited in this respect.

[0104] In exemplary embodiments, SSRM 406 may include a product catalog module that may be configured to dynamically track resilience and recovery attributes of technology products used by the organization, although the disclosure is not limited in this respect.

[0105] In example embodiments, SSRM 406 may include a configuration orchestration engine 508 that may be configured to perform automated repairs and / or modifications, and may additionally initiate fully autonomous recovery of affected environments, although the disclosure is not limited in this respect.

[0106] In exemplary embodiments, SSRM 406 may have a change recording module that may be configured to record pending IT-related changes, affected assets, implementation timeframes, and even human-based risk and impact assessments, although the disclosure is not limited thereto.

[0107] In exemplary embodiments, SSRM 406 may have a business operations module that may be configured to identify business processes, services, products, and their importance to an organization's business operations, although the disclosure is not limited in this respect.

[0108] In exemplary embodiments, but not limited to, SSRM 406 may include a threat intelligence module that may be configured to identify threats the organization faces, has, or that may target or are currently targeting the organization. This information may be used to prepare for, prevent, and identify threats that attempt to misuse valuable resources.

[0109] In example embodiments, SSRM 406 may include a regulatory compliance module that may be configured to identify regulatory agencies and regulations for business data and / or application data, although the disclosure is not limited in this respect.

[0110] The event bus 503 can be configured to provide one or more communication channels between technology assets, application transactions, and threat intelligence to analyze, prepare, and execute remediation activities in real time.

[0111] In example embodiments, SSRM 406 may include a data lake module that may be configured to provide a central location (i.e., data lake 513) for historical recovery incident data as input for predictive recovery analytics. This data lake 513 may map application inventory to business operations modules and determine upstream / downstream availability impacts for improved decision-making. In example embodiments, data lake 513 may store metrics from recovery events and provide real-time recovery data to AI / ML processes, enabling improved automated AI / ML decision-making during the recovery process and enhanced learning and understanding of the outage and recovery process.

[0112] In each example embodiment, SSRM 406 may have a trusted data module that may be configured to retrieve data from a trusted data store to restore one or more applications, although the disclosure is not limited in this respect.

[0113] The process recovery / test reporting process 514 is a recovery process that rebuilds / recovers to a last known good state (trusted state) and simulates (tests) the recovery process.

[0114] 4 and 5. The communications module 414 may be configured to establish communications links between the plurality of data sources 512 and the event bus 503. The implementation module 416 may be configured to implement the machine learning models generated by the AI / ML module 502. The machine learning models may be configured to receive real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure or business process from the plurality of data sources 512 via the event bus 503.

[0115] In example embodiments, prediction module 418 may be configured to use the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data.

[0116] In example embodiments, calculation module 420 may be configured to determine, based on the probability data, the associated risk and impact of the amount of data loss that may occur in the event of a destructive attack requiring the infrastructure or the application or applications supporting the business process to be rebuilt and / or restored.

[0117] In example embodiments, execution module 422 may be configured to dynamically provide estimated data on a total recovery time and / or a total rebuild time for the application or applications based on determining the associated risk and impact of the amount of data loss.

[0118] In exemplary embodiments, the plurality of real-time telemetry data may include, but is not limited to, application data, asset data, data recovery data, business process data, technical product data, application transaction data, threat intelligence data, etc. for calculating resilience and / or recovery metrics for critical infrastructure or business processes and services.

[0119] In example embodiments, the monitoring module 424 may be configured to monitor the critical infrastructure or business processes and services to identify potential business risks where the expected recovery time does not match the actual recovery time due to regulations, business objectives, technology design, or suboptimal business processes.

[0120] In example embodiments, the detection module 426 may be configured to automatically detect the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model. The execution module 422 may be configured to perform automatic functionality repairs and / or functionality modifications based on the inferred data in response to the destructive attack event.

[0121] In example embodiments, the activation module 430 may be configured to initiate a fully autonomous remediation process for the affected environment in response to the destructive attack event.

[0122] In example embodiments, event bus 503 may be configured to analyze and execute remediation activities in real time by providing one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources.

[0123] In example embodiments, upon receiving the plurality of real-time telemetry data, tracking module 432 may be configured to dynamically track information asset data within an organization, dynamically track the application or applications within the organization, and dynamically track resilience and recovery attribute data of technology products used by the organization.

[0124] In example embodiments, in determining the associated risk and the impact of the amount of data loss, a recording module 434 may be configured to record pending IT-related changes, affected assets, implementation timeframes, and even an assessment of human-related risks and impacts.

[0125] In example embodiments, the machine learning module 502 may be configured to enter an alert state based on multiple factors, such as, but not limited to, time of day, time month, and other operating conditions, determine the capacity requirements and / or availability of alternative affected infrastructure, and initiate the orchestration of recovery, the total recovery time, and the amount of data loss.

[0126] In example embodiments, the execution module 422 may be configured to automatically provide simulation results of recovery, and the scheduling module 436 may be configured to automatically schedule recovery tests to verify the simulation results.

[0127] In example embodiments, when a potential availability incident is detected, the machine learning model 502 may be further configured to identify business processes or services to calculate potential business risks to the entire organization or organizational units and to invoke an orchestration engine to execute recovery. Also, in example embodiments, the systems and methods disclosed herein may be capable of analyzing possible changes in the regulatory environment that may force an organization to reorganize its organization or portions of its business units or processes (i.e., splitting up a bank), and identifying the set of affected technologies and / or operations.

[0128] FIG. 6 is a flow diagram illustrating the implementation of the smart simulation recovery module in one exemplary embodiment.

[0129] The method 600 may include establishing communication links between a plurality of data sources and an event bus (step S602).

[0130] Method 600 may include implementing a machine learning model (step S604) configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting a business process from the plurality of data sources via the event bus.

[0131] The method 600 may include automatically predicting probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model (step S606).

[0132] Method 600 may include determining (step S608) based on the probability data the associated risk and impact of the amount of data loss that may occur if the application or applications supporting the infrastructure or business process must be rebuilt and / or restored in the event of a destructive attack.

[0133] Method 600 may include dynamically providing estimated data for a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss (step S610).

[0134] In example embodiments, method 600 may further include monitoring the critical infrastructure or the business processes and services to identify potential business risks where expected recovery times do not match actual recovery times due to regulations, business objectives, technology designs, or suboptimal business processes.

[0135] In example embodiments, method 600 may further include using the machine learning model to automatically detect the availability incident based on the received plurality of real-time telemetry data and the historical event data, and performing automatic functionality repair and / or functionality modification based on the inferred data in response to the destructive attack event.

[0136] In example embodiments, method 600 may further comprise initiating a fully autonomous remediation process for the affected environment in response to the destructive attack event.

[0137] In example embodiments, method 600 may further include, upon receiving the plurality of real-time telemetry data, dynamically tracking information asset data within an organization; dynamically tracking the application or applications within the organization; and dynamically tracking resilience and recovery attribute data of technology products used by the organization.

[0138] In example embodiments, method 600 may further include recording pending business-related changes, affected assets, implementation timeframes, and assessments of human-related risks and impacts when determining the associated risk and impact of the amount of data loss.

[0139] In example embodiments, method 600 may further include automatically providing simulation results of recovery and automatically scheduling recovery tests to verify the simulation results.

[0140] In example embodiments, method 600 may further include, when a potential availability incident is detected, having the machine learning model: identify a business process or service and calculate a potential business risk to the entire organization or an organizational unit; and trigger an orchestration engine to execute recovery.

[0141] In example embodiments, method 600 may further include providing a centralized repository of past recovery incident data as an input for the probabilistic data used in predictive recovery analysis, the centralized repository mapping application inventory to business operations modules to determine upstream and / or downstream availability impacts to improve decision-making, the business operations modules configured to identify business processes, services, products, and their importance to the organization's business operations.

[0142] In each exemplary embodiment, method 600 may further include performing one or more of a recovery process to rebuild the application or group of applications, a recovery process to a last known good state (trusted state) of the application or group of applications, and a simulation / test of the recovery process.

[0143] In exemplary embodiments, the SSRD 402 may include a memory (such as memory 106 shown in FIG. 1 ), which may be a non-transitory computer-readable medium. The medium may be configured to store instructions for implementing the SSRM 406, which automatically performs predictive analysis of the recovery of business processes to a reliable state through real-time telemetry, as disclosed herein. The SSRD 402 may also include a media reader (such as media reader 112 shown in FIG. 1 ), which may be configured to read any one or more instructions (e.g., software, etc.) from any memory described herein. The instructions may be executed by a processor embedded in the SSRM 406 or the SSRD 402 to implement one or more methods or processes described herein. In a specific embodiment, the instructions may reside, in whole or at least in part, in the memory 106 and / or the media reader 112 and / or the processor 104 (see FIG. 1 ) while executing in the SSRD 402.

[0144] For example, when executed, the instructions may cause processor 104 to: establish a communication link between a plurality of data sources and an event bus; implement a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure or business process from the plurality of data sources via the event bus; use the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determine, based on the probability data, an associated risk and impact of a potential amount of data loss in the event of a destructive attack in which the application or applications supporting the infrastructure or business process must be rebuilt and / or restored; and dynamically provide, but the disclosure is not limited to, an estimate of a total recovery time and / or a total rebuild time for the application or applications based on determining the associated risk and impact of the amount of data loss.

[0145] In example embodiments, the instructions, when executed, may cause processor 104 to perform procedures to monitor the critical infrastructure or the business processes and services to identify potential business risks where expected recovery times do not match actual recovery times due to regulations, business objectives, technology designs, or suboptimal business processes.

[0146] In example embodiments, the instructions, when executed, may cause processor 104 to automatically detect the availability incident using the machine learning model based on the received plurality of real-time telemetry data and the historical event data, and to perform automatic functionality repair and / or functionality modification based on the inferred data in response to the destructive attack event.

[0147] In exemplary embodiments, the instructions, when executed, may cause the processor 104 to perform procedures that, in response to the destructive attack event, initiate a fully autonomous recovery process for the affected environment.

[0148] In example embodiments, the instructions, when executed, may cause processor 104 to perform, upon receiving the plurality of real-time telemetry data, steps to dynamically track information asset data within an organization, steps to dynamically track the application or applications within the organization, and steps to dynamically track resilience and recovery attribute data of technology products used by the organization.

[0149] In exemplary embodiments, the instructions, when executed, may cause the processor 104 to perform procedures to record pending IT-related changes, affected assets, implementation timeframes, and even human-induced risk and impact assessments when determining the associated risk and impact of the amount of data loss.

[0150] In exemplary embodiments, the instructions, when executed, may cause processor 104 to perform procedures for automatically providing simulation results of recovery and automatically scheduling recovery tests to verify the simulation results.

[0151] In example embodiments, the instructions, when executed, may cause the processor 104 to perform the following procedures when a potential availability incident is detected: cause the machine learning model to identify business processes or services and calculate potential business risk to the entire organization or an organizational unit; and invoke an orchestration engine to perform recovery.

[0152] In example embodiments, the instructions, when executed, may cause processor 104 to perform procedures to provide a centralized repository of past recovery incident data as input for the probabilistic data used in predictive recovery analysis, the centralized repository mapping application inventory to business operations modules to determine upstream and / or downstream availability impacts for improved decision-making, the business operations modules configured to identify business processes, services, products, and their importance to the organization's business operations.

[0153] In each exemplary embodiment, the instructions, when executed, may cause the processor 104 to perform procedures to perform one or more of a recovery process to rebuild the application or group of applications, a recovery process to a last known good state (trusted state) of the application or group of applications, and a simulation / test of a recovery process.

[0154] According to the exemplary embodiments disclosed so far in FIGS. 1 to 6, technical improvements provided by the present disclosure may include, but are not limited to, a platform that provides a smart simulation recovery module that automatically estimates the total recovery time and amount of data loss required to recover an application, infrastructure, or business process when an application or group of applications supporting a business process, including IT, must be rebuilt and restored in the event of a destructive malware attack.

[0155] While the present invention has been described with reference to certain exemplary embodiments, it is to be understood that these terms are intended to be descriptive and illustrative, rather than limiting. Changes may be made within the scope of the appended claims, as presently defined and as amended, without departing from the scope and spirit of each aspect of the present disclosure. While the present invention has been described with reference to particular means, materials, and embodiments, the present invention is not limited to the specific disclosures, but rather extends to all functionally equivalent structures, methods, and uses, such as fall within the scope of the appended claims.

[0156] For example, a computer-readable medium may be described as a single medium. However, the term "computer-readable medium" encompasses a single medium or multiple media, such as a centralized or distributed data store / center and / or associated caches or servers that store one or more instructions. The term "computer-readable medium" also encompasses any medium capable of storing, encoding, or retaining instructions for execution by a processor or that cause a computer system to implement any one or more of the embodiments disclosed herein.

[0157] The computer-readable medium may comprise one or more non-transitory computer-readable media and / or one or more transient computer-readable media. In one non-limiting embodiment, the computer-readable medium may comprise solid-state memory, such as a package such as a memory card containing one or more non-volatile read-only memories. The computer-readable medium may also comprise volatile rewritable memory, such as random access memory. The computer-readable medium may also include magneto-optical or optical media, such as disks, tapes, or other storage devices that collect carrier signals, such as signals transmitted over a transmission medium. Accordingly, the present disclosure should be construed to encompass all computer-readable media on which data or instructions may be stored, as well as other equivalents and successor media.

[0158] While this application describes particular embodiments that may be embodied as a computer program or code segments in a computer-readable medium, it should be understood that dedicated hardware implementations, such as application specific integrated circuits, programmable logic arrays, or other hardware devices, may also be constructed to perform one or more of the embodiments described herein. Potential applications of the various embodiments described herein may broadly include various electronic and computer systems. Thus, this application may encompass software implementations, firmware implementations, hardware implementations, or combinations thereof. Nothing in this application should be interpreted as being implemented or capable of being implemented solely in software rather than hardware.

[0159] Although components and functions that may be implemented in particular embodiments are described herein with reference to particular standards or protocols, the present disclosure is not limited to such standards or protocols. Such standards are superseded over time by faster or more efficient equivalents of essentially the same functionality. Accordingly, replacement standards or protocols having the same or similar functionality are considered equivalents.

[0160] The drawings of the embodiments described herein are intended to provide a basic understanding of each embodiment. The drawings are not intended to serve as a complete description of all of the features and components of an apparatus or system utilizing the structures and methods described herein. Numerous alternative embodiments may become apparent to those skilled in the art upon review of the present disclosure. Other embodiments may be derived and utilized from the present disclosure, and structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Additionally, the drawings are merely symbolic and may not be drawn to scale. Proportions in the drawings may be exaggerated or reduced. Therefore, the present disclosure and the drawings should be considered illustrative and not limiting.

[0161] Although one or more embodiments of the present disclosure may be referred to herein, singly and / or collectively, as the "present invention," this is done for convenience only and is not intended to intentionally limit the scope of the present application to any particular invention or inventive concept. Furthermore, while specific embodiments are shown and described herein, it should be understood that the specific illustrated embodiments may be replaced by any subsequent arrangements designed to achieve the same or similar purposes. The present disclosure is intended to encompass any and all subsequent adaptations and variations of each embodiment. Combinations of the above embodiments, as well as other embodiments not specifically described herein, will be apparent to those skilled in the art upon review of this specification.

[0162] The Abstract of the Disclosure is submitted with the understanding that it will not interpret or limit the scope or meaning of the claims. Moreover, the foregoing Detailed Description may group or describe various features in a single embodiment for the purpose of streamlining the disclosure. This disclosure should not be interpreted as reflecting an intention that the recited embodiments in each claim require more features than are expressly recited in each claim. Rather, as the appended claims suggest, the inventive subject matter may be directed to fewer components than are disclosed for any embodiment. Accordingly, the appended claims are hereby incorporated into the Detailed Description, with each claim defining independent subject matter in its own right.

[0163] The foregoing disclosed subject matter is illustrative and not limiting, and the appended claims are intended to embrace all such modifications, improvements, and other embodiments that fall within the true spirit and scope of the present disclosure. Accordingly, to the maximum extent permitted by law, the scope of the present disclosure shall be determined by the broadest interpretation of the appended claims and their equivalents, and shall not be limited or restricted by the foregoing detailed description. The present invention includes the following embodiments. [Aspect 1] 1. A method for automatically predicting and analyzing recovery time to a reliable state of an application, infrastructure, or business process using real-time telemetry, using one or more processors and one or more memories, comprising: establishing communication links between a plurality of data sources and an event bus; implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure or business process from the plurality of data sources via the event bus; using the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determining, based on the probability data, the relative risk and impact of the amount of data loss that may occur in the event of a destructive attack requiring the infrastructure or the application or applications supporting the business process to be rebuilt and / or restored; dynamically providing an estimate of a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss; A method comprising: [Aspect 2] 2. The method of claim 1, wherein the plurality of real-time telemetry data includes application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence data for calculating recovery capability and / or recovery indicators of critical infrastructure or business processes and services. [Aspect 3] The method of embodiment 2 further comprises: monitoring said critical infrastructure or said business processes and services to identify potential business risks; wherein the expected recovery time does not match the actual recovery time due to regulations, business objectives, technology design, or suboptimal business processes. [Aspect 4] The method of embodiment 1, further comprising: automatically detecting the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model; performing automatic repairs and / or modifications based on the inferred data in response to the destructive attack event; A method comprising: [Aspect 5] The method of embodiment 4, further comprising: initiating a fully autonomous recovery process for the affected environment in response to the destructive attack event; A method comprising: [Aspect 6] In the method of aspect 1, the event bus provides one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources to analyze and execute remediation activities in real time. A method configured to [Aspect 7] The method according to aspect 1, further comprising, when receiving the plurality of real-time telemetry data, The process of dynamically tracking information asset data within an organization; dynamically tracking the application or applications within the organization; dynamically tracking resilience and recovery attribute data of technology products used by said organization; A method comprising: [Aspect 8] 2. The method of claim 1, further comprising: determining the associated risk and the impact of the amount of data loss; A process for recording pending infrastructure or business-related changes, affected assets, implementation timeframes, and an assessment of human-induced risks and impacts; A method comprising: [Aspect 9] 2. The method of claim 1, wherein the machine learning model is configured to enter an alert state, determine the capacity requirements and / or availability of alternative affected infrastructure, and initiate orchestration of recovery, the total recovery time, and the amount of data loss. [Aspect 10] The method of embodiment 1, further comprising: automatically providing a simulation result of the recovery; automatically scheduling recovery tests to verify the simulation results; A method comprising: [Aspect 11] 2. The method of claim 1, wherein the machine learning model is further configured to, when a potential availability incident is detected, identify a business process or service to calculate a potential business risk to the entire organization or an organizational unit, and launch an orchestration engine to perform recovery. [Aspect 12] The method of embodiment 1, further comprising: providing a centralized repository of historical recovery incident data as an input for said probability data used in predictive recovery analysis; Equipped with the centralized repository maps application inventory to business operations modules and determines upstream and / or downstream availability impacts to improve decision making; The method, wherein the business operations module is configured to identify business processes, services, products, and their importance to the organization's business operations. [Aspect 13] The method of embodiment 1, further comprising: performing one or more of a recovery process to rebuild the application or applications, a recovery process to a last known trusted state of the application or applications, and a simulation / testing of the recovery process; A method comprising: [Aspect 14] A system that uses real-time telemetry to automatically predict and analyze the recovery time to a reliable state of an application, infrastructure, or business process, a processor; a memory operatively connected to the processor via a communications interface and storing computer-readable instructions; Equipped with The instructions, when executed, cause the processor to: establishing communication links between a plurality of data sources and an event bus; implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure or business process from the plurality of data sources via the event bus; using the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determining, based on said probability data, the relative risk and impact of the amount of data loss that may occur in the event of a destructive attack requiring the infrastructure or said application or applications supporting said business processes to be rebuilt and / or restored; and dynamically providing an estimate of a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss; A system that executes the following. [Aspect 15] In the system described in aspect 14, the plurality of real-time telemetry data includes application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence data for calculating recovery capability and / or recovery indicators of critical infrastructure or business processes and services. [Aspect 16] In the system of aspect 15, the processor is further configured to monitor the critical infrastructure or the business processes and services to identify potential business risks, where the expected recovery time does not match the actual recovery time due to regulations, business objectives, technical design, or suboptimal business processes. [Aspect 17] In the system described in aspect 16, the processor is further configured to use the machine learning model to automatically detect the availability incident based on the received plurality of real-time telemetry data and the historical event data, and to perform automatic functional repairs and / or functional modifications based on the inferred data in response to the destructive attack event. [Aspect 18] 18. The system of claim 17, wherein the processor is further configured to initiate a fully autonomous recovery process for the affected environment in response to the destructive attack event. [Aspect 19] In the system described in aspect 18, the event bus is configured to analyze and execute recovery activities in real time by providing one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources. [Aspect 20] a non-transient computer-readable medium configured to store instructions for automatically predictively analyzing recovery time to a reliable state of an application, infrastructure, or business process through real-time telemetry; The instructions, when executed, cause the processor to: establishing communication links between a plurality of data sources and an event bus; implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure or business process from the plurality of data sources via the event bus; using the machine learning model to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data; determining, based on said probability data, the relative risk and impact of the amount of data loss that may occur in the event of a destructive attack requiring the infrastructure or said application or applications supporting said business processes to be rebuilt and / or restored; and dynamically providing an estimate of a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss; A medium that allows the execution of the following:

Claims

1. 1. A method for automatically predictively analyzing recovery time to a reliable state of an application or infrastructure using real-time telemetry, using one or more processors and one or more memories, the method comprising: realizing or implementing a Smart Simulation and Recovery Module (SSRM) that automatically estimates the total recovery time and data loss that may occur in recovering an application or infrastructure in the event of a destructive malware attack or an event resulting from a destructive malware attack, wherein the SSRM has a communication module, an implementation module, a prediction module, a calculation module, a detection module, and an execution module, each of which is invoked by a corresponding application programming interface (API); establishing communication links between a plurality of data sources and an event bus by invoking the communication module via a first application programming interface (API); implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure from the plurality of data sources via the event bus by invoking the implementation module via a second application programming interface (API); calling the prediction module via a third application programming interface (API) to automatically predict probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model; automatically detecting the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model by invoking the detection module via a corresponding application programming interface (API); and determining, based on the probability data, the associated risk and impact of the amount of data loss that may occur if the application or applications supporting the infrastructure must be rebuilt and / or restored in the event of a destructive malware attack by invoking the calculation module via a fourth application programming interface (API). dynamically providing a total recovery time and / or total rebuild time estimate for the application or group of applications based on determining the associated risk and impact of the amount of data loss by invoking the execution module via a fifth application programming interface (API); detecting the destructive malware attack event by invoking the detection module; performing automated repairs and / or modifications based on the inferred data in response to detecting the destructive malware attack event to rebuild and / or restore the application or applications supporting the infrastructure; A method comprising:

2. 10. The method of claim 1, wherein the plurality of real-time telemetry data includes application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence data for calculating resilience and / or recovery metrics of critical infrastructure or business processes and services.

3. The method of claim 1 further comprising: Initiating a fully autonomous remediation process for the affected environment in response to the destructive malware attack event; A method comprising:

4. 10. The method of claim 1, wherein the event bus is configured to analyze and execute remediation activities in real time by providing one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources.

5. 2. The method of claim 1, further comprising, when receiving the plurality of real-time telemetry data: The process of dynamically tracking information asset data within an organization; dynamically tracking the application or applications within the organization; dynamically tracking resilience and recovery attribute data for technology products used by said organization; A method comprising:

6. 10. The method of claim 1, further comprising: determining the associated risk and the impact of the amount of data loss; A process for recording pending infrastructure or business-related changes, affected assets, implementation timeframes, and an assessment of human-induced risks and impacts; A method comprising:

7. 10. The method of claim 1, wherein the machine learning model is configured to enter an alert state, determine capacity requirements and / or availability of alternative affected infrastructure, and initiate orchestration of recovery, the total recovery time, and the amount of data loss. There is a way.

8. The method of claim 1 further comprising: automatically providing a simulation result of the recovery; automatically scheduling recovery tests to verify the simulation results; A method comprising:

9. 10. The method of claim 1, wherein the machine learning model is further configured to, when a potential availability incident is detected, identify business processes or services to calculate potential business risks to the entire organization or organizational units, and invoke an orchestration engine to perform recovery.

10. The method of claim 1 further comprising: providing a centralized repository of historical recovery incident data as an input for said probability data used in predictive recovery analysis; Equipped with The centralized repository maps application inventory to business operations modules and determines upstream and / or downstream availability impacts for improved decision-making. and The method, wherein the business operations module is configured to identify business processes, services, products, and their importance to the organization's business operations.

11. The method of claim 1 further comprising: performing one or more of a recovery process to rebuild the application or applications, a recovery process to a last known trusted state of the application or applications, and a simulation / testing of the recovery process; A method comprising:

12. 1. A system for automatically predicting and analyzing recovery time to a reliable state of an application or infrastructure through real-time telemetry, comprising: a processor; a memory operatively connected to the processor via a communications interface and storing computer-readable instructions; Equipped with The instructions, when executed, cause the processor to: a method for realizing or implementing a Smart Simulation and Recovery Module (SSRM) that automatically estimates the total recovery time and amount of data loss that may occur in recovering an application or infrastructure that requires rebuilding and recovering an application or group of applications comprising information technology in the event of a destructive malware attack or the event of such an attack, wherein the Smart Simulation and Recovery Module (SSRM) has a communication module, an implementation module, a prediction module, a calculation module, a detection module, and an execution module, each of which is invoked by a corresponding application programming interface (API); establishing communication links between a plurality of data sources and an event bus by invoking the communication module via a first application programming interface (API); implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure from the plurality of data sources via the event bus by invoking the implementation module via a second application programming interface (API); automatically predicting probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model by invoking the prediction module via a third application programming interface (API); automatically detecting the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model by invoking the detection module via a corresponding application programming interface (API); and invoking the calculation module via a fourth application programming interface (API) to determine, based on the probability data, the associated risk and impact of the amount of data loss that may occur if the application or applications supporting the infrastructure must be rebuilt and / or restored in the event of a destructive malware attack. dynamically providing estimated data for a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss by invoking the execution module via a fifth application programming interface (API); detecting the destructive malware attack event by invoking the detection module; and performing automated repairs and / or modifications based on the inferred data in response to detecting the destructive malware attack event to rebuild and / or restore the application or applications supporting the infrastructure; A system that executes the following.

13. 13. The system of claim 12, wherein the plurality of real-time telemetry data includes application data, asset data, data recovery data, business process data, technical product data, application transaction data, and threat intelligence data for calculating resilience and / or recovery metrics for critical infrastructure or business processes and services.

14. 13. The system of claim 12, wherein the processor is further configured to initiate a fully autonomous remediation process for the affected environment in response to the destructive malware attack event.

15. 13. The system of claim 12, wherein the event bus is configured to analyze and execute remediation activities in real time by providing one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources.

16. a non-transient computer-readable medium configured to store instructions for automatically predictively analyzing recovery time to a reliable state of an application or infrastructure through real-time telemetry; The instructions, when executed, cause the processor to: a method for realizing or implementing a Smart Simulation and Recovery Module (SSRM) that automatically estimates the total recovery time and amount of data loss that may occur in recovering an application or infrastructure that requires rebuilding and recovering an application or group of applications comprising information technology in the event of a destructive malware attack or the event of such an attack, wherein the Smart Simulation and Recovery Module (SSRM) has a communication module, an implementation module, a prediction module, a calculation module, a detection module, and an execution module, each of which is invoked by a corresponding application programming interface (API); establishing communication links between a plurality of data sources and an event bus by invoking the communication module via a first application programming interface (API); implementing a machine learning model configured to receive a plurality of real-time telemetry data and historical event data associated with an application or applications supporting an infrastructure from the plurality of data sources via the event bus by invoking the implementation module via a second application programming interface (API); automatically predicting probability data of an availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model by invoking the prediction module via a third application programming interface (API); automatically detecting the availability incident based on the received plurality of real-time telemetry data and the historical event data using the machine learning model by invoking the detection module via a corresponding application programming interface (API); and invoking the calculation module via a fourth application programming interface (API) to determine, based on the probability data, the associated risk and impact of the amount of data loss that may occur if the application or applications supporting the infrastructure must be rebuilt and / or restored in the event of a destructive malware attack. dynamically providing estimated data for a total recovery time and / or a total rebuild time for the application or group of applications based on determining the associated risk and impact of the amount of data loss by invoking the execution module via a fifth application programming interface (API); detecting the destructive malware attack event by invoking the detection module; and performing automated repairs and / or modifications based on the inferred data in response to detecting the destructive malware attack event to rebuild and / or restore the application or applications supporting the infrastructure; A medium that allows the execution of the following:

17. 17. The non-transitory computer-readable medium of claim 16, wherein the instructions, when executed, cause a processor to perform a procedure for initiating a fully autonomous recovery process for an affected environment in response to the event of a destructive malware attack.

18. 17. The non-transient computer-readable medium of claim 16, wherein the event bus is configured to analyze and execute remediation activities in real time by providing one or more communication channels between technical asset data sources, application transaction data sources, and threat intelligence data sources.

Citation Information

Patent Citations

  • Troubleshooting supporting system

    JP2018136656A

  • Failure handling program and failure handling method

    JP2019215813A

  • Platform for facilitating development of intelligence in an industrial internet of things system

    US20200348662A1