Ransomware mitigation using versioning and entropy increment-based recovery

By identifying multiple versions in the data repository and utilizing entropy increment calculations, the unencrypted data version can be identified and recovered, solving the challenge of data recovery in ransomware attacks, achieving fast and reliable data recovery, and reducing business interruption.

CN121195259APending Publication Date: 2025-12-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480034958.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-29
Filing Date
2024-06-13
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively recover encrypted data from ransomware attacks, leading to business disruptions and losses for enterprises.

Method used

By employing versioned storage and an entropy increment-based recovery method, rapid recovery is achieved by identifying multiple versions of data in the data repository and utilizing entropy increment calculations to identify the latest unencrypted version.

Benefits of technology

It enables rapid and reliable data recovery from ransomware attacks, reducing the impact on business operations and preventing data overwriting and malicious encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121195259A_ABST
    Figure CN121195259A_ABST
Patent Text Reader

Abstract

A mitigation system protects data in a data repository that has not been successfully encrypted by ransomware attacks from being encrypted. The data is stored in a data repository as a set of versions identified in the data tree, and the versions can only be updated by writing a new version to the tree. Access control that prevents tree modification is also in place. After the attack, a recovery function is performed to attempt recovery. The function calculates an entropy increment that compares the entropy of the encrypted version with the entropy of the data version that has not been encrypted. Based on the calculated entropy increment, the recovery function identifies a latest plaintext version of the data, and then initiates a recovery operation for that version.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] BACKGROUND TECHNICAL FIELD

[0002] The present disclosure relates generally to protecting data and, in particular, to a system that is able to mitigate ransomware attacks and recover from ransomware that is successful in encrypting data critical to the operation of an enterprise. BACKGROUND

[0003] In today's e-commerce ecosystem, cyber threats pose real risks and sometimes introduce situations where enterprises cannot recover from attacks. Large and small enterprises are increasingly concerned about the risk that ransomware poses to business continuity. Ransomware is a type of malicious software (malware) that uses encryption to prevent victims from reading their own data. Typically, such attacks are performed on-site at the victim and affect local data, for example, in a data center. Once encrypted, the attacker demands a ransom in exchange for decryption instructions and keys.

[0004] Despite advances in the ability to identify and possibly prevent ransomware attacks, there remains a real threat that ransomware can manage to encrypt some target data. Therefore, there remains a need to provide technology that enables recovery from a ransomware attack that has successfully encrypted data of a target entity. SUMMARY

[0005] The present disclosure provides a technique to mitigate a ransomware attack that has successfully encrypted some portion of data of a target entity. In a system implementing the method, the data of the entity has been stored as a set of versions in a data store, with versions identified in a data tree. The versioning functionality also enforces one or more access controls at defined points in the data tree to ensure that the tree cannot be manipulated. According to a first aspect of the present disclosure, after a ransomware attack has occurred, an entropy-based recovery function is executed to attempt recovery from the ransomware attack. To this end, the recovery function computes an entropy delta that compares the entropy of a version encrypted by the ransomware with the entropy of one or more versions of data that have not been encrypted. The entropy delta computation effectively exploits the attack, as the entropy of an encrypted version of data will necessarily be higher than the entropy of a plaintext version of that data. Based on the computed entropy delta, the recovery function then identifies the most recent plaintext version of data in the data store and then initiates a recovery operation with respect to that version.

[0006] According to another aspect of the disclosure, an apparatus includes a processor and a computer memory. The computer memory holds computer program instructions for execution by the processor to mitigate ransomware attacks. The computer program instructions include program code configured to perform a set of operations after a ransomware attack that has encrypted a set of versions identified in a data tree, where one or more versions of the set remain unencrypted. As described above, the operations compute an entropy delta that compares the entropy of the versions encrypted by the ransomware attack to the entropy of the one or more versions that remain unencrypted. Based on the computed entropy delta, a given version of the set of versions that remains unencrypted is identified, and then a recovery operation is initiated for the identified given version.

[0007] According to another aspect of the disclosure, a computer program product in a non-transitory computer readable medium is provided. The computer program product holds computer program instructions for execution by a processor in a host processing system configured to mitigate ransomware attacks. The computer program instructions include program code configured to perform operations such as the steps described above.

[0008] The mitigation method herein provides a reliable and effective technique that enables a target entity to recover quickly from any successful attack and with minimal disruption to its business. The method turns the attack against the adversary by using entropy computation to detect malicious data injection that corrupts the system state of the graph, enabling recovery to proceed efficiently. The entropy delta between data versions enables the system to determine (independent of time) what data has been encrypted by the ransomware, enabling recovery to proceed efficiently.

[0009] The foregoing has outlined some of the more pertinent features of the disclosed subject matter. These features should be construed to be merely illustrative. Many other beneficial results can be attained by applying the disclosed subject matter in a different manner or by modifying the subject matter as will hereinafter be described. BRIEF DESCRIPTION OF DRAWINGS

[0010] For a more complete understanding of the subject matter of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:

[0011] Figure 1 An exemplary block diagram of a data processing system in which illustrative embodiments of aspects can be implemented is described;

[0012] Figure 2 A representative cloud-based object storage system in which the techniques of the present disclosure can be practiced is depicted;

[0013] Figure 3 A high-level block diagram of a ransomware mitigation system of the present disclosure is depicted;

[0014] Figure 4a state of a version tree in a data store after a ransomware attack has occurred; and

[0015] Figure 5 A process flow depicting operation of a ransomware mitigation system of the present disclosure is depicted. DETAILED DESCRIPTION

[0016] Various aspects of the present disclosure are described in terms of narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program products (CPPs). With respect to any flowcharts, the operations might be performed in an order different than that shown in a given flowchart. For example, two operations shown in consecutive flowchart blocks might be performed in reverse order, as two separate integrated operations, as a single integrated operation, or at least partially contemporaneously with one another. In some embodiments, operations might be performed in an order different than that shown in a given flowchart.

[0017] A computer program product embodiment (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively including one or more sets of machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device that can hold and store instructions for use by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices including these media include: a magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or punch holes / platforms formed in the main surface of a disk, or any suitable combination of the foregoing. A computer-readable storage medium, as that term is used in the present disclosure, is not to be construed as being a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses through an optical fiber cable, electronic signals through a wire, and / or other transmission media. As those skilled in the art will readily appreciate, data is typically moved between storage devices at some incidental point in time during normal operation, such as during access, defragmentation, or garbage collection, but this does not render the storage device transitory, as the data is not transitory when it is stored.

[0018] The computing environment 100 contains an example of an environment for executing at least some of the computer code involved in performing the methods of the present invention, such as the ransomware mitigation code 200 of the present disclosure. In addition to the block 200, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 102, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a set of processors 110 (including processing circuitry 120 and a cache 121), a communication fabric 111, a volatile memory 112, a persistent storage 113 (including an operating system 122 and the block 200, as described above), a set of peripheral devices 112 (including a set of user interface (UI) devices 123, a storage 122, and a set of Internet of Things (IoT) sensors 125), and a network module 115. The remote server 102 includes a remote database 130. The public cloud 105 includes a gateway 120, a cloud coordination module 121, a set of host physical machines 122, a set of virtual machines 123, and a set of containers 122.

[0019] The computer 101 can take the form of a desktop computer, a laptop computer, a tablet computer, a smart phone, a smart watch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or hereafter developed that is capable of running a program, accessing a network, or querying a database such as the remote database 130. As is well known in the computer arts, and depending on the technology, the performance of the computer-implemented methods can be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of the computing environment 100, the detailed discussion is focused on a single computer, particularly the computer 101, to keep the presentation as simple as possible. The computer 101 can be located in the cloud, even though not shown in the cloud in Figure 1 On the other hand, the computer 101 need not be in the cloud, unless it can be positively indicated at any level.

[0020] The set of processors 110 includes one or more computer processors of any type now known or hereafter developed. The processing circuitry 120 can be distributed across multiple packages, such as multiple cooperating integrated circuit chips. The processing circuitry 120 can implement multiple processor threads and / or multiple processor cores. The cache 121 is a memory located in the processor chip package and is typically used for data or code that should be quickly accessible by threads or cores running on the set of processors 110. The cache is typically organized into multiple levels according to relative proximity to the processing circuitry. Alternatively, some or all of the cache of the set of processors can be located "off-chip." In some computing environments, the set of processors 110 can be designed to work with qubits and perform quantum computations.

[0021] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowchart and / or the narrative description of the computer-implemented method included in this document (collectively, the “method of the invention”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the method of the invention. In computing environment 100, at least some of the instructions for performing the method of the invention may be stored in persistent storage device 113 in block 200.

[0022] Communication structure 111 is a signal transmission path that allows various components of computer 101 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths that form buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.

[0023] Volatile memory 112 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 101.

[0024] The persistent storage device 113 is any form of non-volatile memory for a computer, now known or to be developed in the future. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 may be a read-only memory (ROM), but typically at least a portion of the persistent memory allows data to be written, deleted, and rewritten. Some common forms of persistent storage include hard disks and solid-state storage devices. The operating system 122 may take several forms, such as Linux, various known proprietary operating systems, or open-source portable operating system interface-type operating systems employing a kernel. The code included in block 200 generally includes at least some of the computer code involved in performing the methods of the present invention.

[0025] Peripheral device set 112 includes a set of peripheral devices for computer 101. Data communication connections between peripheral devices and other components of computer 101 can be implemented in various ways, such as Bluetooth connectivity, near field communication (NFC) connectivity, connections made by cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Memory 122 is external memory, such as an external hard drive, or pluggable memory, such as an SD card. Storage device 122 can be permanent and / or volatile. In some embodiments, storage device 122 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 requires a large amount of storage (e.g., where computer 101 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.

[0026] Network module 115 is a collection of computer software, hardware, and firmware that allows computer 101 to communicate with other computers via WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in network module 115.

[0027] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology known now or developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced by and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0028] End User Equipment (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in conjunction with computer 101. EUD 103 typically receives useful and helpful data from the operation of computer 101. For example, assuming computer 101 is designed to provide recommendations to the end user, these recommendations are typically transmitted from network module 115 of computer 101 to EUD 103 via WAN 102. In this way, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 can be client equipment, such as a thin client, heavy client, mainframe, desktop computer, etc.

[0029] Remote server 102 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 102 can be controlled and used by the same entity operating computer 101. Remote server 102 represents a machine that collects and stores useful and useful data for use by other computers such as computer 101. For example, if computer 101 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 101 from a remote database 130 of remote server 102.

[0030] Public cloud 105 is any computer system that can be used by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (especially data storage (cloud storage) and computing power) without direct active management by users. Cloud computing typically leverages resource sharing to achieve scalability consistency and economy. Direct and active management of the computing resources of public cloud 105 is performed by the computer hardware and / or software of cloud coordination module 121. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers constituting host physical machine set 122, which is the entire domain of physical computers in and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 123 and / or containers from container set 122. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud coordination module 121 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. Gateway 120 is a collection of computer software, hardware, and firmware that allows public cloud 105 to communicate via WAN 102.

[0031] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." New active instances of a VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically appear as actual computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a characteristic known as containerization.

[0032] Private cloud 106 is similar to public cloud 105, except that computing resources are available only to a single enterprise. While private cloud 106 is depicted as communicating with WAN 102, in other embodiments, private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardization or proprietary technology that enables coordination, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 105 and private cloud 106 are both part of a larger hybrid cloud. ransomware

[0033] The most common type of ransomware, called ransomware or crypto-ransomware, holds the victim's data hostage by encrypting it. The attacker then demands a ransom in exchange for the encryption key needed to decrypt the data. Crypto-ransomware begins by identifying and encrypting the files. Once the files are encrypted, the ransomware typically warns the victim of the infection via a .txt file placed on the computer desktop or through other notifications. This ransom note contains instructions on how to pay the ransom, usually in cryptocurrency or a similar untraceable method, in exchange for the decryption key or restoration of standard operations. Cloud object storage and versioning

[0034] For organizations looking to store their data, there are many different options available. Object storage is a primary storage solution used as a central storage platform for unstructured data in both cloud and on-premises solutions. Object storage is highly scalable, easy to manage, and allows users to balance storage costs, location, and compliance control requirements across datasets and core applications. Object storage is typically accessed directly from applications via a RESTful application programming interface (API). In this approach, objects are stored in a flat namespace, with all other objects in the same namespace. Object names are used to write to and read objects from the object storage, which often also provides the ability to add custom metadata to application data. This type of system typically provides a software-defined storage solution that operates on-premises and integrates with data at the edge, in the data center, or in a private cloud.

[0035] Figure 2 A representation of this type of cloud-based object storage system is depicted. In this example, the cloud-based object storage system is a database 230 located in a computing environment 200 including host 210 and network 240. Each of host 210 and database 230 may include one or more computer systems, such as... Figure 1The data processing system 100. Host 210 can be a standalone computing device, management server, web server, mobile computing device, or any other electronic device or computing system capable of receiving, sending, and processing data. In other embodiments, host 210 can represent a server computing system utilizing multiple computers as server systems, such as in a cloud computing environment (e.g., cloud computing environment 105 or 106). In some embodiments, host 210 includes application 212. In some embodiments, host 210 can send a request for a cloud service including credentials. Credentials may contain / include one or more attributes defining data related to the request. Attributes may relate to account, user, account level, time of request, location of request, etc.

[0036] Application 212 can be any combination of hardware and / or software configured to perform functions on a computing device (e.g., host 210). In some embodiments, application 212 is a web application. In some embodiments, application 212 can be configured to access / process data stored in database 230. Application 212 can be an account-based application. Users or user groups can use one or more accounts to access data. Access to processes / data can be based on one or more attributes of each account. In some embodiments, application 212 can be configured to run on a cloud computing network (e.g., Figure 1 One or more functions are performed on the cloud computing environment 105 or 106 shown. In some embodiments, application 212 may generate a request for a cloud service. The request may include a certificate consistent with the certificate of host 210. Application-based credentials may also exist.

[0037] Database 230 can be any combination of hardware and / or software configured to store data. Database 230 includes an access control system (ACS) 220, metadata 234, accessor 231, bucket 232, objects 233(1), objects 233(2), objects 233(3), and up to objects 233(n), where n can be an integer representing an object. For the purposes of this application, objects 233(1), objects 233(2), and up to objects 233(n) can be collectively, separately, and / or individually referred to as object 233. Thus, as depicted, database 230 includes object storage or an object-based storage architecture (or object storage). As described above, object storage is a data storage architecture that manages data as objects. Object storage can allow large amounts of unstructured data to be stored as a single unit / object.

[0038] ACS 220 can be any combination of hardware and / or software configured to control access to applications and / or data. The application can be application 212. Data can be stored in database 230. In some embodiments, ACS 220 can be defined and / or updated by the data owner. ACS 220 may include rules guiding how data and / or resources are utilized. In some embodiments, ACS 220 includes accounts 221, policies 222, roles 223, and an encryption engine 224. In some embodiments, ACS 220 can be fully integrated into database 230 and / or accessor 231. For example, the data owner can access database 230 to update and / or define security policies. This integration can provide many benefits of this disclosure. It eliminates the need for access to a remote ACS, which increases scalability costs and can potentially lead to bottlenecks and other availability and service issues.

[0039] Account 221 can be a set of credentials used to perform functions and / or access data. Any number of accounts 221 can exist for ACS 220. New accounts can be created, and / or old accounts can be disabled / deleted. Each account in account 221 can be associated with a specific user, group, entity (e.g., enterprise), department, etc. Account 221 can be accessed by correctly entering the credentials associated with each account. In some embodiments, account 221 is associated with one or more attributes. In some embodiments, information related to the account / user can be used as one or more credentials.

[0040] Role 223 can be an attribute of account 221. There can be any number of roles within ACS 220. In some embodiments, each role 223 has a defined set of attributes / credentials. Attributes allow access to one or more databases, data types, functions, applications, etc. In some embodiments, each account 221 can be associated with at least one role 223. Depending on the configuration, an account 221 can have two or more roles or just a single role associated with the account. In some embodiments, roles are hierarchical. Each role 223 can have at least as many permissions / attributes as the roles below it. For example, there can be a basic user, a supervisory user, and an administrator. Administrators can perform all functions and access all data, supervisory roles, a subset of administrators, and a smaller subset of basic users and supervisory roles. In some embodiments, role 223 can be associated with actions / functions. For example, there can be one role for each type of data manipulation and / or application using that data. In some embodiments, roles can be associated with categories of users of the data. For example, a role can be a data owner and a data user, where the data owner can add and remove data, while the data user can only view data.

[0041] Policy 222 can be a rule / definition for managing data. Policy 222 can include any number of attributes. Policy 222 can control when and how data is accessed. Attributes can be based on one or more of the following: location, account, user, job title, time of day, date (e.g., weekends, holidays, and weekdays may have different attributes), intended use, access frequency, etc. In some embodiments, policy 222 can define attributes that must be present in the access request to allow access to the cloud system.

[0042] In some embodiments, there is overlap between role 223 and policy 222. In some embodiments, role 223 and policy 222 may be combined into a single attribute list. The attribute list may include each condition that an account / user can access various data. In some embodiments, attributes may be defined at any level of granularity.

[0043] The encryption engine 224 can be any combination of hardware and / or software configured to encrypt and decrypt data. In some embodiments, the encryption engine 224 may reside in the host 210 and / or database 230 along with the ACS 220. In some embodiments, the encryption engine 224 may use attribute-based encryption (ABE). ABE is a public-key encryption method where the key depends on the attributes of the requester.

[0044] Metadata 234 can be a set of metadata configured to encapsulate access control to data within database 230. In some embodiments, each level within database 230 may include a unique set of metadata 234. For example, metadata may exist separately for database 230 as a whole, for bucket 232, and for each of objects 233. In some embodiments, metadata 234 may be incorporated into one or more buckets 232, objects 233, or any other level / function within database 230. In other words, each component within database 230 may include a set of metadata 234 specifically configured for the components to which it is attached. Metadata may include policy attributes that define how access to resources is granted.

[0045] Accessor 231 can be any combination of hardware and / or software configured to allow access to data within database 230. In some embodiments, accessor 231 can compare received credentials with an access request to determine if access is valid, wherein the access request is for a requested set of data / function access attributes. This comparison includes / is performed by an authorization function. If all attributes match the credentials, access can be granted. In some embodiments, if any attribute does not match, access can be denied. In some embodiments, accessor 231 includes an authorization function 235. In some embodiments, accessor 231 can identify one or more target resources based on the request and identify corresponding attributes and / or authorization functions.

[0046] Authorization function 235 can be any combination of hardware and / or software configured to determine whether a request will be allowed or denied. In some embodiments, authorization function 235. In some embodiments, accessor 231 includes one or more authorization functions 235. For each unique set of attributes defining access to any resource in database 230, there can be a unique authorization function. Thus, more than one authorization function can be used to obtain access to a single object. For example, there can be a unique set of attributes required to access database 230, a second set for bucket 232, and a third set for objects 233. Accessor 231 can perform three comparisons of a certificate against the individual authorization functions, where each set of attributes corresponds to a separate authorization function. In some embodiments, two or more resources can share the same set of attributes. For example, a first bucket and a second bucket can use the same set of attributes and authorization functions. In another example, a bucket and one or more objects within a bucket can share the same set of attributes.

[0047] Bucket 232 can be a data organization tool. In some embodiments, database 230 may include one or more buckets 232. In some embodiments, bucket 232 may be a container for storing one or more data objects. Bucket 232 may be defined based on one or more attributes. Attributes can be used to allow / disallow execution / access to objects within the bucket. This may include permissions for adding and / or deleting objects from bucket 232. In some embodiments, attributes are included in the metadata of bucket 232. Metadata can be received from ACS 220. Metadata may be based on one or more of account 221, policy 222, and role 223.

[0048] Object 233 can be a collection of data representing files. Each object may include a unique identifier, a certain amount of metadata, and the data that forms the object. Objects can be any type of file (e.g., word processing documents, spreadsheets, photos, videos, etc.). Each object 233 may include attribute metadata. Attribute metadata can be received from ACS 220. Attribute metadata can define when access to data in the object is permitted. Metadata can be based on one or more of account 221, policy 222, and role 223.

[0049] Object storage systems, such as those described above, can provide versioning. In such systems, versioning allows multiple revisions of a single object to exist within the same bucket. Each version of an object can be queried, read, restored from its archived state, or deleted. Enabling versioning on a bucket can mitigate data loss due to user error or unintentional deletion. Generally, in data storage systems that allow versioning (whether object storage or others), versions are identified in a version tree, and objects can only be updated by creating new versions; deleting objects from the version tree is not permitted. Any new data write to the system creates only one or more new versions, and those new versions are written to the data repository by manipulating the version tree. Ransomware mitigation systems using versioning and entropy increment-based recovery.

[0050] With the foregoing background, the technology of the present invention will now be described. In a representative embodiment, such as Figure 3 As shown, the ransomware mitigation system 300 includes a set of advanced functions, namely a versioning mechanism 302 and a recovery mechanism 304. Typically, each mechanism is implemented as computer software, i.e., a set of computer program code instructions executable by one or more hardware processors. Mechanisms 302 and 304 are depicted as separate from each other, but this is not necessary, as the functions described below can be integrated into a single mechanism, or one or both functions can be implemented within or associated with other systems, devices, processes, or programs within the computing system, practicing the techniques described herein. In this representative embodiment, the ransomware mitigation system 300 is associated with... Figure 2 The description herein relates to the cloud-based object storage solutions described above. However, this implementation is not intended to be restrictive, as ransomware mitigation systems (or their functionality) can be implemented in any type of data repository, whether on-premises or cloud-based, and for any type of versioned storage solution, whether object-based, block-based, or file-based.

[0051] The method presented in this paper assumes that a ransomware attack has occurred or is occurring, and that a portion of the data to be protected has been encrypted. As mentioned above, the system's goal is to mitigate the attack. Versioning mechanism 302 provides the first line of defense against ransomware. Specifically, and as previously mentioned, it is assumed that the data to be protected is stored in a data repository using versioning. One type of versioning is combined above... Figure 2 The description uses cloud-based object storage as an example, but this type of bucket-based object versioning is merely representative. The content constituting a "version" will typically vary depending on the nature of the data and the type of data repository. In a versioned data repository, the data in question (e.g., objects, collections of objects, files, collections of files, data blocks, collections of data blocks, etc.) is stored as a collection of one or more versions, each identified by a version identifier.

[0052] A collection of one or more versions is identified by a version tree 303. In cloud-based data repositories, and as mentioned above, versions are stored and accessible, typically using REST-based APIs. In the case of object storage, each time an object is modified and transferred to object storage, the new version of the object is identified in the version tree 303. Therefore, Figure 3 A version tree 303 is described to identify multiple versions of data. The version tree is sometimes more generally referred to herein as a data tree. Generally, and in systems implementing the techniques of this paper, the data repository storing the data is running versioning, and the way objects identified in the tree are updated is by creating a new version of them (and preferably, objects are not allowed to be deleted from the tree). Any new data write to the system creates only one or more new versions, and those new versions must be written to the data repository by manipulating the version tree. Versioning ensures that an adversary cannot simply replace the data version identified in the tree.

[0053] Therefore, versioning is implemented in the data repository before a ransomware attack occurs. After a ransomware attack, versioned data (i.e., unencrypted data) is not overwritten. Specifically, with versioning running, new writes only create new versions, thus ensuring that unaffected data remains protected from attack.

[0054] In particular, Figure 3The data tree in the diagram describes a set of data versions. Here, a ransomware attack has occurred, so some versions 305 of the data are shown as encrypted by the ransomware, while a set of versions 307 in the tree are shown as plaintext, i.e., unencrypted. In this example, the "plaintext" versions are those corresponding to the data writes after the ransomware attack, and as mentioned above, versioning protects the good data from being overwritten. Versions encrypted by ransomware are sometimes referred to as "bad" in this document, while unencrypted versions are sometimes referred to as "good".

[0055] The 302 versioning mechanism not only enforces that any new data write will generate a new version, but it also enforces one or more access controls designed to further prevent adversaries from overwriting good data (good versions) in the data repository with their maliciously encrypted code (bad versions). As with the ongoing versioning process, access controls to prevent the version tree from being modified are also in place (regardless of whether a ransomware attack has occurred). Figure 2 In the cloud-based object storage described herein, this functionality can be implemented in an access control system (ACS). Although access control is implemented in association with versioning functionality in the described embodiments, this is not a limitation, as existing access control systems, devices, programs, or processes can be redesigned for this functionality.

[0056] The following provides additional details regarding access control for versioning mechanism 302, which imposes several restrictions on the version tree itself. As already noted, and prior to a ransomware attack, the use of versioning restricts the storage system's ability to update or replace versions already identified on the tree. Furthermore, access control is also in place (at all times) to prevent modification of the version tree. This restricts the accessor's ability to write a new version of data to the data repository to reach one or more points that an adversary might attack. It also ensures that the write functionality of this new version cannot be exploited by ransomware (even at those points). As mentioned above, access control preventing modification of the tree is always in place. In one embodiment, the system locks away any credentials with the ability to manipulate the version tree from standard users or automated processes with the ability to write to the data repository. One way to achieve this is to use role-based access control (RBAC) based on the version tree. RBAC is a mechanism that grants users access to specific resources based on the roles they play in a protected environment. Here, versioning mechanism 302 is configured to enable authorized entities to supply RBAC on version tree 303, and specifically, to supply RBAC on their new version write functionality. Then, for any accessor attempting to write a new version to the data repository after a ransomware attack is detected, version tree write RBAC is implemented. Therefore, this method assumes that the only way to write a new version is by manipulating the version tree. Other types of access control that can be used for this purpose include Access Control Lists (ACLs), Attribute-Based Access Control (ABAC), object locking, etc. Alternatively, or to enhance access control, versioning mechanisms can also enforce other types of security policies or controls that restrict the ability to modify the version tree. Unrestrictedly, one such policy is a multi-factor modification feature requiring cross-account collaboration to manipulate the version tree. For example, suppose the version tree, or parts thereof, is encrypted and must first be decrypted before it can be changed. In this case, the multi-factor modification feature implements a secret-sharing scheme, where multiple factors need to be shared with a secret key (decryption key) before the key can be applied. The latter approach ensures that no single actor (e.g., ransomware) can modify the version tree by writing a new infected version. Regardless of how version tree access control is implemented, this method protects the version tree from manipulation (and therefore from writing new versions) after a ransomware attack is detected. In other words, versioning mechanisms enable the mitigation system to protect data that has not yet been encrypted by ransomware from being overwritten.

[0057] Once the unencrypted data is protected through access control operations using the versioning mechanism described above, the system mitigates the risk by using recovery mechanism 304 to recover data (or portions thereof) encrypted by ransomware. Of course, the size and scale of the data—whether from a small or large enterprise—can easily be far greater than what a large number of people could categorize, let alone in a timely manner. In reality, data recovery may take longer than the enterprise can afford. Therefore, to enhance the versioning functionality described above, recovery mechanism 302 provides a reliable way for targets of ransomware attacks to recover efficiently from an attack. To this end, the recovery mechanism performs an "entropy-based" analysis of good and bad versions of the data in the data tree to attempt to locate the most recent (most recent) "good" version of the data. Typically, the latest good version is the version with the latest clean data. Once this latest good version is located, the system preferably performs version recovery from that point in the version tree. Specifically, recovery can involve one of several different techniques, such as copying the latest good version of the data to a new location, moving the top of the version tree to the identified latest good version, clearing the version tree of the encrypted data versions (leaving only the good versions), or some combination thereof.

[0058] Entropy-based analysis preferably works as follows. With a brief background, in information theory, the entropy of a random variable is the average level of "information" inherent in the possible outcomes of the variable. Thus, for example, in the context of a data segment transmitted over a network, the information value of that data depends on how surprising its content is. If a highly probable event occurs regarding the transmission of that data, the data carries very little information; conversely, if a highly improbable event occurs, the data provides more information. Entropy represents a measure of the information content of data. In the context of this disclosure, the recovery mechanism utilizes entropy analysis, which exploits the adversary by leveraging the mechanisms used to attack the data in the first instance. In particular, one of the main functions of ransomware encryption is to "randomize" the data, making it unrecognizable. Thus, the encrypted version of a data segment shows a difference in entropy compared to the unencrypted form of the same (or similar) data segment. In fact, the encrypted version of a given data exhibits a higher entropy than the plaintext entropy of that data. By comparing the relative entropies of the data versions in the version tree, the recovery mechanism 304 determines the last good version of the data retained after the ransomware attack. As described above, once this determination is made, recovery can be performed from the latest identified plaintext version.

[0059] This operation is in Figure 4 The diagram describes a data repository 400 with a version tree 403, which has a set of versions. This corresponds to the above in... Figure 3The version tree described herein, data repository 400 has used versioning and associated access controls to prevent data from being overwritten in the manner described above. In this case, after a ransomware attack, the data tree has encrypted (bad) version 405 and unencrypted (good) version 407. As reflected, and based on entropy analysis, bad version 405 has high entropy, and the plaintext version has relatively low entropy. In operation, the recovery mechanism performs a scanning function to determine the relative entropy of version(s)(s) in the tree. In one embodiment, the scanner identifies encrypted version 405 with the highest entropy value, and then compares that value with the entropy calculated for each unencrypted (plaintext) version 407; then, the plaintext version with the lowest relative entropy (compared to the encrypted version with the highest entropy) is determined to be the latest plaintext version. This is as follows: Figure 4 Version 409 is described in the text.

[0060] One measure of relative entropy that can be used in this paper is the Kullback-Leibler divergence (also known as relative entropy and I divergence), denoted as... Furthermore, it is a statistical distance, a measure of how a probability distribution P differs from a second reference probability distribution Q. Other techniques for comparing version entropy include calculating cross-entropy values, calculating mutual information, etc. As described above, and once the system identifies the latest plaintext version 409 from entropy increment analysis, it retrieves that version from the data repository and places it back into the target environment, where it can then be used.

[0061] Versioning and recovery features provide multi-layered (or multi-stage) defense against ransomware attacks, and the solution addresses the question of how best to recover from such an attack once it is detected. Figure 5 The processing flow of a ransomware mitigation system in operation is described. Versioning is in progress (preferably throughout the data's lifecycle) prior to any attack (or any detection of such an attack), and access controls to prevent modification of the version tree are in place. This is step 500. In step 502, a ransomware attack is indicated as having occurred. Specific methods or mechanisms for detecting ransomware attacks are not an aspect of this disclosure, as many known techniques and skills exist for performing this function. Therefore, in a typical embodiment, the mitigation system of this disclosure operates in association with cybersecurity solutions such as SOAR, SIEM, EDR, or XDR, or some other security component or system. In step 504, a recovery function is activated to calculate the entropy increments of the various versions existing in the data tree. In step 506, the latest plaintext version of the data is identified. The process then continues at step 508 to initiate recovery from the identified latest plaintext version found based on the entropy increments. The specific manner in which recovery is then performed is not limiting, and one of many different recovery methods may be implemented relative to the version tree.

[0062] The aforementioned techniques offer significant advantages. By leveraging versioning and recovery capabilities together, the mitigation system provides a reliable and efficient technique that allows targets to recover quickly from any successful attack with minimal disruption to their operations. This method uses entropy calculations to detect malicious data injections attempting to corrupt the system's state, thus redirecting the attack towards the adversary. The entropy increments between data versions allow the system to determine (time-independently) what data has been encrypted by ransomware, enabling efficient recovery. Therefore, the method presented in this paper provides attack mitigation by utilizing versions that cannot be compromised by ransomware and are maintained by the system. Another benefit of the method presented in this paper is the ability to accurately recover objects without resorting to backup copies. In fact, using the method here, the most recent good copy still exists in the data repository.

[0063] More generally, in the context of the disclosed subject matter, each computing device is a data processing system comprising hardware and software (such as...). Figure 1 (As shown in the diagram), and these entities communicate with each other via networks such as the Internet, intranets, extranets, private networks, or any other communication medium or link. Applications on the data processing system provide native support for the Web and other known services and protocols, including but not limited to support for HTTP, FTP, SMTP, SOAP, XML, WSDL, UDDI, and WSFL. Information on SOAP, WSDL, UDDI, and WSFL is available from the World Wide Web Consortium (W3C), which develops and maintains these standards; further information on HTTP, FTP, SMTP, and XML is available from the Internet Engineering Task Force (IETF). Familiarity with these known standards and protocols is assumed.

[0064] Similarly, Figure 1 As shown, the solution described in this paper can be implemented in various server-side architectures or in combination with various server-side architectures, including simple n-tier architectures, portal websites, and federated systems. The techniques in this paper can also be implemented, either wholly or partially, in loosely coupled server (including cloud-based) environments.

[0065] More generally, the subject matter described herein may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment that includes both hardware and software elements. In a preferred embodiment, the functionality is implemented in software, including but not limited to firmware, resident software, microcode, etc. Furthermore, as described above, the analytics engine functionality may take the form of a computer program product accessible from a computer-usable or computer-readable medium, which provides program code used by or in conjunction with a computer or any instruction execution system. For the purposes of this specification, a computer-usable or computer-readable medium can be any means capable of containing or storing a program used by or in conjunction with an instruction execution system, apparatus, or device. This medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device). Examples of computer-readable media include semiconductor or solid-state memory, magnetic tape, removable computer disks, random access memory (RAM), read-only memory (ROM), hard disks, and optical discs. Current examples of optical discs include optical disc-read-only memory (CD-ROM), optical disc-read / write (CD-R / W), and DVDs. Computer-readable media are tangible articles.

[0066] In a representative embodiment, the ransomware mitigation system functions described above are implemented in a dedicated computer, preferably in software executed by one or more processors. The software is maintained in one or more data repositories or memories associated with one or more processors, and the software can be implemented as one or more computer programs. Commonly, the dedicated hardware and software comprise the system described above.

[0067] While a particular order of operations performed by certain embodiments of the disclosed subject matter has been described above, it should be understood that such an order is exemplary, as alternative embodiments may perform operations in a different order, combine certain operations, overlap certain operations, etc. References to a given embodiment in the specification indicate that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment may not necessarily include that particular feature, structure, or characteristic.

[0068] Finally, although the given components of the system have been described individually, those skilled in the art will understand that some functions can be combined or shared in a given instruction, program sequence, code section, etc.

[0069] The techniques described herein provide improvements to another technology or technology field, namely cybersecurity products, services, and solutions, as well as improvements to the operational capabilities of such systems when used in the manner described.

[0070] The nature of the data protected in a storage system depends on the application and is not intended to be restricted. Furthermore, as previously stated, the techniques described herein are not limited to any particular data repository or storage architecture.

[0071] The subject matter has been described, and the content for which protection is sought is as follows.

Claims

1. A method for mitigating ransomware attacks on data stored in a data repository, comprising: The data in the data repository is versioned into a set of versions identified in the data tree, wherein new data writes to the data repository create only one or more new versions; Following a ransomware attack that has encrypted one or more versions of the version set identified in the data tree, wherein one or more versions of the version set in the data tree remain unencrypted, the entropy increment is calculated by comparing the entropy of the versions encrypted by the ransomware attack with the entropy of the one or more versions that remain unencrypted. Based on the calculated entropy increment, identify the given versions in the version set that remain unencrypted; and Initiate a recovery operation for the identified given version.

2. The method according to claim 1, further comprising: Implement one or more access controls at one or more definition points in the data tree.

3. The method according to claim 1, wherein, The data repository is one of the following: a database, a file-based data repository, and a cloud object data repository.

4. The method according to claim 1, wherein, The identified version is the latest unencrypted version of the data.

5. The method according to claim 2, wherein, The access control is one of the following: role-based access control for the data tree, an access control list for the data tree, attribute-based access control for the data tree, and a policy that allows only a set of cooperating entities to manipulate the data tree.

6. The method according to claim 1, wherein, Initiating the recovery operation is one of the following operations: retrieving the identified given version to a new location in the data repository, adjusting the identified given version to the top of the data tree, and clearing the version tree of one or more encrypted versions.

7. The method according to claim 1, wherein, The entropy increment is calculated as a relative entropy.

8. An apparatus comprising: processor; A computer memory storing computer program instructions, which are executed by the processor to mitigate ransomware attacks against data stored as version sets in a data repository, the version sets being identified in a data tree, the computer program instructions including program code configured to perform the following operations: The data in the data repository is versioned into a set of versions identified in the data tree, wherein new data writes to the data repository create only one or more new versions; Following a ransomware attack that has encrypted one or more versions of the version set identified in the data tree, wherein one or more versions of the version set in the data tree remain unencrypted, the entropy increment is calculated by comparing the entropy of the versions encrypted by the ransomware attack with the entropy of the one or more versions that remain unencrypted. Based on the calculated entropy increment, identify the given versions in the version set that remain unencrypted; and Initiate a recovery operation for the identified given version.

9. The apparatus according to claim 8, wherein, The program code is further configured to implement one or more access controls at one or more definition points in the data tree.

10. The apparatus according to claim 8, wherein, The data repository is one of the following: a database, a file-based data repository, and a cloud object data repository.

11. The apparatus according to claim 8, wherein, The identified version is the latest unencrypted version of the data.

12. The apparatus according to claim 9, wherein, Regarding the data tree, the access control is one of the following: role-based access control for the data tree, an access control list for the data tree, attribute-based access control for the data tree, and a policy that allows only a set of cooperating entities to manipulate the data tree.

13. The apparatus according to claim 8, wherein, The program code configured to initiate the recovery operation includes: program code configured to perform one of the following: retrieving the identified given version to a new location in the data repository, adjusting the identified given version to the top of the data tree, and clearing the version tree of one or more encrypted versions.

14. The apparatus according to claim 8, wherein, The entropy increment is calculated as a relative entropy.

15. A computer program product in a non-transitory computer-readable medium, the computer program product holding computer program instructions that, when executed by a processor in a host processing system, mitigate ransomware attacks on data stored as a set of versions in a data repository, the set of versions being identified in a data tree, the computer program instructions including program code configured to perform the following operations: The data in the data repository is versioned into a set of versions identified in the data tree, wherein, Writing new data to the data repository only creates one or more new versions; Following a ransomware attack that has encrypted one or more versions of the version set identified in the data tree, wherein one or more versions of the version set in the data tree remain unencrypted, the entropy increment is calculated by comparing the entropy of the versions encrypted by the ransomware attack with the entropy of the one or more versions that remain unencrypted. Based on the calculated entropy increment, identify the given versions in the version set that remain unencrypted; and Initiate a recovery operation for the identified given version.

16. The computer program product according to claim 15, wherein, Implement one or more access controls at one or more definition points in the data tree.

17. The computer program product according to claim 15, wherein, The data repository is one of the following: a database, a file-based data repository, and a cloud object data repository.

18. The computer program product according to claim 15, wherein, The identified version is the latest unencrypted version of the data.

19. The computer program product according to claim 15, wherein, The access control is one of the following: role-based access control for the data tree, an access control list for the data tree, attribute-based access control for the data tree, and a policy that allows only a set of cooperating entities to manipulate the data tree.

20. The computer program product according to claim 15, wherein, The program code configured to initiate the recovery operation includes: program code configured to perform one of the following: retrieving the identified given version to a new location in the data repository, adjusting the identified given version to the top of the data tree, and clearing the version tree of one or more encrypted versions.

21. The computer program product according to claim 15, wherein, The entropy increment is calculated as a relative entropy.

22. A computing system for mitigating ransomware attacks, comprising: The mechanism includes a versioning mechanism for first program code executed on hardware, the versioning mechanism being configured to: store data as a set of versions identified in a data tree in a data repository, wherein versions in the set of versions can only be updated by writing a new version to the data tree, and to implement access control to prevent modification of the data tree; and A recovery mechanism including second program code executed on hardware, the recovery mechanism being configured to perform a recovery function that identifies the latest plaintext version of the data from the data tree and initiates a recovery operation on the latest version.

23. The computing system according to claim 22, wherein, The recovery function calculates an entropy increment, which compares the entropy of the version identified in the data tree and encrypted by the ransomware attack with the entropy of the one or more versions identified in the data tree that remain unencrypted, and identifies the latest plaintext version of the data based on the calculated entropy increment.

24. The computing system according to claim 23, wherein, The recovery operation is configured to: clear the version tree of the encrypted version.

25. A method of operating in a computing system, wherein, The method identifies data versions in a data repository based on a data tree, wherein a version can only be updated by writing a new version to the data tree. In response to a ransomware attack that has encrypted one or more versions in the set of versions identified in the data tree, an entropy increment is calculated, which is obtained by comparing the entropy of the version encrypted by the ransomware attack with the entropy of the one or more versions that remain unencrypted. Based on the calculated entropy increment, identify the latest version in the version set that remains unencrypted; and Initiate a restore operation for the aforementioned latest version.